What we measured, including the runs that went against us.
Write-ups of the benchmarks we run on our own products. Each one carries its method, its numbers, and the workloads where we come off worse than the thing we are compared against — on the same page as the result, not in a footnote. If a figure here is not one we measured, it is not here.
-
28 August 2026 · Governance
We tried to get past our policy gate. It failed 50 times.
A pre-inference refusal matters only if it survives writing intended to evade it. Across 2,251 adversarial requests, fifty crossed our gate; five of fifty benign controls were refused. The post documents both errors, the fixes that changed the result, and what the study still cannot establish.
2,251 adversarial cases 50 crossed the gate 5 of 50 benign controls refusedRead the write-up → -
25 August 2026 · Benchmark
Compression was the easy part.
Six agent problems through four ways of handing work between agents. Tileward Context cut billable tokens by 72.1% against an uncompressed control and wrote the best code in the set — then came last of the four on research fan-out, the multi-agent case we sell it for. The cause is measured rather than guessed at.
−72.1% billable tokens 0.968 coding, best of four 5.33 research panel, last of fourRead the write-up →
This page exists because the next write-up is already being run, and because a result worth publishing should not have to wait for a place to put it. What lands here is whatever the harness produces — point it at a workload you actually run and we will publish what it does. Shorter measurements, and the limits on each, stay on Evals.