What we measured

Each card is one workload we ran, what it cost with and without Tileward, and whether the answers held up. Filter to the one that looks like yours.

4.3×
Fewer context tokens
$341.97
Saved per 1,000 runs
100%
Of required context still delivered
Agent workflowsToken costQuality

Engineering pipeline sequential hand-off

planner → builder-migration → builder-api → validator → tester → reviewer → release
Without Tileward80,031 · $0.080/run
With Tileward36,854 · $0.037/run
2.2×
Fewer tokens
$43.18
Saved per 1,000 runs
7
Agents · 5 rounds

11,667 tokens of shared state over 21 artifacts. Every artifact each agent was declared to need still arrived, and the routed run scored 0.95 on graded answers against 0.81 for the agent handed everything.

Agent workflowsToken cost

Research fan-out wide parallel gather

five parallel gatherers → synthesizer → critic → editor
Without Tileward364,306 · $0.364/run
With Tileward65,513 · $0.066/run
5.6×
Fewer tokens
$298.79
Saved per 1,000 runs
8
Agents · 4 rounds

45,712 tokens of shared state over 32 artifacts. Every required artifact still arrived. Answers were not graded on this flow.

Conversation memoryToken costQuality

Long conversation 142 turns, asked back

Whole transcript resent3,122 per question
With Tileward103–127 per question
25–30×
Fewer tokens
20 of 20
Answered, either way
0 of 20
Answered with no history

12 facts mentioned once each, 20 questions, 8 needing two or three facts at once. Qwen3.5-9B, exact substring match. With decomposition turned off it answers 16 of 20, so the question-splitting is doing that work rather than the retrieval.

Model compressionQuality

Compressed model against the publisher’s own build

Publisher’s official FP835.7 GB · 82.60% MMLU
Tileward 35B-A3B24.5 GB · 82.58% MMLU
Community GPTQ Int423.3 GB · 83.30% MMLU
Community AWQ 4-bit24.3 GB · 83.17% MMLU
1.46×
Smaller on disk
0.02
Points apart on MMLU
±0.31
Error bar on each
FieldTilewardGPTQAWQFP8
Humanities77.7778.3679.0278.47
Social sciences89.7690.0689.7089.37
STEM80.4081.7080.4979.58
Other84.9785.7185.7185.23

Level with FP8, behind both community 4-bit builds. The FP8 gap is inside the error bar, which is what a substitute should look like. GPTQ leads by 0.72 and AWQ by 0.59, at 1.4 to 1.7 times the combined error, so neither is settled.

  • Five-shot, 14,042 questions, lm-evaluation-harness, no sampling. Same base model, one machine.
  • Bars are weights on disk. Serving VRAM is 25.8 GB against FP8's 40.4 GB.
  • MMLU is multiple choice: no signal on generation, instruction following or long context.
  • Matching FP8 is consistent with the builds behaving alike, not proof. Per-token agreement not run.

Agent flows measured on gpt-oss-120b, tokens counted by the serving model rather than estimated. Costs are that model’s published rate of $1.00 per million tokens, on uncached serving: a provider that discounts a repeated prefix narrows the gap, because the run without Tileward sends every agent the same block.

Bring your own workload. Point the harness at something you actually run. Talk to an engineer.