Each card turns a platform promise into a workload you can inspect: what it costs with and without Tileward, whether the output holds up, and where the current boundary sits. Start with the workload that looks like yours.
11,667 tokens of shared state over 21 artifacts. Tileward delivered
86% of the material each step declared it needed; hand-configured routing, which is a
person tagging artifacts per problem, delivered 95%. This flow
was graded twice, four hours apart: Tileward 0.881 then 0.952, the uncompressed control
0.952 then 1.000. The scores move further between runs than between arms, so treat the
token figures as the reliable half. The fuller run.
five parallel gatherers → synthesizer → critic → editor
Without Tileward364,306 tokens/run
Hand-configured routing65,513 tokens/run
With Tileward Context81,388 tokens/run
4.5×
Fewer tokens
283M
Tokens saved per 1,000 runs
8
Agents · 4 rounds
This is the flow we do worst on, and the saving is the largest.
45,712 tokens of shared state over 32 artifacts. Hand-configured routing delivered every
artifact each agent declared it needed; Tileward delivered 66%. Answers were not
graded on this flow, but a later run judged them on a blind panel and Tileward came last of
four. The whole result.
Agent backlog 104 AgentBench tasks with earlier work in history
3.2×
Less context than full history
37 / 24
Tasks solved, Context / full history
14–19
Tasks identical runs can flip
Tileward sent 523,028 context tokens against 1,651,372 for full histories. The context reduction is stable; the 13-task lead is not, so the repeated ordering matters more than the margin.
LOCOMO long conversations answerable questions, judged by a separate model
18.9×
Less context than full history
55%
Answered with Tileward Context
63–67%
Answered with full history
Tileward sent 1,499 tokens per question against 28,365 for the full transcript. This keeps most, not all, of the answer rate; a separate model judged every answer.
A recall upgrade, measured LongMemEval-S, our own run on a 150-question sample, self-judged
Before the upgrade67 of 140 · 47.9%
After the upgrade81 of 140 · 57.9%
+10.0
Points, paired A/B
14 / 0
Fixed / broken
9→16
LongMemEval-S temporal-reasoning, of 40
The arms differ by one upgrade to how Tileward Context recalls
conversations; store, questions, model and judge are identical. No question that was
right before went wrong after. Paired 95% interval on the overall gain: +5.0 to +15.0
points. The upgrade spends a little of the recall budget on itself, so the after arm
sometimes carries slightly less material and still wins.
Our own harness run of LongMemEval-S on Tileward’s recall path at default settings:
a single run of a 150-question sample, 140 scored identically in both arms.
Compressed model against the publisher’s own build
Publisher’s official FP835.7 GB · 82.6% MMLU
Tileward 35B-A3B24.5 GB · 82.6% MMLU
Community GPTQ Int423.3 GB · 83.3% MMLU
Community AWQ 4-bit24.3 GB · 83.2% MMLU
1.46×
Smaller on disk
±0.31
Error bar on each
Level with FP8, behind both community 4-bit builds. Re-run with both arms
in one session, the FP8 gap is not significant. GPTQ’s lead over us holds on a
paired test; AWQ’s does not have one.
Five-shot, 14,042 questions, lm-evaluation-harness, no sampling. The four bars are one run, one machine, same base model.
Quoted to the tenth: five runs of our own checkpoint spread 0.24, wider than most gaps here.
Bars are weights on disk; serving VRAM is 25.8 GB against FP8's 40.4 GB. MMLU is multiple choice and we lead on one of its four categories; matching FP8 is consistent with the builds behaving alike, not proof, and per-token agreement was not run.
Compressed model, in an agent loop four builds of one base, one machine
Qwen3.8-27B, the base (BF16)64 + 64 of 142 · 45.1%
Tileward-Qwen3.8-27b61 + 63 of 142 · 43.7%
Publisher’s official FP860 + 65 of 142 · 44.0%
Bonsai 2 27B, ternary57 + 64 of 142 · 42.6%
10 / 13
Tasks won / lost against the base
142
Tasks · 2 runs per build
A tie on finished tasks: all three compressed builds finish within noise of the base. The ternary build is the smallest file and works hardest for it.
Paired by task over two runs, Tileward-Qwen3.8-27b wins 10 and loses 13 against the BF16
base, a gap smaller than the test can see. The publisher’s FP8
build finishes 1.1 points fewer (11 won, 16 lost) and the ternary build 2.5 fewer (13 won,
19 lost); neither gap is one this test can call. Every compressed build spent more prompt
tokens than the base to get there: ours 7% more, the ternary build 28% more.
AgentBench OS: Linux tasks in a real container, pass or fail decided by the environment; the 142 tasks every arm got a turn on. Greedy decoding, thinking off, 32k context, one NVIDIA 80 GB GPU, the same tasks and settings for all four.
An agent loop does not repeat itself exactly, so 142 tasks resolve a gap of about ten points; every gap on this card is smaller than that, a direction rather than a size.
The three Qwen-format builds ran on vLLM; the ternary build ran on its publisher’s own server, the only one that loads it, with reasoning off like the others, so its publisher’s thinking-on figures do not compare to this row. On this Hopper-generation card the FP8 build ran its native 8-bit path.
Agent
flows measured on a 120B open-weight model, tokens counted by the serving model rather than estimated. A
provider that discounts a repeated prefix narrows what the saving is worth.
Bring your own workload.
Point the harness at something you actually run.
Talk to an engineer.