Tileward / Evals
Evals

See what Tileward changes before you deploy it.

Each card turns a platform promise into a workload you can inspect: what it costs with and without Tileward, whether the output holds up, and where the current boundary sits. Start with the workload that looks like yours.

Workload
Benchmark
Measurement
3.6×
Fewer tokens
322M
Tokens saved per 1,000 runs of each flow
66–86%
Of required context delivered
Agent workflowsToken costQuality

Engineering pipeline sequential hand-off

planner → builder-migration → builder-api → validator → tester → reviewer → release
Without Tileward80,031 tokens/run
Hand-configured routing35,226 tokens/run
With Tileward Context40,486 tokens/run
2.0×
Fewer tokens
39.5M
Tokens saved per 1,000 runs
7
Agents · 5 rounds

11,667 tokens of shared state over 21 artifacts. Tileward delivered 86% of the material each step declared it needed; hand-configured routing, which is a person tagging artifacts per problem, delivered 95%. This flow was graded twice, four hours apart: Tileward 0.881 then 0.952, the uncompressed control 0.952 then 1.000. The scores move further between runs than between arms, so treat the token figures as the reliable half. The fuller run.

Write-up: Compression was the easy part →
Agent workflowsToken costQuality

Research fan-out wide parallel gather

five parallel gatherers → synthesizer → critic → editor
Without Tileward364,306 tokens/run
Hand-configured routing65,513 tokens/run
With Tileward Context81,388 tokens/run
4.5×
Fewer tokens
283M
Tokens saved per 1,000 runs
8
Agents · 4 rounds

This is the flow we do worst on, and the saving is the largest. 45,712 tokens of shared state over 32 artifacts. Hand-configured routing delivered every artifact each agent declared it needed; Tileward delivered 66%. Answers were not graded on this flow, but a later run judged them on a blind panel and Tileward came last of four. The whole result.

Write-up: Compression was the easy part, including this loss →
Agent workflowsAgentBenchToken costQuality

Agent backlog 104 AgentBench tasks with earlier work in history

3.2×
Less context than full history
37 / 24
Tasks solved, Context / full history
14–19
Tasks identical runs can flip

Tileward sent 523,028 context tokens against 1,651,372 for full histories. The context reduction is stable; the 13-task lead is not, so the repeated ordering matters more than the margin.

Report: Give agents useful context, not the whole backlog →
Conversation memoryLOCOMOToken costQuality

LOCOMO long conversations answerable questions, judged by a separate model

18.9×
Less context than full history
55%
Answered with Tileward Context
63–67%
Answered with full history

Tileward sent 1,499 tokens per question against 28,365 for the full transcript. This keeps most, not all, of the answer rate; a separate model judged every answer.

Report: Keep most of the answer rate on 19× less context →
Conversation memoryToken costQuality

Long conversation 142 turns, asked back

Whole transcript resent3,122 per question
With Tileward38–83 per question
38–82×
Fewer tokens
20 of 20
Answered, either way
0 of 20
Answered with no history

12 facts mentioned once each, 20 questions, 8 needing two or three facts at once. Qwen3.5-9B, exact substring match.

All write-ups →
Conversation memoryLongMemEvalQuality

A recall upgrade, measured LongMemEval-S, our own run on a 150-question sample, self-judged

Before the upgrade67 of 140 · 47.9%
After the upgrade81 of 140 · 57.9%
+10.0
Points, paired A/B
14 / 0
Fixed / broken
9→16
LongMemEval-S temporal-reasoning, of 40

The arms differ by one upgrade to how Tileward Context recalls conversations; store, questions, model and judge are identical. No question that was right before went wrong after. Paired 95% interval on the overall gain: +5.0 to +15.0 points. The upgrade spends a little of the recall budget on itself, so the after arm sometimes carries slightly less material and still wins.

  • Our own harness run of LongMemEval-S on Tileward’s recall path at default settings: a single run of a 150-question sample, 140 scored identically in both arms.
  • Tileward 35B-A3B wrote and judged every answer.
Report: Improve recall without losing correct answers →
Model compressionMMLUQuality

Compressed model against the publisher’s own build

Publisher’s official FP835.7 GB · 82.6% MMLU
Tileward 35B-A3B24.5 GB · 82.6% MMLU
Community GPTQ Int423.3 GB · 83.3% MMLU
Community AWQ 4-bit24.3 GB · 83.2% MMLU
1.46×
Smaller on disk
±0.31
Error bar on each

Level with FP8, behind both community 4-bit builds. Re-run with both arms in one session, the FP8 gap is not significant. GPTQ’s lead over us holds on a paired test; AWQ’s does not have one.

  • Five-shot, 14,042 questions, lm-evaluation-harness, no sampling. The four bars are one run, one machine, same base model.
  • Quoted to the tenth: five runs of our own checkpoint spread 0.24, wider than most gaps here.
  • Bars are weights on disk; serving VRAM is 25.8 GB against FP8's 40.4 GB. MMLU is multiple choice and we lead on one of its four categories; matching FP8 is consistent with the builds behaving alike, not proof, and per-token agreement was not run.
Report: A smaller model should not become a different model →
Model compressionAgent workflowsAgentBenchQuality

Compressed model, in an agent loop four builds of one base, one machine

Qwen3.8-27B, the base (BF16)64 + 64 of 142 · 45.1%
Tileward-Qwen3.8-27b61 + 63 of 142 · 43.7%
Publisher’s official FP860 + 65 of 142 · 44.0%
Bonsai 2 27B, ternary57 + 64 of 142 · 42.6%
10 / 13
Tasks won / lost against the base
142
Tasks · 2 runs per build

A tie on finished tasks: all three compressed builds finish within noise of the base. The ternary build is the smallest file and works hardest for it. Paired by task over two runs, Tileward-Qwen3.8-27b wins 10 and loses 13 against the BF16 base, a gap smaller than the test can see. The publisher’s FP8 build finishes 1.1 points fewer (11 won, 16 lost) and the ternary build 2.5 fewer (13 won, 19 lost); neither gap is one this test can call. Every compressed build spent more prompt tokens than the base to get there: ours 7% more, the ternary build 28% more.

  • AgentBench OS: Linux tasks in a real container, pass or fail decided by the environment; the 142 tasks every arm got a turn on. Greedy decoding, thinking off, 32k context, one NVIDIA 80 GB GPU, the same tasks and settings for all four.
  • An agent loop does not repeat itself exactly, so 142 tasks resolve a gap of about ten points; every gap on this card is smaller than that, a direction rather than a size.
  • The three Qwen-format builds ran on vLLM; the ternary build ran on its publisher’s own server, the only one that loads it, with reasoning off like the others, so its publisher’s thinking-on figures do not compare to this row. On this Hopper-generation card the FP8 build ran its native 8-bit path.
All write-ups →
GovernanceQuality

Policy bypass battery meaning-preserving attacks

Written-to-evade requests2,251 tested
Stopped before governed inference2,201 observed
Crossed the policy gate50 observed
2,251
Adversarial cases
50
Residual bypasses
45 of 50
Benign controls allowed
Report: Put policy in front of the model—not inside the prompt →

Agent flows measured on a 120B open-weight model, tokens counted by the serving model rather than estimated. A provider that discounts a repeated prefix narrows what the saving is worth.

Bring your own workload. Point the harness at something you actually run. Talk to an engineer.