Keep most of the answer rate on 19× less context.
After excluding LOCOMO’s unanswerable questions, Tileward Context answered 29% with 1,499 context tokens per question. The full transcript answered 33–35% with 28,365. Keeping only the newest turns answered 3%, and Tileward beat that baseline question by question, 418 wins to 16 losses. This is one model scored by text matching, not a claim of parity with full history.
See the matching eval card for the result and limits in one screen.
29% answered with 1,499 tokens; full history reached 33–35% with 28,365.
LOCOMO (Maharana et al.) contains ten conversations that run for hundreds of turns across several sittings, with 1,986 questions about what was said. We asked the answerable questions four ways: with the whole transcript, the newest turns, what Tileward Context returned, and no conversation at all.
| What the model was given | Questions answered | Context per question | Against sending everything |
|---|---|---|---|
| Everything, as-is | 33% | 28,365 tokens | Baseline |
| Tileward Context | 29% | 1,499 tokens | 4 points fewer, on 19× less context |
| The newest turns that fit the same budget | 3% | 2,165 tokens | 30 points fewer, on 13× less |
| Nothing (no conversation at all) | 0% | 138 tokens | 33 points fewer, on 206× less |
The token saving is large; the answer-rate cost is real.
This run supports using selective context when sending the whole transcript is too expensive or too large for the window. It does not support calling the two approaches equivalent: Tileward used 18.9× less context and finished four to six points behind the full-transcript score. Test whether that trade is acceptable on your own conversations.
This is not a parity claim.
We are not claiming to beat sending everything. On this test we come out behind it, and the gap is not rounding. Our measurement also understates the full-transcript arm: on the longest questions it sometimes ran out of room to answer at all. Setting those aside it scores 35%, so read the honest range as 29% for us against 33–35% for sending everything.
These numbers are not comparable to other companies' published numbers on the same benchmark. We score answers by matching text; the published figures elsewhere are scored by asking a model to judge. Same benchmark, different ruler.
One fifth of this benchmark rewards saying nothing. Some questions are unanswerable and the correct response is to decline, so an approach that returns no context at all scores well on them. Every number on this page excludes those questions. Any figure quoted from this benchmark without that exclusion (ours or anyone else's) is partly measuring how little was sent.
One model, and it is not a frontier one. A stronger model may use a compressed history better, or worse. We have not tested that.