Governed · Compressed: stop paying to resend the whole conversation, agent history or coding session.
Every turn resends the whole thread. Tileward Context sends only the part that answers.
Works with the clients you already run
28 have a recipe to paste, and anything else that speaks MCP works too.
Our own usage, read off the same dashboard you'd get.
of one long-running Claude Desktop thread never resent, 21.6M tokens of 23.8M.
median cut per conversation on Claude Code sessions over 50k tokens, 45–97% across those 61. Counting all 179 sessions, short ones included, the median is 63%
across every conversation, lifetime, 23.3M saved of 26.1M
If you already run prompt caching, read 90.7% against a cached baseline.
A cache hit is a discount on tokens you still send: the cached prefix travels and occupies the window. Tileward's are not sent at all, so the window comes back, and against a well-cached baseline the dollar saving shrinks by roughly an order of magnitude while the tokens stay withheld.
Where it pays, conversation by conversation
| Conversation | Client | Tokens without Context cumulative resend | Tokens with Context actually sent | Never resent share saved |
|---|---|---|---|---|
| long-running assistant thread | Claude Desktop | 23.8M | 2.2M | 90.7% |
| agent session · repo work | Claude Code | 516.8k | 114.4k | 77.9% |
| agent session · refactor | Claude Code | 470.6k | 146.5k | 68.9% |
| agent session · skill authoring | Claude Code | 288.0k | 56.0k | 80.6% |
| agent session · debugging | Claude Code | 194.5k | 51.1k | 73.7% |
| agent session · long review | Claude Code | 146.2k | 6.4k | 95.7% |
| agent session · docs pass | Claude Code | 123.2k | 16.2k | 86.9% |
Short threads have nothing to distill.
Less is sent, and what matters still arrives.
Each time you ask something, Tileward picks the turns that bear on that question and sends only those, so the model answers from what was actually said, however long ago.
Nothing is thrown away
Every turn stays stored as it was said; only what is sent each time is small.
The relevant turns go
A question about a decision from last week gets the turns where that decision was made, in the words that were used.
You can see what went
The console shows which turns were sent for each question and what they cost, so you can check that the model saw the right things.
In our own test, a model given only those turns answered every question it could answer from the whole conversation, from a fraction of the tokens. The test. And when the answer isn’t in the conversation, it makes one up far less often: 26% of the time, against 63% with the whole transcript. How we measured it.
Less to read, so the model starts responding sooner.
Before a model writes a word it reads everything it was sent, and that reading takes longer the longer the history gets. Sending only the part that answers cuts the wait at the source: a 255k-token history sent as 51k takes about 13× less GPU time to read, 49 s against 3.9 s.
| History sent to Tileward 35B-A3B | Wait until it starts responding |
|---|---|
| 8k tokens | 0.42 s |
| 32k tokens | 2.1 s |
| 64k tokens | 5.4 s |
| 130k tokens | 15.6 s |
| 203k tokens | 33 s |
| 255k tokens | 49 s |
On a model that answers without reasoning first, the whole reply lands about 20% sooner, even against a cached prompt.
Three steps, about two minutes.
Context is an MCP server, so it plugs into the client you already use rather than replacing it.
Create a free account
No card. You get $25 of credit, 2,000 recalls and 100 MB of documents (enough to run a real thread through it).
Start freePick your client in the docs
Its recipe is a few lines to paste, signed in through the browser or with a key.
Open the connect guideAsk it to remember something
Then recall it a turn later. The console watches for that first call to land, and from then on your own thread's savings read off the same dashboard as the figures above.
Anything else that speaks MCP or the OpenAI wire format works too. The docs carry the endpoint table, which authentication each client can do, and the conversation id you send with every call. Tileward Governance runs against the same key, and the docs mark which clients have the hook it needs.
Start free. Compare plans in one place.
Context is included in every complete Tileward plan and is also available on its own. Paid Context plans start at $19/month; complete plans start at $39/month.
Bring us a long conversation.
We'll distill it, compare against sending the whole thing, and show you what it retrieved. Thirty minutes, an engineer, no deck.