Run large models on hardware you own.
Tileward is a governed inference platform that lets teams run large models on hardware they own: we compress models to a fraction of their served memory, send only the relevant slice of a conversation instead of the whole transcript, and refuse restricted topics before the model writes a token.
Three problems most teams solve with four vendors, run as one system.
A model host, a retrieval layer, a document store, and a policy tool bolted on afterwards — that's the usual stack. Running them together is what makes the numbers below possible: a small context window stops being a limitation once you stop filling it with the transcript, and a guard that sits outside the model costs a fraction of one that asks the model to police itself.
What the compression costs
Qwen3.6-35B-A3B, at roughly 4 bits per weight, with a 256K context window. It scores the same on MMLU as the publisher's own FP8 build and needs no change above the model, so anyone already serving FP8 can swap it in and keep the memory: longer context, more requests at once, or a smaller machine. The full MMLU run, including the two builds that score above us.
Tileward 397B is next — a reasoning mixture-of-experts model
Learn more about Tileward Models →What never gets resent
A long thread is resent from the top on every turn. Tileward Context keeps the history and hands over only the part that answers the question — about 270 tokens, a page.
Learn more about Tileward Context →What a refusal costs
Lock a topic and the request is declined by a guard outside the model, however it is reworded — and every decision is written down for audit.
Learn more about Tileward Governance →Answers from your own files
Upload what the model should know — handbooks, contracts, runbooks — and answers come back grounded in them, with the passages cited. Your files stay private to your account: never pooled with other tenants, never used for training.
Learn more about Tileward Documents →Same models on every tier. Pick your ceiling.
Tileward Context and Governance ship in full on every plan — paid tiers only raise volume, seats, deployment, and request rate. Explore starts at $0 with no card; paid tiers are clean multiples on top.
- 100K model tokens
- 2K Context recalls
- 200 guard checks
- 100 MB documents
- 3 tiles · 2 keys
- 7-day audit window
- 100 requests/min
- 1M model tokens
- 20K Context recalls
- 2K guard checks
- 50 GB documents
- 25 tiles · 10 keys
- 30-day audit window
- 300 requests/min
- 20M model tokens
- 400K Context recalls
- 40K guard checks
- 200 GB documents
- Unlimited tiles · 50 keys
- 3 seats · 30-day audit
- 1,000 requests/min
- 100M model tokens
- 2M Context recalls
- 200K guard checks
- 1 TB documents
- Unlimited tiles · 100 keys
- 1-year audit · SSO
- No rate limit
- Committed-use token pool
- Storage is your own disk
- Or a flat air-gapped licence
- SSO + SCIM
- SIEM export
- On-prem · no rate limit
Bring us a long conversation, or a topic you can’t let the model discuss.
We’ll distill the first and lock the second, try to get past both, and show you the record. Thirty minutes, an engineer, no deck — or skip the call, the free tier does all three today.
Pick a time right here on Google Calendar — no form, no waiting on a reply. Prefer email? hello@tileward.com