Run large models on hardware you own.
Tileward is an inference platform for teams running large models on hardware they own. Four products, one system: model compression, conversation memory, governance, and answers grounded in your own files.
Same MMLU as the FP8 build it replaces — 82.58% against 82.60%.
About 270 tokens handed over instead of the transcript — a page.
64 regulated topics of 215 lockable tiles, every decision recorded.
Passages cited, and never pooled with other tenants.
Four problems most teams solve with four vendors, run as one system.
Running the four together is what makes the numbers below possible: a small context window stops being a limitation once you stop filling it with the transcript, and a guard that sits outside the model costs a fraction of one that asks the model to police itself.
What the compression costs
Qwen3.6-35B-A3B, at roughly 4 bits per weight, built for a 256K context window with 65,536 served on the hosted API. It scores the same on MMLU as the publisher's own FP8 build and needs no change above the model, so anyone already serving FP8 can swap it in and keep the memory: longer context, more requests at once, or a smaller machine. The full MMLU run, including the two builds that score above us.
Tileward 397B is next — a reasoning mixture-of-experts model
Learn more about Tileward Models →What never gets resent
A long thread is resent from the top on every turn. Tileward Context keeps the history and hands over only the part that answers the question — about 270 tokens, a page.
Learn more about Tileward Context →What a refusal costs
Lock a topic and the request is declined by a guard outside the model, however it is reworded — and every decision is written down for audit.
Learn more about Tileward Governance →Answers from your own files
Upload what the model should know — handbooks, contracts, runbooks — and answers come back grounded in them, with the passages cited. Your files stay private to your account: never pooled with other tenants, never used for training.
Open Documents in the app — Tileward Documents →Same models on every tier. Pick your ceiling.
Everything ships on every plan — paid tiers raise volume, seats, deployment and request rate, never model quality. Switch the tabs to price a single product instead. Explore starts at $0 with no card.
- 100K model tokens
- 100 requests/min
- 2K recalls
- 200 guard checks
- 10 tiles · 7-day audit
- 100 MB storage
- 1M model tokens
- 300 requests/min
- 20K recalls
- 2K guard checks
- Unlimited tiles · 10 custom
- 30-day audit window
- 10 GB storage
- 20M model tokens
- 1,000 requests/min
- 400K recalls
- Unlimited guard checks
- Unlimited tiles · 100 custom
- 30-day audit window
- 200 GB storage
- 100M model tokens
- No rate limit
- 2M recalls
- Unlimited guard checks
- Unlimited tiles · 100 custom
- 1-year audit · SSO
- 1 TB storage
- Committed token pool
- On-prem or air-gapped
- Unlimited recalls
- Unlimited guard checks
- Private tiles · SIEM export
- Storage is your own disk
- Unlimited conversations
- 2K recalls / mo
- 100 MB storage
- Kept as long as the account is active
- Unlimited conversations
- 20K recalls / mo
- 5 GB storage
- Kept as long as the account is active
- Unlimited conversations
- 400K recalls / mo
- 100 GB storage
- Kept as long as the account is active
- Unlimited conversations
- Unlimited recalls
- Storage is your own disk
- On-prem · air-gapped
Priced per account, not per seat. Storage counts distilled conversations and uploaded documents, not the raw transcript — which is why it goes further than it looks. Adding Governance at the same rung costs $19 or $99; Everything is a dollar more than the pair.
- 200 guard checks
- 10 tiles
- 100 requests/min
- 7-day audit window
- 2K guard checks
- Unlimited tiles · 10 custom
- 1,000 requests/min
- 30-day audit window
- 40K guard checks
- Unlimited tiles · 100 custom
- 20K requests/min
- 30-day audit window
- Committed volume
- Private tiles of your own
- SIEM export
- On-prem audit retention
Every rung gets the full roster of 215 tiles and the same classifier — what changes is volume, custom tiles of your own, rate limit and how long the record is kept. Adding Context at the same rung costs $19 or $49; Everything is a dollar more than the pair.
Bring us a long conversation, or a topic you can’t let the model discuss.
We’ll distill the first and lock the second, try to get past both, and show you the record. Thirty minutes, an engineer, no deck — or skip the call, the free tier does all three today.
Pick a time right here on Google Calendar — no form, no waiting on a reply. Prefer email? hello@tileward.com