New Tileward 35B-A3B: 2.74× smaller, and the same MMLU as the FP8 build it replaces

Run large models on hardware you own.

Tileward is a governed inference platform that lets teams run large models on hardware they own: we compress models to a fraction of their served memory, send only the relevant slice of a conversation instead of the whole transcript, and refuse restricted topics before the model writes a token.

No credit card · no time limit · your keys, your hardware, your data

Three problems most teams solve with four vendors, run as one system.

Public beta, everything here is live and metered. Tell us what breaks at hello@tileward.com.
The Tileward platform

A model host, a retrieval layer, a document store, and a policy tool bolted on afterwards — that's the usual stack. Running them together is what makes the numbers below possible: a small context window stops being a limitation once you stop filling it with the transcript, and a guard that sits outside the model costs a fraction of one that asks the model to police itself.

01 · Models

What the compression costs

2.74×
smaller served, Qwen3.6-35B-A3B. Weights on disk, measured rather than estimated. The expert layers go to 4 bits on a group of 32; attention, the router and the embeddings stay at full precision.
0.02 pts
from Full MMLU, 5-shot, 14,042 questions. Ours 82.58%, the publisher's official FP8 82.60%, both ±0.31, so the two cannot be separated, at 11.2 GB less. Two community 4-bit builds do score above us, GPTQ 83.30% and AWQ 83.17%, which is a change in behaviour rather than a free upgrade., so it swaps in unchanged

Qwen3.6-35B-A3B, at roughly 4 bits per weight, with a 256K context window. It scores the same on MMLU as the publisher's own FP8 build and needs no change above the model, so anyone already serving FP8 can swap it in and keep the memory: longer context, more requests at once, or a smaller machine. The full MMLU run, including the two builds that score above us.

Tileward 397B is next — a reasoning mixture-of-experts model

Learn more about Tileward Models
02 · Tileward Context

What never gets resent

90.7%
of one long thread never resent

A long thread is resent from the top on every turn. Tileward Context keeps the history and hands over only the part that answers the question — about 270 tokens, a page.

Learn more about Tileward Context
03 · Governance

What a refusal costs

~14
tokens to decide, no answer written

Lock a topic and the request is declined by a guard outside the model, however it is reworded — and every decision is written down for audit.

Learn more about Tileward Governance
04 · Documents

Answers from your own files

1 TB
included on Team, 50 GB from $39

Upload what the model should know — handbooks, contracts, runbooks — and answers come back grounded in them, with the passages cited. Your files stay private to your account: never pooled with other tenants, never used for training.

Learn more about Tileward Documents

Same models on every tier. Pick your ceiling.

Tileward Context and Governance ship in full on every plan — paid tiers only raise volume, seats, deployment, and request rate. Explore starts at $0 with no card; paid tiers are clean multiples on top.

No card required
Explore
$0/mo
$1 credit, no card
  • 100K model tokens
  • 2K Context recalls
  • 200 guard checks
  • 100 MB documents
  • 3 tiles · 2 keys
  • 7-day audit window
  • 100 requests/min
Start free
Build
$39/mo
$390/yr
or $390/yr · 2 months free
$32.50/mo · saves $78
a solo developer · 10× Explore
  • 1M model tokens
  • 20K Context recalls
  • 2K guard checks
  • 50 GB documents
  • 25 tiles · 10 keys
  • 30-day audit window
  • 300 requests/min
Start Build
Team
$349/mo
$3,490/yr
or $3,490/yr · 2 months free
$290.83/mo · saves $698
5 seats · $35 per extra seat
  • 100M model tokens
  • 2M Context recalls
  • 200K guard checks
  • 1 TB documents
  • Unlimited tiles · 100 keys
  • 1-year audit · SSO
  • No rate limit
Start Team
Enterprise
Custom
committed use or air-gapped licence
  • Committed-use token pool
  • Storage is your own disk
  • Or a flat air-gapped licence
  • SSO + SCIM
  • SIEM export
  • On-prem · no rate limit
Talk to an engineer

Bring us a long conversation, or a topic you can’t let the model discuss.

We’ll distill the first and lock the second, try to get past both, and show you the record. Thirty minutes, an engineer, no deck — or skip the call, the free tier does all three today.

Pick a time right here on Google Calendar — no form, no waiting on a reply. Prefer email? hello@tileward.com