Tileward / Models
01 · Tileward Models

Governed · Compressed: run a bigger model in your VPC, on an air-gapped network, or on the GPUs you have.

A model that does not fit on the card you have costs a second card, a smaller model, or a cloud bill. Tileward Models fits it.

Works with the clients you already run

28 have a recipe to paste, and anything else that speaks MCP works too.

Quality

A compressed model has three jobs. Two are measured, one is where it runs.

Fit on the hardware you have, stay the model you chose, and run wherever you run it. The first two are measured against the publisher’s own release.

Fits the card you have

24.5 GB

of weights for Tileward 35B-A3B, from 68.6 GB as published: 2.8× smaller, and 1.46× smaller than the publisher’s own FP8 release.

The write-up → · the eval

Still the model you chose

±0.3

is the error bar on each score, and the gap between our build and the publisher’s FP8 release sits inside it: level on a multiple-choice test.

The write-up → · the eval

Runs where you run it

3

places it runs unchanged: our hosted API, your VPC, or air-gapped.

Where each piece runs →

Catalogue

Every model we have compressed.

Three are served today and one is no longer served. The rest are compressed and ready to serve.

Served todayCompressed by TilewardMixture of expertsPer-tile locking

Tileward 35B-A3B

Compressed from Qwen3.6-35B-A3B. Serving requests today.

Weights
68.6 GB → 24.5 GB, 64% smaller
Active
~3B of 35B parameters per token
Context
262,144 tokens
Results against the FP8 build →
Served todayCompressed by TilewardDense

Tileward-Qwen3.8-27b

Compressed from Qwen3.8-27B. Serving requests today.

API id
Tileward-Qwen3.8-27b
Input
Text and images
What we serve →
Served todayAs publishedMixture of experts

gpt-oss 20B

Served as its publisher shipped it.

Active
~3.6B of 20B parameters per token
Context
8,192 tokens
Locking
Not available on this model
What we serve →
Ready to serveCompressed by TilewardMixture of experts

GLM-5.3-Flash

Compressed, not served. Hosted or on your hardware, on request.

Ready to serveCompressed by TilewardDense

Llama-3.1-Nemotron-Ultra-253B

Compressed, not served. Hosted or on your hardware, on request.

Ready to serveCompressed by TilewardMixture of experts

Nemotron 3 Nano 30B

Compressed, not served. Hosted or on your hardware, on request.

Ready to serveCompressed by TilewardMixture of experts

Qwen3.5-397B

Compressed, not served. Hosted or on your hardware, on request.

Ready to serveCompressed by TilewardMixture of experts

Qwen3-Next-80B-A3B

Compressed, not served. Hosted or on your hardware, on request.

Ready to serveCompressed by TilewardMixture of experts

Qwen3-Coder-30B-A3B

Compressed, not served. Hosted or on your hardware, on request.

No longer servedAs publishedMixture of experts

gpt-oss 120B

No longer served.

Running one of the unserved ones? We can serve it for you, hosted or on your hardware.

Served builds

Every model we serve, spec by spec.

Weights, memory, context window and locking for each served build.

Tileward 35B-A3B: Qwen3.6-35B-A3B, compressed to just over a third

Serving requests today. A mixture-of-experts: 256 experts per layer, 8 of them consulted on any one token. It costs about what a 3B model costs to run, and knows what a 35B model knows..

MeasureResult
BaseQwen3.6-35B-A3B
Parameters35B total, ~3B active
Size on diskWeights on disk, measured rather than estimated, on the catalog’s MiB/1000 basis.
Smaller by64%, a 2.8× ratio
Context window262,144 tokens
Per-tile lockingAvailable

Measured against the publisher’s FP8 build: the eval · the write-up.

What you get

Your hardware, your VPC, the same governance.

1.8× to 2.8× smaller

Tileward 35B-A3B, a mixture-of-experts, lands 2.8× smaller than the weights it was compressed from. The dense Tileward-Qwen3.8-27b lands 1.8×, because more of a dense model stays at full precision.

OpenAI-compatible

Change the base URL and your existing client works.

Governed by the same tiles

The 215 Tileward Governance tiles apply to the self-hosted model exactly as they do to the hosted API, with audit records retained locally.

client = openai.OpenAI(base_url="https://api.tileward.com/v1", api_key=TILEWARD_KEY)
client.chat.completions.create(model="Tileward-Qwen3.6-35B-A3B", messages=[...])
# also answers to tileward-35b-a3b and Qwen3.6-35B-A3B-TW, at the same rate
# same call self-hosted, only the base_url changes

Three ways to run it

Hosted API
Minutes
  • Change the base URL
  • Default tile pack
  • Zero infrastructure
Your VPC
Days
  • Single node, one GPU
  • Tiles and policies of your own
  • Data never leaves
On-prem, air-gapped
Weeks
  • Fully offline
  • Runs on your hardware
  • Local audit retention
Run it on your own hardware

Tell us the hardware you have. We'll tell you what it will do.