Tileward / Use cases / Compare
Live comparison

Ask one question. Watch two builds answer it.

Pick a use case and ask. Tileward Qwen3.8-27B in TWX48 and the original Qwen3.8-27B in BF16 answer side by side. Same question, same documents, same moment. The numbers are measured live.

Live comparison

Answers come from a bank’s internal policy and procedure documents.

Tileward Qwen3.8-27B

TWX48
Time to first token
—
Reasoning time
—
Time to last token
—
Tokens per second (decode)
—
Tokens (prompt / completion / total)
—
The answer appears here.

Qwen3.8-27B

BF16
Time to first token
—
Reasoning time
—
Time to last token
—
Tokens per second (decode)
—
Tokens (prompt / completion / total)
—
The answer appears here.
What the numbers mean

Both sides run on one clock.

Each question is checked once and grounded once. Both models then receive the same messages and settings at the same moment, and timing starts at the gateway, so your own network sits outside every figure.

Time to first token
From the request leaving the gateway to the first token back, reasoning or answer.
Reasoning time
From the first reasoning token to the first answer token, when Reasoning is on.
Time to last token
From the request leaving the gateway to the last token back.
Tokens
The engine’s own count for the prompt and the completion. The completion includes any reasoning.
Tokens per second
Completion tokens after the first, divided by the time from first token to last: the speed of generation.
On your own documents

Run the same test on the questions your team asks.

Bring the documents and the questions. An engineer runs them with you and shows you the numbers. Thirty minutes, no deck.