Benchmarks
The Benchmarks page runs a concurrency sweep against a running server and records what it measured. It answers the question a self-hosted evaluation has to answer: can this hardware serve enough tokens, fast enough, for enough concurrent users, to beat what you are paying now?

Running one
Section titled “Running one”Pick a running server, give it a list of concurrency levels, and start. Levels are measured one at a time, in ascending order — running them together would make every level contend with every other, so no level would measure what it claims to.
A benchmark against a stopped server is refused rather than run. A sweep with nothing serving reports zeros, and zeros read as bad hardware rather than as a stopped server.
What each column means
Section titled “What each column means”| Column | Meaning |
|---|---|
| Throughput tok/s | Tokens per second across all requests at that level. The fleet number. |
| Per-request tok/s | The mean of each request’s own rate — what one user feels. It falls as concurrency rises even while throughput climbs. |
| TTFT p50 / p95 | Time to the first token. Carries the prefill cost, so it grows with prompt length and with queueing. |
| TPOT p50 | Time per output token — the decode cadence, measured between tokens. Excludes the first token deliberately: folding prefill into it makes a fast decoder look slow on a long prompt. |
| e2e p50 | Whole-request wall time. |
| Failed | Requests that errored. A non-zero count invalidates the level’s averages — read it first. |
| GPU mem | Memory in use during that level, sampled from the accelerator. Feeds the cost model. |
Two settings that change whether the number is true
Section titled “Two settings that change whether the number is true”Output tokens per request. Below about 128, per-request fixed overhead dominates and a real difference can vanish into it. Measured on one deployment, a 43% single-stream gain was invisible at 128 tokens and obvious at 256. The default is 256, and a lower value is recorded as a warning on the run rather than silently accepted.
Thinking. Models that reason before answering are run with thinking off by default, so two models are comparable and a short run is not spent reasoning without reaching an answer. Reasoning tokens are counted as tokens either way — a harness that counted only answer tokens would score a perfectly good run as zero.
Each level is preceded by a warm-up request that is excluded from every reported number. The first request into a cold slot set pays costs no steady-state user sees.
Cost per million tokens
Section titled “Cost per million tokens”Supply an hourly hardware cost and the Platform card reports $/1M output tokens; add a baseline rate and it reports the saving against it.
The basis is stated on the card rather than implied: output tokens only, at the measured peak sustained for the whole hour. It excludes prompt tokens, idle time and redundancy — all of which make real-world cost higher. Compare it against a like-for-like quoted rate, and quote it with its assumptions attached.
Leave the hourly cost blank and no cost figure is produced. A number derived from a guessed rate is worse than no number, because it gets quoted.
The three cards
Section titled “The three cards”Results are reported as three cards because they answer different questions and travel separately:
- Model card — the weights: single-stream rate, TTFT, TPOT.
- Engine card — the serving stack: peak throughput, the concurrency it peaked at, and the full scaling curve.
- Platform card — the hardware: GPU, memory, hourly cost, cost per million tokens and the basis for it.
“Copy cards as JSON” puts all three on the clipboard for pasting into a report.
What this does not do
Section titled “What this does not do”It measures performance, not quality. Whether the served model gives good answers — accuracy, hallucination rate, task success — is a separate exercise needing a graded dataset, and belongs in an evaluation tool rather than here. A model can top every column on this page and still be the wrong model.