Skip to content
Docs

Performance

Every number on this page was measured on the hardware named beside it. None is extrapolated, and none is a vendor figure.

Read the measurement basis before quoting any of them — a throughput figure without its model, quantization and concurrency is not a fact about hardware.

Your hardware Run this Why
Apple Silicon Mac pip install 'lifeboat[hub,mlx]' or the desktop app Metal for quantized models, MLX for full-precision ones
NVIDIA data-center GPU the container image The only path with the tensor engine and the GPU optimization layer
NVIDIA Jetson (JetPack) pip install 'lifeboat[hub]' The CUDA container targets data-center drivers and is ~17 GB
AMD Instinct MI300X or newer the container image, AMD tag Native 8-bit float, so the tensor engine is the right choice
AMD Instinct MI210 / MI250 the container image, AMD tag, quantized models No 8-bit float silicon — see the CDNA2 trap
CPU-only server the Lite image or pip install 'lifeboat[hub]' No accelerator runtime to install
Mini-PC, NUC, integrated GPU the Lite image with LIFEBOAT_GPU_DEVICE One binary offloads to Intel, AMD or NVIDIA
Raspberry Pi class, 4–8 GB pip install 'lifeboat[hub]' See tiny edge devices

The installer picks for you — it detects the accelerator, detects an edge device, and recommends the path that fits, including declining to pull a container image the board cannot use:

Terminal window
curl -fsSL https://raw.githubusercontent.com/IterateAI/lifeboat-releases/main/install/get-lifeboat.sh | bash

Add --check to see what it decides without installing anything.

Single-stream baseline, same model, same harness

Section titled “Single-stream baseline, same model, same harness”

The comparable set: one 4-bit GGUF of a 0.5B model, one request at a time, so what varies is the machine and nothing else.

Machine Path TTFT median Decode Repeated prefix
Mac Studio M4 Max Metal 0.008 s 380.5 tok/s 31.4×
EPYC 9555P (Rocky 9), CPU only CPU 0.015 s 164.1 tok/s 50.8×
Jetson Orin Nano Super 8 GB CUDA (shipped default) 0.038 s 88.2 tok/s 49.9×
Jetson Orin Nano Super 8 GB Vulkan 0.057 s 47.2 tok/s —
Jetson Orin Nano Super 8 GB CPU 0.087 s 41.6 tok/s —

Two things worth taking from that table.

A 64-core server CPU beats an 8 GB edge GPU, at this model size. Decode is memory-bandwidth bound, and an EPYC’s DRAM bandwidth is real. Do not assume a small accelerator is faster than the CPU beside it — measure.

Repeated prefixes are 30–50× cheaper than cold ones. That is the single largest effect on this page, and it is free. An agentic client resending a conversation each turn pays it constantly; see prefix reuse.

Single-stream numbers do not predict a loaded server. These are aggregate throughput across all in-flight requests.

Mixture-of-experts is what makes a modest GPU fast

Section titled “Mixture-of-experts is what makes a modest GPU fast”

AMD Instinct MI210 (64 GB), one card, a 30B model with 128 experts of which 8 run per token — so roughly 3.3B parameters move per token instead of 30B:

Concurrent requests Quantized (GGUF engine) 8-bit float (tensor engine)
1 133.8 tok/s 8.4
4 225.8 31.7
16 475.8 123.8
32 480.4 243.0

The same card serving a dense 27B reaches 37.7 tok/s single-stream and peaks at 139.4. Same hardware, same engine, same quantization — 3.5× the throughput purely from the model’s architecture.

If you can choose the model, choose a mixture-of-experts one. It is the cheapest large win available, and it costs nothing at serving time.

An NVIDIA data-center GPU with the tensor engine

Section titled “An NVIDIA data-center GPU with the tensor engine”

RTX PRO 6000 Blackwell, dense 27B, 8-bit float checkpoint:

Concurrent requests Aggregate TTFT median
1 32.6 tok/s —
4 109.9 —
16 249.3 —
32 307.2 0.80 s

Note the shape rather than the numbers: the tensor engine keeps scaling and holds time-to-first-token nearly flat as load rises, because it batches continuously. The GGUF engine has fixed slots — it peaks at its slot count and then degrades, with TTFT growing linearly. On the same MI210, the GGUF engine went from 0.039 s TTFT at one request to 4.46 s at 32 while the tensor engine stayed between 0.29 s and 0.34 s.

So the two engines trade off, and the trade is concurrency. One or two users on a small model: the GGUF engine is usually faster and much cheaper to start. Dozens of concurrent users: the tensor engine’s flat latency is worth more than peak tokens per second.

Same MI210, a 35B mixture-of-experts model with a built-in draft head, 77% of drafted tokens accepted:

Concurrent requests Off Draft window 3 Draft window 6
1 83.9 tok/s 103.8 120.2
4 133.9 145.8 230.0
16 268.9 407.5 359.7

The window is a trade-off, not a dial to maximize. A wide window wins at low concurrency and loses at high concurrency, because rejected draft tokens are wasted compute that a busy server cannot spare. Leave it on Auto unless you know which end of that table your traffic sits at.

Already the fastest per-watt machine on this page. Three things still matter.

  1. Install the MLX extra — pip install 'lifeboat[hub,mlx]', or use the desktop app, which bundles it. Without it, full-precision safetensors models cannot be served on a Mac at all; Lifeboat will tell you so before you download 60 GB rather than after.
  2. Prefer 4-bit quantized models unless you specifically need full precision. Unified memory is shared with everything else on the machine.
  3. Leave the thread count alone. Lifeboat sets it to the performance core count, not the total. Including the efficiency cores makes generation slower, because the pool advances at the pace of its slowest member.

Check what the machine actually offers with lifeboat doctor.

This is the only hardware where the full optimization layer applies — cache compression, adaptive cache management, fair scheduling, memory tiering, admission control, dynamic precision — because all of it hooks the tensor engine’s scheduler.

  • Driver 580.65.06 or newer. This is exact, not approximate; see hardware requirements.
  • Serve an 8-bit or 4-bit float checkpoint rather than full precision. Halving the bytes read per token roughly doubles the decode ceiling, and Blackwell GPUs have native 4-bit float tensor cores.
  • Turn speculative decoding to Auto. It picks the fastest method the model’s own weights allow, with no setup.
  • Leave memory sizing on Auto. It derives the fraction from live GPU state at every start and learns each model’s real footprint. A hand-pinned value is measured once and then wrong forever.
  • Pick the routing mode for the traffic, not for the benchmark — see routing modes.

Install with pip, not the container. The CUDA image is ~17 GB and targets data-center drivers; JetPack ships its driver inside the board’s own OS, so that image is the wrong shape for the hardware — the installer detects a Jetson and says so rather than pulling it.

Terminal window
sudo apt install -y python3-venv
python3 -m venv ~/lifeboat && ~/lifeboat/bin/pip install -U pip
~/lifeboat/bin/pip install 'lifeboat[hub]'
~/lifeboat/bin/lifeboat engine install
~/lifeboat/bin/lifeboat up

A Jetson gets the CUDA engine automatically — lifeboat engine install detects the board and fetches a CUDA build rather than the general Vulkan one. That is worth roughly 2x: 47.2 → 88.2 tok/s single-stream and 112 → 180 tok/s peak on an Orin Nano Super, measured on the same model and the same engine source with only the backend changed. Nothing to configure, and lifeboat doctor shows which one you got (CUDA0: Orin rather than Vulkan0:).

The CUDA asset is larger — 81 MB against 27 MB — which is the whole cost.

Stay at or below a 4B model at 4-bit on an 8 GB board. The board shares its memory with the GPU, so weights, cache and the operating system all come from the same pool.

Instinct cards need the ROCm engine, and lifeboat engine install fetches it automatically — one build covers MI210, MI250, MI300X, MI325X and MI350X.

This is not an optimization but the difference between using the card and not: the portable engine reaches GPUs through Vulkan, whose open driver targets graphics parts, and a CDNA compute card has no graphics engine. It is therefore never enumerated. Measured on an MI210, the portable engine reported no GPU at all and served on the CPU; the ROCm engine reports ROCm0: AMD Instinct MI210 and is 1.89x faster on a single request (375.4 against 198.2 tok/s).

Worth knowing before you assume the GPU always wins: on that host, aggregate throughput at high concurrency was higher on the 64-core CPU (1330 against 883 tok/s) for a 0.5B model. A small model on a large server CPU parallelises extremely well, and the GPU’s advantage is latency there. The GPU’s margin grows with model size.

Radeon cards are unaffected and keep the portable engine — the open Vulkan driver does support RDNA.

AMD Instinct CDNA3 and newer (MI300X, MI325X, MI350X)

Section titled “AMD Instinct CDNA3 and newer (MI300X, MI325X, MI350X)”

Native 8-bit float, so the tensor engine and the full optimization layer are the expected configuration — the same guidance as an NVIDIA data-center GPU. Pull the image tag that matches the generation: the kernels are compiled for one generation per tag, and the wrong tag fails inside a kernel launch rather than at startup.

Serve quantized models here. Do not serve 8-bit float.

CDNA2 has no 8-bit float tensor hardware, so that path is emulated in software. Measured on the same card and the same dense 27B model:

Path Single-stream decode
4-bit quantized, GGUF engine 37.7 tok/s
8-bit float, tensor engine 1.7 tok/s

That is a 22× penalty, and no amount of tuning recovers it — kernel tuning was measured at about 8%. The card is not slow; the format is wrong for it. The same card runs a 30B mixture-of-experts at 480 tok/s aggregate.

Two settings matter on the GGUF engine here:

  • Keep per-server concurrency at 16 or below. Above the engine’s slot count, surplus requests queue inside the engine where Lifeboat cannot see them, so the console reports no wait while requests are waiting.
  • Confidential computing is not available on AMD GPUs and turns itself off rather than refusing every launch.

Genuinely viable, and faster than people expect — 164.1 tok/s single-stream on a 64-core EPYC at 0.5B, 55.5 tok/s on a smaller x86 host.

  • Use the Lite image or the pip package. Neither installs an accelerator runtime; the Lite image is about 720 MB against 17–28 GB.
  • Quantize. 4-bit reads a quarter of the bytes per token that full precision does, and decode is bandwidth-bound.
  • Expect time-to-first-token to dominate, not generation. Prefill is the compute-bound half. A smaller model improves it disproportionately.
  • Memory bandwidth is the specification that matters, not core count.

Integrated GPUs (Intel Iris/Arc iGPU, AMD APU)

Section titled “Integrated GPUs (Intel Iris/Arc iGPU, AMD APU)”

Detected automatically, but the container needs the render device passed through. Without it Lifeboat correctly reports the GPU as present and unreachable, and everything runs on the CPU at full speed with no error anywhere:

LIFEBOAT_GPU_DEVICE=/dev/dri:/dev/dri

One binary offloads to Intel, AMD or older NVIDIA hardware through Vulkan, and falls back to the CPU when there is no usable device.

Lifeboat runs on boards far below the server requirements. What changes is model sizing, not the software.

Available memory Largest comfortable model Realistic expectation
4 GB 0.5B–1.5B at 4-bit Classification, extraction, short replies
8 GB up to 4B at 4-bit Assistant-quality replies; 8B is the ceiling
16 GB 8B at 4-bit comfortably General-purpose serving

A 14B model does not fit in 8 GB — its weights alone are about 8.1 GB before the operating system takes its share.

Install with pip and the GGUF engine. The installer recognises a Jetson, a Raspberry Pi class board, or any host with 8 GB or less and few cores, and recommends that path directly instead of a container image such a board cannot run.

Read out of the published artifacts rather than inferred from the build flags, because on this class of hardware the difference is a crash:

  • The Linux engines link GLIBC_2.27 / GLIBCXX_3.4.25. That is Ubuntu 18.04, Debian 10 and RHEL 8 onward — so Raspberry Pi OS, JetPack 5 and 6, Ubuntu 22.04 and Debian 12 all run the prebuilt engine. Upstream’s own Linux builds need glibc 2.38 and refuse to start on every one of those.
  • Every CPU variant is compiled in and selected at runtime. The aarch64 build ships an armv8.0 backend — Raspberry Pi 4 class — through to armv9.2; the x86-64 build ships a plain x64 baseline that needs no AVX, through to zen4. Nothing is -march=native, so one artifact spans an Atom to a current server part.
  • Vulkan is compiled in but not required. With no usable device the CPU backends take over, so the same binary serves a headless board and one with an integrated GPU.

The one ARM instruction above the armv8.0 baseline in any always-loaded object is a GCC outline atomic behind a runtime check, with a load-exclusive fallback — so it executes correctly on a Pi 4.

Measured on a Jetson Orin Nano: the whole control plane — console, load balancer, model registry, licensing, the OpenAI and Anthropic API surface — holds 63 MB of RSS. On a small board essentially all of the memory budget goes to the model, which is what makes the sizing table above the operative constraint rather than a rule of thumb.

Three settings that matter more here than anywhere else:

  • Keep the context window small. Cache memory scales with context × concurrent requests, and on a small board that is the binding constraint rather than the weights. Lifeboat sizes it from the hardware; lowering it further buys headroom.
  • Serve one model, not several. Each loaded model holds its weights resident.
  • Ask the machine rather than guessing. GET /api/hardware/profile reports what this specific host can run, including a container memory limit if one is set — which the raw system figures do not reflect.

These apply on every machine, in rough order of how much they are worth.

Measured at 30–50× on time-to-first-token for a repeated prefix. Keep the stable part of a prompt — system instructions, tool definitions, retrieved context — at the front and constant, and put what varies at the end. Reordering a prompt costs nothing and is frequently the largest available win.

Two caveats: some model architectures disable prefix caching (the console’s server row says so), and sticky-session routing is what keeps a conversation landing on the backend that already holds its cache.

Decode reads every active weight once per token, so bytes per token is the ceiling. 4-bit against full precision is roughly a 4× difference in that ceiling and usually a small quality difference. See model formats.

The default favours time-to-first-token, which is right for interactive use and leaves throughput on the table under load. hybrid switches automatically as a pool heats up. See routing modes.

A larger context window costs cache memory on every request, whether or not the request uses it. Extending beyond a model’s native window also degrades short-prompt quality, because the scaling applies to every request. The default is the model’s native window on purpose.

Lifeboat reserves memory for stopped servers that might start, capped at 40% of the card. Three stale rows on one deployment held back 24 of 96 GB and dropped the server actually serving traffic from 0.80 to 0.64 of the GPU.

How this compares to other inference servers

Section titled “How this compares to other inference servers”

Everything above measures Lifeboat against itself on different hardware. The comparison against other servers — llama.cpp and ollama, same machine, same model bytes, one client — is published with its harness and raw results at github.com/IterateAI/lifeboat-releases/tree/main/benchmark/comparison, so it can be re-run rather than taken on trust.

Out of the box, with nothing tuned on any side:

Hardware Lifeboat peak Next best
64-core server CPU 1330 tok/s llama.cpp 686 1.94x
Jetson Orin Nano 180 tok/s ollama 89 2.01x
Apple M4 Max 166 tok/s ollama 112 1.48x

Two things that page states plainly and are worth repeating here. On a Mac, Lifeboat and llama.cpp are level — both saturate memory bandwidth, and no serving layer beats physics. And the aggregate wins come from one decision: Lifeboat sizes its concurrent slots from the host, where a bare engine takes a fixed default because it cannot inspect the machine it was launched on. That is worth 2x on a large box and nothing at all on a small one, which is exactly the shape you should expect.

Do not plan against this page. Run the sweep on your hardware, with your model:

  • Console → Benchmarks runs a concurrency sweep against a running server and reports tokens per second, time-to-first-token, inter-token latency and cost per million tokens. See benchmarks.
  • lifeboat doctor reports what a pip or desktop install can actually use — which engines are present, whether the GPU is reachable, and what is missing.
  • GET /api/hardware/profile reports the sizing this host supports.

Three rules that make a measurement trustworthy, each learned by getting it wrong:

  1. Generate at least 256 tokens. Below about 128, fixed overhead dominates and real differences disappear — a 43% speculative-decoding gain measured as 0% at 128 tokens and was obvious at 256.
  2. Discard the first request. A cold server pays graph capture and allocator growth that no steady-state user ever sees.
  3. Count reasoning output as output. A thinking model emits its reasoning in a separate field; a harness counting only the final answer scores a perfectly healthy server as zero tokens with no first-token time at all.

The single-stream table: Qwen2.5-0.5B-Instruct, 4-bit GGUF, one request at a time, 64 generated tokens, temperature 0. Time-to-first-token is the median of five streamed requests measured to the first content delta. Decode rate is (tokens − 1) / (end-to-end − time-to-first-token), median of the same five. The repeated-prefix figure is the cold first-token time over the warm one on an identical ~1500-token prefix. Produced by the acceptance harness that ships with the pip package.

The concurrency tables: real models on real deployments, thinking disabled so two models are comparable, warmup excluded, aggregate throughput across all in-flight requests. Model, quantization and card are named with each table because none of those numbers transfers to a different combination.

What is not measured here: the GPU optimization layer’s effect. It earns its keep under memory pressure, so a small model on a large card — the only configuration in which an on/off comparison is cheap to run — shows nothing. Measured on an MI210 at 0.5B it was neutral within noise at every concurrency level, which is the expected result and not a claim about its value. To see what it is worth, measure a model that fills the card.