Skip to content
Docs

Helm Intelligent Routing

Most requests do not need your largest model. Helm Intelligent Routing serves them on a small pilot model in the same family and escalates the hard ones to the larger target — and when it escalates, it carries the pilot’s already-computed KV cache across instead of making the target read the prompt again.

It is off by default, beta, and a pair may only be used after it has been calibrated and has cleared a retention gate.

Why carrying the cache across is the hard part

Section titled “Why carrying the cache across is the hard part”

Two models in a family do not store the same prompt the same way. A key also carries the position it was produced at, folded in as a rotation. So the transfer is not a copy:

  1. the pilot’s keys are un-rotated back into content space,
  2. a per-layer, per-head linear map — fitted once, offline — converts the pilot’s representation into the target’s,
  3. the result is re-rotated into the position it will occupy in the target’s sequence.

Step 3 is not optional and there is no switch to skip it. Skipping it produces a cache of exactly the right shape, full of plausible-looking numbers, that decodes to noise.

Calibration fits that map for ONE pair of models, offline, on a GPU:

Terminal window
lifeboat helm calibrate \
--source-model Qwen/Qwen3-0.6B \
--target-model Qwen/Qwen3-1.7B \
--corpus fineweb-edu --k 4

It samples a probe corpus, prefills both models over it, fits the per-(layer, head) maps, and then scores the result on probes the fit never saw. Exit code 0 means the pair cleared the gate, 7 means it fitted and did not, 4 means there is no accelerator on this host.

You can also run it from the console: Server Settings → Performance & Routing → Helm Intelligent Routing → Calibrate, which streams the same progress.

Calibration is minutes for a small pair, and the artifact it writes is a few hundred MB. It is a one-off per pair: you do not re-run it per request, per server or per restart.

Number What it means
Retention How often the target predicts the SAME next token with the transferred cache as it would have predicted on its own. This is the gate — default 0.90.
KL How far the whole output distribution moved, not just the top choice. A sampling deployment feels this; a greedy one may not.
KV cosine How close the mapped cache is to the target’s own. Diagnostic: low cosine AND low retention means the map is wrong; high cosine and low retention means this target is simply sensitive to small cache error, and refitting will not fix the pair.

A pair that fails is recorded as TRANSFER_INELIGIBLE with its score, not discarded — so the next person can see it was tried and what it scored, instead of trying it again.

Not every family pair transfers well. Measured on an AMD MI210, Qwen3-0.6B → Qwen3-1.7B retains 0.77 against a 0.90 gate and is correctly refused — and ten times more calibration data moved that by 0.06, so it is a property of the pair rather than of the fit. The control in the other direction: calibrating Qwen3-1.7B against itself retains 0.98, which is what tells you the refusal above is about those two models and not about the machinery. Treat calibration as a measurement, not a formality.

The point of carrying the cache across is to not pay for the prompt twice. Measured on an MI210, the target’s own prefill against the transfer that replaces it:

Prompt tokens Target (1.7B) prefill Transfer Transfer is
512 36 ms 11 ms 3.4x cheaper
2,048 103 ms 26 ms 4.0x cheaper
8,192 414 ms 91 ms 4.6x cheaper
16,384 968 ms 179 ms 5.4x cheaper

That comparison assumes the pilot has already answered, which is the premise of routing: the request was served on the small model, judged hard, and escalated. If you count the pilot’s prefill from scratch as well, the whole path is roughly break-even at 512 tokens and about 1.15x cheaper from 2,048 tokens up — the longer the prompt, the more the transfer wins.

What happens when a transfer cannot be used

Section titled “What happens when a transfer cannot be used”

It falls back to a normal full prefill on the target, and the request is served exactly as it would have been without this feature. That is true for every refusal: no calibrated pair, an unreadable artifact, a mapper fitted for a different pair, a block of tokens too small to be worth moving. A transfer is an optimisation, and an optimisation is never allowed to fail a request.

Hardware NVIDIA (compute capability 8.0+) or AMD CDNA (Instinct). Radeon/RDNA is refused.
Model pair Same family, same tokenizer, same number of KV heads and the same per-head dimension. Layer count, hidden size and parameter count may differ freely.
Calibration One run per pair, on a GPU. Without it the pair is NOT_CALIBRATED and never transfers.

The pair requirements are checked when the pair is registered, so an impossible pair is refused with the specific reason rather than failing later.

  • An explicit model= in a request is never overridden.
  • Transfer never crosses model families, whatever the shapes say.
  • Per-tenant enablement, so a tenant can be excluded.
  • On hardware the gate refuses, the whole feature is off rather than partly on.

Current limitation — read this before planning a rollout

Section titled “Current limitation — read this before planning a rollout”

The transfer is implemented and measured end to end offline: the pilot’s cache is captured, mapped and installed, and the target decodes from it. The serving-path installer into the inference engine’s paged cache is written and unit-tested, but no live server calls it yet — wiring it into the running scheduler is the next step.

So today the console toggle, the calibration and the pair registry are real and usable, and a served request is not yet accelerated by a transfer. The capability is published this way deliberately: calibrating pairs and seeing their retention scores is the part you need before a rollout would make sense.