Speculative decoding
/AI/5 min read
A small model guesses the next few tokens and a large one checks them all in a single pass. Done properly the output is not an approximation of the large model — it is exactly what the large model would have produced.
Generating a token requires reading every weight in the model out of memory. Doing the arithmetic on those weights takes far less time than fetching them, so for most of each token the compute units sit idle waiting on memory.
That idle capacity is what speculative decoding spends. A small model guesses several tokens ahead; the large model checks all of the guesses in one pass, which costs it barely more than checking one.
The round
The draft model generates its guesses one at a time, as any model must. But the large model sees all of them at once, as a sequence, and produces its own prediction for every position in a single forward pass — the same pass it would have spent on one token.
Tokens are accepted up to the first disagreement. That one is replaced with the large model's choice, and everything after it is thrown away — those guesses were conditioned on a token that turned out to be wrong.
So a round of guesses yields between 1 and tokens.
Why the output is not an approximation
The obvious worry is that this trades quality for speed. It does not, and the reason is worth spelling out.
Each draft token carries two probabilities: , what the draft model gave it, and , what the large model gives it. The rule is:
If the large model likes the token at least as much as the draft model did, it is always accepted. If it likes it less, it is accepted in proportion.
On a rejection, the replacement is not simply the large model's top choice. It is sampled from the residual distribution, renormalised — what is left of the large model's opinion after subtracting what the draft model already had a chance to propose.
Those two rules together make the output distribution exactly . Not close to it. To check rather than assume, here is the whole scheme run twenty million times over four tokens:
| token | measured | error | ||
|---|---|---|---|---|
| alpha | 0.50 | 0.15 | 0.150152 | 1.5e−4 |
| beta | 0.20 | 0.35 | 0.349640 | 3.6e−4 |
| gamma | 0.20 | 0.10 | 0.100065 | 6.5e−5 |
| delta | 0.10 | 0.40 | 0.400143 | 1.4e−4 |
The draft model gave alpha half its probability mass and the large model gave it 0.15 — and 0.15 is what comes out. The draft model's opinion has no influence on the result whatsoever. It only affects speed.
There is a tidy consequence. The acceptance rate is , the overlap between the two distributions. For the numbers above that is 0.55, and the simulation measured 0.55018. A draft model helps exactly as much as it agrees.
What the speedup actually is
Let be the per-token acceptance rate, the number of guesses per round, the time for one large-model pass and the time per draft token.
Tokens per round is a truncated geometric series, and it comes out clean:
With ms and ms:
| best | |||||
|---|---|---|---|---|---|
| 0.95 | 3.23× | 3.77× | 4.11× | 4.42× | 15 → 4.48× |
| 0.90 | 2.93× | 3.26× | 3.40× | 3.39× | 10 → 3.43× |
| 0.80 | 2.40× | 2.47× | 2.40× | 2.15× | 6 → 2.47× |
| 0.70 | 1.98× | 1.91× | 1.78× | 1.50× | 4 → 1.98× |
| 0.50 | 1.38× | 1.24× | 1.11× | 0.91× | 2 → 1.46× |
| 0.30 | 1.02× | 0.89× | 0.79× | 0.65× | 1 → 1.18× |
Read down a column and the acceptance rate dominates everything. Read across a row and the draft length has an optimum that moves: at 95% agreement it pays to guess fifteen tokens ahead, at 50% it pays to guess two, and guessing twelve at 50% agreement is slower than not bothering at all.
That is the trap. A long draft is not cautious, it is a bet — every token after the first rejection is work that gets thrown away, and the further ahead you guess the more of it there is.
With the whole thing breaks even at about 38% acceptance. Below that the drafting costs more than it saves.
At a realistic 80% and : 3.95 tokens per 64 ms round, against 158 ms for the same tokens from the large model alone. Over a 300-token reply, 12.0 seconds becomes 4.9.
What it costs
The draft model sits in GPU memory alongside the large one, and it earns nothing except by agreeing.
Both models must share a tokeniser. Agreement is compared token by token, so two models that split text differently cannot be compared at all.
The gain also depends on there being idle compute to use. At low batch sizes memory bandwidth is the limit and the spare capacity is real. Under heavy batching the GPU is already busy, verification is no longer nearly free, and the speedup shrinks.
And on very short replies the fixed overhead of running two models is a larger share of the total.
An alternative avoids the second model entirely: let the large model draft for itself, either through extra output heads predicting several positions at once, or by drafting from its own internal representations. Same round structure, same verification, no separate set of weights to host.
The short version
- Generating one token is limited by reading the weights, not by the arithmetic, so compute sits idle.
- A small model drafts several tokens; the large model verifies them all in one pass.
- Tokens are accepted until the first disagreement; everything after it is discarded.
- Accept with probability and resample rejections from , and the output distribution is exactly the large model's.
- The acceptance rate is the overlap between the two models' distributions.
- Speedup is — and the best draft length falls as agreement falls.
- Draft too far ahead with a poorly matched model and it is slower than doing nothing.