// guess ahead, check together
Generation is serial — each token depends on the one before it. Verification is not: a model can score several proposed tokens in a single forward pass. Speculative decoding trades one against the other.
Everything after the first rejected token goes too, because it was written on a premise that turned out to be false. Then the cycle repeats from there.
The output is the target model's own
This is the strong claim, and it is not a hand-wave about quality being "close enough". The modified rejection-sampling scheme preserves the distribution of the target model within hardware numerics — you are sampling from the big model, just arriving there faster.
So the usual instinct — "a smaller model is involved, so the quality must dip" — is wrong here. The draft never decides anything. It only proposes, and every proposal is either confirmed by the big model or thrown away.
The speed is not guaranteed
Here is where the headline numbers stop transferring. A study on consumer hardware measured five configurations end to end:
Two causes, and neither is exotic. A draft whose per-step framework overhead exceeds the entire decode step of an already-small target — you pay for the guess and save nothing. And a quantized Metal backend that ran "parallel" verification serially below batch size 8: verify cost grew almost linearly, 51 ms for a 2-token batch and 218.5 ms for 7, before flattening to 141–142 ms at 8 and above. The mechanism assumes verification is batch-parallel. When the backend does not deliver that, the whole argument collapses.
Guessing further ahead backfires
The obvious knob is K, how many tokens the draft proposes per round. Bigger is not better, because acceptance falls as the draft runs further from the target's own reasoning:
Every rejected token is work the big model did and threw away, so K has an interior optimum rather than a direction. Note also that their best throughput sat at 37.8% acceptance — the acceptance rate on its own is not the thing to maximise.
So: measure it
The practice is settled even where the results are not. Draft from the same model family and keep it small — roughly 0.5B to 3B against a much larger target; mismatched tokenizers need the heterogeneous-vocabulary variants rather than the plain algorithm. Then benchmark on your machine, with your backend and your quantization, because those are the variables that decided the sign in every configuration above.
And budget the memory. The draft model is resident too — peak usage across those five configurations ranged 2.93 GB to 5.49 GB — which on this hardware is the scarce resource, not an afterthought. the KV cache, which your weights do not show you and active is not what you keep cover the rest of that ledger.
Sources: Leviathan et al. — Fast Inference from Transformers via Speculative Decoding · Chen et al. — Accelerating Large Language Model Decoding with Speculative Sampling (arXiv 2302.01318) · Lossless but Not Free: An Empirical Anatomy of Speculative Decoding on Consumer Hardware (arXiv 2607.17283) · Accelerating LLM Inference with Lossless Speculative Decoding Algorithms for Heterogeneous Vocabularies (arXiv 2502.05202). Verified 2026-09-02.
One concept a week. Free.
The deeper, copy-paste version of each ToolCall short — in your inbox.
// total: 0.00 · spam: void · unsubscribe: one click
