Speculative decoding
Speculative decoding speeds up generation without changing the output · a cheap draft model guesses several tokens and the main model checks them all at once.
Text reviewed October 5, 2026
Speculative decoding speeds up generation without changing the output · a cheap draft model guesses several tokens and the main model checks them all at once.
Basic
Large models generate one token at a time, which is slow. With speculative decoding, a smaller draft model proposes the next few tokens; the large model verifies them in a single forward pass and keeps the ones it agrees with. When the guesses are good, several tokens come out for the price of one large-model step.
Deep
With the standard acceptance rule, the output follows the same distribution as the large model alone, so quality is unchanged. Speed-ups depend on how often the draft is right, which is higher for predictable text such as code.
Expert
Variants replace the separate draft model with extra prediction heads or with lookup from the prompt. Providers use these techniques behind the scenes, which is one reason the same model can run at very different speeds on different hosts.
Depending on why you're here
- ·A trick that makes AI answer faster
- ·Faster responses at the same quality; check provider speed numbers
- ·Lower serving cost per token
- ·Output distribution preserved with the standard acceptance rule