ConceptsReading · ~3 min · 38 words deep

Speculative decoding

Speculative decoding speeds up generation without changing the output · a cheap draft model guesses several tokens and the main model checks them all at once.

Text reviewed October 5, 2026

TL;DR

Speculative decoding speeds up generation without changing the output · a cheap draft model guesses several tokens and the main model checks them all at once.

Level 1

Large models generate one token at a time, which is slow. With speculative decoding, a smaller draft model proposes the next few tokens; the large model verifies them in a single forward pass and keeps the ones it agrees with. When the guesses are good, several tokens come out for the price of one large-model step.

Level 2

With the standard acceptance rule, the output follows the same distribution as the large model alone, so quality is unchanged. Speed-ups depend on how often the draft is right, which is higher for predictable text such as code.

Level 3

Variants replace the separate draft model with extra prediction heads or with lookup from the prompt. Providers use these techniques behind the scenes, which is one reason the same model can run at very different speeds on different hosts.

The takeaway for you
If you are a
Curious · Normie
  • ·A trick that makes AI answer faster
If you are a
Builder
  • ·Faster responses at the same quality; check provider speed numbers
If you are a
Investor
  • ·Lower serving cost per token
If you are a
Researcher
  • ·Output distribution preserved with the standard acceptance rule
Not with the standard method: the verification step keeps the output equivalent to the large model.