Speculative Decoding Explained: EAGLE-3 vs Medusa vs Lookahead
A modern GPU can do roughly three hundred floating-point operations for every byte it reads from memory, yet generating one token from a large language model asks it to do about one. That mismatch is why a single user’s chat stream crawls along while most of the silicon sits idle. Speculative decoding exploits the gap: a cheap drafter guesses several tokens ahead, the big model checks all of them in one forward pass, and a rejection-sampling rule guarantees the output distribution is exactly what the big model would have produced on its own.
It matters now because inference cost and latency, not training, dominate the bill for most production LLM systems, and the method family has matured fast. Draft-model speculation has given way to head-based schemes such as Medusa, feature-level drafters such as EAGLE-3, and draft-free approaches such as lookahead decoding, all of which now ship in serving engines like vLLM and SGLang.
You will leave with the speedup formula and what each term means, a mechanical explanation of why batch size can erase the gain, and a practical way to choose among the methods for your workload.
What this covers: the memory-bound argument, the draft-and-verify math, the four method families side by side, serving-engine support, failure modes, and a decision checklist.
Context and Background
Autoregressive decoding produces one token per forward pass of the model. Each pass must stream every weight from high-bandwidth memory (HBM) into the compute units, then read the growing key-value (KV) cache for attention. For a 70-billion-parameter model in 16-bit precision, that is about 140 GB of weights per token, before counting the KV cache. The arithmetic done on those weights is tiny: roughly two floating-point operations per parameter for a single sequence. The step is memory-bandwidth-bound, not compute-bound.
That observation is not new. The Medusa paper states it directly: each decoding step “necessitates moving the full model parameters from High-Bandwidth Memory (HBM) to the accelerator’s cache” (Cai et al., arXiv:2401.10774). The lookahead decoding authors make the same point, calling autoregressive decoding memory-bandwidth bounded (Fu et al., arXiv:2402.02057). Speculative decoding is the technique that turns this wasted compute into latency.
Two papers introduced the core algorithm independently in early 2023. Leviathan, Kalman and Matias at Google reported a 2X-3X acceleration on T5-XXL “without any changes to the outputs” (arXiv:2211.17192). Chen and colleagues at DeepMind called it speculative sampling and reported a 2-2.5x decoding speedup on the 70-billion-parameter Chinchilla model in a distributed setup (arXiv:2302.01318). Both used a smaller draft model from the same family and a modified rejection-sampling scheme that preserves the target model’s distribution.
The idea sits next to, but is distinct from, other serving optimizations. Continuous batching raises throughput by packing many requests into each step, which we covered in continuous batching for LLM inference. Speculative decoding instead shortens the latency of each individual stream. The two interact, and that interaction is the most important thing to understand before you turn the feature on in production. For a view of how serving engines compare in general, see our vLLM, TGI, SGLang and Triton benchmark.
Since 2023 the field has split along one question: where do the draft tokens come from? A separate small model (classic), extra prediction heads bolted onto the target (Medusa), a tiny autoregressive head that reads the target’s own hidden features (EAGLE), or the target itself running a parallel iteration that mines n-grams (lookahead). Each answer trades training cost, memory, acceptance rate and operational complexity differently, and the rest of this post works through those trade-offs.
How Speculative Decoding Works: The Draft-and-Verify Architecture
Speculative decoding runs a cheap drafter to propose gamma candidate tokens, then runs the target model once over all of them in parallel. A rejection-sampling rule accepts the longest valid prefix and adds one extra token, so each target pass yields between one and gamma plus one tokens while leaving the output distribution unchanged.

Figure 1: The draft-and-verify loop. The drafter proposes gamma tokens, the target scores all positions in one pass, and rejection sampling keeps the longest valid prefix plus one bonus or corrected token.
The loop in Figure 1 repeats until the sequence ends. The drafter is cheap by construction, so the cost of an iteration is dominated by one target forward pass, the same pass that would otherwise yield a single token. Whatever the drafter gets right is a free gain, and whatever it gets wrong costs only the small drafting overhead plus some wasted verification FLOPs.
Why verification is nearly free
A transformer processes all positions of an input sequence in parallel during prefill. Verification reuses that capability. Given the prompt plus gamma draft tokens, one forward pass produces the target model’s next-token distribution at every one of the gamma plus one positions simultaneously. Position i’s distribution is conditioned on the draft tokens before it, which is exactly the context the target would have seen had it generated those tokens itself.
The cost of that wider pass is small only because the step is memory-bound. The weights are streamed once regardless of whether the pass handles one token or five. In a memory-bound regime the extra FLOPs for the additional positions land on idle tensor cores. This is the entire economic basis of the technique, and it is also exactly what breaks at large batch sizes, as we will see.
The acceptance rule and why it is lossless
Let p(x) be the target model’s probability for token x at a given position and q(x) the drafter’s. The drafter samples a token x from q. The verifier accepts it with probability min(1, p(x)/q(x)). If p(x) is at least q(x), the token is always accepted, because the target likes it at least as much as the drafter did. If p(x) is smaller, the token is accepted with probability p(x)/q(x).
On rejection, the algorithm samples a replacement from the normalized residual distribution proportional to max(0, p(x) minus q(x)). A short calculation shows that the marginal distribution of the emitted token is exactly p. This is the property both original papers prove: the DeepMind paper describes it as a “modified rejection sampling scheme which preserves the distribution of the target model within hardware numerics.”
That qualifier matters in practice. The vLLM documentation states that speculative decoding sampling “is theoretically lossless up to the precision limits of hardware numerics,” and warns that floating-point differences, plus batch-size-dependent kernels, can cause slight variations in logprobs. In other words, the guarantee is distributional, not bitwise. Two runs with different speculation settings can differ token by token while both being valid samples from the same model.
If every draft token is accepted, the verifier has already computed the target distribution for position gamma plus one, so it samples one more “bonus” token for free. This is why the best case is gamma plus one tokens per target pass, and the worst case is one token: the first draft token is rejected, a corrected token is sampled from the residual, and you have made exactly the progress plain decoding would have.
The acceptance rate alpha
Define alpha as the expected probability that a single draft token is accepted. Leviathan and colleagues show it equals one minus the total variation distance between p and q, averaged over contexts. Intuitively it measures how often the drafter and target agree. A drafter that merely imitates the target’s most likely continuation scores well on predictable text such as boilerplate, code syntax and quoted passages, and poorly on high-entropy creative text.
Alpha is not a constant of the method. It varies with the task, the sampling temperature, the position within the draft and the context. Greedy decoding at temperature zero typically gives higher agreement than sampling at temperature one. We return to this in the failure-modes section, because quoting a single alpha for a method hides most of what matters.
The Speedup Formula, Derived
Assume each draft token is accepted independently with probability alpha. The number of tokens produced per target pass is then a truncated geometric count. With gamma draft tokens, the expected number is
E[tokens per pass] = (1 − alpha^(gamma+1)) / (1 − alpha).
To see why, note that the first token is always produced, the second only if the first draft was accepted (probability alpha), the third only if the first two were (alpha squared), and so on up to alpha to the power gamma. Summing the geometric series 1 + alpha + alpha^2 + … + alpha^gamma gives the expression above.
Now add cost. Let c be the cost of one drafter step relative to one target step. A full iteration costs one target pass plus gamma drafter steps, which is 1 + gamma times c in units of target-pass time. Plain decoding yields one token per unit. The expected wall-clock speedup is therefore
Speedup = (1 − alpha^(gamma+1)) / ((1 − alpha) × (1 + gamma × c)).
This is the form given by Leviathan et al., and it makes the design space legible. The numerator rewards high acceptance and long drafts with diminishing returns. The denominator punishes expensive drafters and long drafts linearly.
Worked numbers (illustrative, not measured)
The figures below are computed directly from the formula to show its shape. They are not benchmark results.
| alpha | gamma | c | Tokens per pass | Speedup |
|---|---|---|---|---|
| 0.5 | 4 | 0.05 | 1.94 | 1.62x |
| 0.6 | 4 | 0.05 | 2.31 | 1.92x |
| 0.7 | 4 | 0.05 | 2.77 | 2.31x |
| 0.8 | 4 | 0.05 | 3.36 | 2.80x |
| 0.8 | 4 | 0.20 | 3.36 | 1.87x |
| 0.9 | 7 | 0.05 | 5.70 | 4.22x |
Three lessons fall out. First, alpha dominates: moving from 0.6 to 0.8 at fixed gamma lifts the speedup from 1.92x to 2.80x. Second, drafter cost matters: a drafter four times more expensive (c of 0.20 instead of 0.05) drops the same alpha-0.8 configuration from 2.80x to 1.87x. Third, longer drafts help only when alpha is high.
The optimal draft length
For alpha of 0.8 and c of 0.05, evaluating the formula for gamma from 1 to 12 gives speedups of 1.71, 2.22, 2.57, 2.80, 2.95, 3.04, 3.08, 3.09 (at gamma of 8), 3.08, 3.05, 3.00 and 2.95. The curve peaks near gamma equals 8 and is flat within a few percent from 6 to 10. Past the peak, each added draft token is increasingly unlikely to survive to be used, yet it still adds drafting cost.
The independence assumption is a simplification. In real text, acceptance of token i+1 is conditional on token i having been accepted, and alpha falls with depth because the drafter drifts from the target the further it extrapolates. That is why every modern method either shortens drafts adaptively or branches into a tree, which we examine in the next section.

Figure 2: Where draft tokens come from in each method family. All four converge on the same verification step in the target model.
The figure shows that the families differ only upstream of verification. Downstream, the target model verifies a chain or tree of candidates using the same acceptance logic, which is what lets serving engines expose them behind one configuration switch.
The Four Method Families Compared
The method you pick determines your training burden, your memory footprint and, above all, your acceptance rate. The sections below take each in turn, using the numbers the authors themselves report. Those figures come from different models, hardware, tasks and batch sizes, so treat them as indicators of each method’s ceiling under favourable conditions, not as a ranking.
Classic draft-model speculation
The original approach pairs a large target with a small model from the same family, for example a 1B drafter for a 70B target. The two must share a tokenizer and vocabulary, because the verifier compares probabilities over the same token IDs. TensorRT-LLM’s documentation describes the flow plainly: the draft model is queried to generate up to K draft tokens, which are then forwarded to the target model for verification.
The strengths are simplicity and zero training: if a well-matched smaller sibling exists, you can deploy it immediately. The weaknesses are real. The drafter needs its own weights in GPU memory and its own KV cache. It runs autoregressively, so drafting gamma tokens takes gamma sequential small-model passes, which inflates c. And a generic small model is a mediocre imitator of a large model that was post-trained differently, so alpha is often lower than you would like unless the drafter is distilled against the target.
Medusa: extra heads on the target
Medusa removes the separate model. It adds several lightweight decoding heads on top of the target’s final hidden state, where head k predicts the token k steps ahead. One forward pass of the target plus the heads yields candidates for several future positions at once. The heads are small multi-layer perceptron blocks, so memory overhead is modest and there is no second KV cache.
The candidates are arranged into a tree and verified with tree attention, a mask that lets many branches be scored in a single pass while each branch attends only to its own ancestors. The paper reports two training regimes. Medusa-1 trains the heads on a frozen backbone and achieves over 2.2x speedup while remaining lossless. Medusa-2 fine-tunes the heads jointly with the backbone and reports 2.3 to 3.6x, at the cost of a more delicate training recipe that must avoid degrading the base model. The paper also describes self-distillation for cases where the original training data is unavailable, and a “typical acceptance” scheme that raises acceptance by relaxing the strict rule.
Note the last point carefully. Typical acceptance accepts tokens that are merely plausible under the target, which is not the exact rejection-sampling rule. It trades the strict distributional guarantee for a higher acceptance rate. If you need lossless sampling, confirm which acceptance mode your serving engine uses.
The structural weakness of Medusa is that its heads predict independently from the same hidden state. Head 3 does not see the token that head 2 chose, so it must predict position t+3 without knowing what occupies t+2. The joint probability of a candidate path is only loosely modelled, which limits accuracy at depth and is the problem EAGLE set out to fix.
EAGLE, EAGLE-2 and EAGLE-3: autoregression on features
EAGLE keeps a small autoregressive drafter, but instead of working on tokens it works on the target’s hidden features. The drafter is a single transformer layer that takes the target’s second-to-top-layer features plus the already-sampled tokens and predicts the next feature, which the target’s own output head converts into a token distribution. The EAGLE paper argues that autoregression at the feature level is easier than at the token level, and that feeding in the advanced token sequence resolves the sampling uncertainty that otherwise hurts feature prediction. It reports a latency speedup of 2.7x to 3.5x on LLaMA2-Chat 70B and roughly doubled throughput, with output quality preserved (arXiv:2401.15077).
EAGLE-2 changed the tree, not the drafter. EAGLE-1 used a static tree shape that assumed token acceptance depends only on position. EAGLE-2 observed that acceptance depends on context and that the drafter’s confidence scores approximate acceptance rates with small errors, so it builds the tree dynamically, expanding branches the drafter is confident about. It reports speedup ratios of 3.05x to 4.26x across three LLM series and six tasks, a 20 to 40 percent improvement over EAGLE-1, and remains lossless (arXiv:2406.16858).
EAGLE-3, submitted in March 2025 by the same group, makes two changes that target a scaling problem. The authors found that EAGLE’s reliance on feature prediction limited how much a drafter could benefit from more training data. EAGLE-3 drops the feature-prediction objective in favour of direct token prediction, replaces top-layer features with a fusion of low-, middle- and high-layer features from the target, and uses a “training-time test” technique in which the drafter is trained while consuming its own predictions, so that training matches the way it is used at inference. The paper reports a speedup ratio of up to 6.5x, about 1.4x better than EAGLE-2, and a 1.38x throughput improvement at batch size 64 in SGLang, evaluated on chat and reasoning models across five tasks (arXiv:2503.01840).
The 6.5x figure is a best case on the authors’ favourable tasks and settings. The batch-64 figure, a 1.38x throughput gain, is the more useful number for capacity planning, because it shows the method still helps when the server is busy. It is one paper’s measurement, not a guarantee for your workload.
Operationally, EAGLE-family drafters must be trained per target model. Because the drafter reads the target’s internal features, a drafter trained for one checkpoint will not transfer to another. This is the main reason the approach has a supply-chain dimension: you depend on someone publishing a matched drafter for the exact model version you serve, or you train one yourself.
Lookahead decoding: no drafter at all
Lookahead decoding, from Fu, Bailis, Stoica and Zhang, takes a different route. It needs neither a draft model nor an external data store. It runs Jacobi iteration, a fixed-point method that treats the next several positions as unknowns and refines all of them in parallel at each step. Over successive steps the guesses at each position form trajectories, and n-grams harvested from those trajectories go into a pool.
Two branches run in the same forward pass. The lookahead branch maintains a two-dimensional window controlled by a window size W (how many future positions) and an n-gram size N (how many past iterations to reference). The verification branch takes candidate n-grams from the pool whose first token matches the current position and confirms them against the model’s actual output. Because both branches share a pass, no second model is needed.
The reported results are 1.5x to 2.3x latency reduction in the LMSYS write-up, with code generation seeing 2x or more on CodeLlama, and up to 1.8x on MT-bench plus up to 4x on multi-GPU code completion in the paper’s abstract. The price is stated candidly by the authors: the technique “trades per-step log(FLOPs)” for fewer steps, so it needs a GPU with surplus compute and gets worse if W and N are set too large. It shines in latency-sensitive, small-batch settings on powerful hardware and loses its advantage where compute is already contended.
N-gram and prompt-lookup drafting
A cheaper cousin deserves a mention because it often wins in practice. N-gram or prompt-lookup drafting proposes the tokens that followed the same n-gram earlier in the prompt or generation. TensorRT-LLM notes it works best where input and output overlap heavily, such as summarization, document question answering, multi-turn chat and code editing. There is no model, no training and almost no overhead. When the output copies from the input, which retrieval-augmented and editing workloads do constantly, it delivers a real speedup at essentially zero cost, and vLLM’s documentation lists it as a method that offers modest speedups without increasing workload during peak traffic.
Side-by-side summary
| Method | Draft source | Extra training | Extra memory | Reported ceiling (authors) | Best fit |
|---|---|---|---|---|---|
| Draft model | Separate small LLM | None if sibling exists | Second model plus KV cache | 2-3x (T5-XXL, Chinchilla) | Matched family, quick start |
| Medusa | Parallel heads on target | Heads, Medusa-2 also backbone | Small head weights | 2.2x (M1), 2.3-3.6x (M2) | Single-model deployments, simple |
| EAGLE-3 | One-layer feature-fused drafter | Per-target drafter | Small drafter plus KV | Up to 6.5x, 1.38x at batch 64 | Latency-critical chat and reasoning |
| Lookahead | Jacobi n-gram pool | None | Pool and larger step | 1.5-2.3x, 4x multi-GPU code | No-training, spare compute |
| N-gram lookup | Prompt and history | None | Negligible | Workload dependent | RAG, summarization, editing |
The “reported ceiling” column mixes models, hardware and benchmarks, so it is not a head-to-head comparison. Treat it as each paper’s own best claim.
Why Batch Size Decides Everything
The single most misunderstood fact about speculative decoding is that it is a latency tool whose benefit shrinks as the server gets busier. The reason follows directly from the memory-bound argument.

Figure 3: At small batch the GPU is memory-bound and verification is nearly free. At large batch the tensor cores are saturated and every rejected draft token is paid for in real compute.
A roofline estimate
Take an H100 SXM-class GPU with roughly 989 teraFLOPs of dense BF16 compute and roughly 3.35 TB/s of HBM3 bandwidth, from NVIDIA’s published specifications. The ratio is about 295 FLOPs per byte, the “ridge point” above which a kernel stops being bandwidth-limited. In BF16 each weight is two bytes and is used for two FLOPs per token in the batch, so the arithmetic intensity of a decode step is roughly the batch size. Real kernels reach the ridge at a lower batch than the idealized figure because of attention, KV reads and overheads, but the logic holds: at batch 1 you are using about one three-hundredth of the available compute.
Now apply speculation. A server running 16 concurrent sequences with gamma of 4 verifies 16 times 5, which is 80 token positions per step instead of 16. Both numbers sit far below the ridge, so the step time barely changes and the full speedup is available. A server running 256 sequences verifies 1,280 positions per step. That is well past the ridge, the step is now compute-bound, and the extra 1,024 positions cost real time. Every rejected draft token is wasted work the server could have spent on another user’s request.
Latency versus throughput
This produces a crossover. At low load, speculation improves per-request latency and also total throughput, because idle compute is converted into progress. As load rises, per-request latency still improves for the speculating request but total system throughput can fall below the non-speculative baseline, because compute is spent on tokens that get thrown away. The vLLM documentation’s performance matrix reflects this: EAGLE and multi-token prediction deliver strong gains at low queries per second, while n-gram and suffix decoding give modest speedups without increasing workload during peak traffic.
The EAGLE-3 batch-64 result is consistent with this picture. A 1.38x throughput gain at batch 64 shows the crossover has not arrived by then for that model and hardware, but the gain is far below the single-stream speedup, which is exactly the shape the roofline predicts. Where your own crossover lies depends on model size, GPU, sequence length and acceptance rate, so it must be measured, not assumed.
Adaptive speculation
Because the right draft length depends on instantaneous load and acceptance, the engines have grown dynamic controls. vLLM lists dynamic speculative decoding and adaptive verification among its features, and SGLang auto-tunes its step, top-k and draft-token parameters when you omit them. The principle is simple: when the server is idle, draft aggressively; when it is saturated, shorten drafts or switch speculation off per request. If your traffic is bursty, a fixed gamma tuned at 3 a.m. will hurt you at noon.
Serving-Engine Support and Configuration
All three major open serving stacks now treat speculation as a first-class feature, although their method lists and flag names differ. Check the version you run, because this surface is changing quickly and the lists below reflect the documentation read for this post in October 2026.

Figure 4: One speculative step inside a serving engine. The scheduler asks the proposer for draft tokens, the target verifies k plus 1 positions, the rejection sampler returns an accepted prefix, and rejected KV entries are rolled back.
The sequence in Figure 4 hides a subtle cost: KV-cache bookkeeping. Verification writes KV entries for every draft position, but only the accepted prefix is valid. The engine must roll back the rest and, for draft-model methods, also keep the drafter’s cache consistent. This is why speculation interacts with paged attention, prefix caching and chunked prefill, and why each engine’s feature-compatibility matrix deserves a read.
vLLM
The vLLM documentation lists supported approaches including EAGLE, multi-token prediction (MTP), draft model, parallel draft model, multi-layer perceptron heads, n-gram and suffix decoding, with additional dynamic speculative decoding and adaptive verification features. Configuration is a small JSON object passed through the speculative config. These examples are taken from the official documentation:
{ "method": "eagle3", "num_speculative_tokens": 5 }
{ "method": "ngram", "num_speculative_tokens": 4,
"prompt_lookup_min": 2, "prompt_lookup_max": 5 }
{ "method": "draft_model", "model": "<draft-model>",
"num_speculative_tokens": 5 }
The num_speculative_tokens value is gamma. Note that the exact way you pass this JSON (command-line flag versus Python argument) has changed across vLLM releases, so follow the documentation for your installed version rather than copying an older blog post. The vLLM docs also carry the batch-size caveat quoted earlier: changes in batch size can cause variations in logprobs and output probabilities.
SGLang
SGLang’s guide lists EAGLE-2, EAGLE-3, MTP, a LoRA-adapter method it calls UNO, a linear-block-verification method, standalone draft models, n-gram (CUDA only) and a speculative decoding V2 variant built on the overlap scheduler. Method selection is a single flag, with draft weights supplied separately where required:
--speculative-algorithm EAGLE3
--speculative-draft-model-path <path>
--speculative-num-steps 3
--speculative-eagle-topk 1
--speculative-num-draft-tokens 4
The three tuning flags map onto the tree: num-steps is the autoregressive drafting depth, eagle-topk is the branching factor per step, and num-draft-tokens is the maximum number of candidates sent to verification. The documentation says these are auto-tuned if omitted. Its published example, on LLaMA-Instruct 3.1 8B with MT-bench on a single H100, shows a 158.34 tokens/s baseline, 244.10 tokens/s with EAGLE-2 (+54%) and 373.25 tokens/s with EAGLE-3 (+136%). That is a vendor-maintained single-GPU, single-stream number, which is the most favourable regime for the technique.
A useful sanity check on that baseline: an 8-billion-parameter BF16 model is about 16 GB, and 16 GB at 3.35 TB/s takes roughly 4.8 ms, a ceiling near 209 tokens/s. The measured 158 tokens/s baseline sits below that ceiling, as expected once KV reads, attention and launch overheads are added, and it confirms the step is bandwidth-bound. Reaching 373 tokens/s means averaging well over two tokens per target pass.
The SGLang guide also warns about memory. The draft tree and extra CUDA graphs consume VRAM, and for out-of-memory errors it recommends, in order, lowering the static memory fraction, reducing the maximum decode CUDA-graph batch size, and shrinking the draft tree.
TensorRT-LLM
NVIDIA’s TensorRT-LLM documentation lists draft-target-model, n-gram, Medusa, ReDrafter, EAGLE and lookahead decoding. ReDrafter, a method from Apple, uses recurrent prediction in which each draft token depends on the previous one, with beam search inside the engine. The docs describe Medusa’s TopK approach as forming a tree of candidate paths, and EAGLE as a single-layer transformer predicting draft tokens from previous hidden states and decoded tokens. Support for multi-token prediction in this specific document was not stated, so check the release notes of the version you deploy.
Multi-token prediction and built-in heads
A notable recent trend is models that ship with their own speculation heads, trained during pre-training as multi-token-prediction (MTP) modules. Both vLLM and SGLang list MTP as a method. Where a model provides one, it removes the supply-chain problem of the EAGLE family, because the drafter is trained against exactly that checkpoint by the same team that trained the model. Which specific models ship such heads, and with what measured acceptance, is something to verify per model card rather than assume.
Deeper Analysis: Trees, Acceptance Depth and Drafter Quality
The headline speedups rest on three levers that interact. Understanding them explains why EAGLE-3 outperforms Medusa on the same target and why tuning matters more than method choice.
Chains versus trees
A chain proposes one candidate per position. If the drafter’s first choice at position 2 is wrong, everything after it is wasted. A tree proposes several candidates per position and verifies all branches at once using a tree attention mask, so the verifier can pick whichever branch matches.
The expected accepted length of a tree is higher than a chain with the same depth, but the verification width grows. A tree with 60 candidate nodes is verified in one pass, so it adds 60 positions of work, and that is cheap at batch 1 and expensive at batch 64. This is the lever the SGLang flags expose: top-k widens the tree and num-draft-tokens caps its size. EAGLE-2’s contribution was making width follow confidence rather than a fixed template, so compute goes where acceptance is likely.
Acceptance depends on position and temperature
Acceptance is not uniform along the draft. The first draft token is the easiest to predict because the drafter has the most context, and each later token compounds the previous uncertainty. In practice measured per-position acceptance curves slope downward, so the independence assumption in our formula overstates the benefit of long chains. That is precisely why a dynamic tree beats a static chain.
Temperature has an equally strong effect. At temperature zero the target’s distribution is a point mass on its argmax, and a good drafter matches it often. At temperature one on open-ended text the target distribution is broad, the drafter’s distribution is broad, and their overlap, one minus total variation distance, is low. This is why the headline numbers, which often come from benchmark prompts at temperature zero, can overstate what creative-writing traffic will see. Measure alpha on your own prompts at your own sampling settings.
Drafter quality and the training-data lesson
The EAGLE-3 paper’s central observation is a scaling one: a drafter trained to predict features did not improve with more data the way a token predictor does. By training directly on token prediction with multi-layer features and the training-time-test procedure, EAGLE-3 turns additional training data into acceptance. The practical corollary is that drafter training is a data-and-distribution problem. A drafter trained on generic chat data underperforms on your domain, and one trained on the target’s own outputs for your traffic can do much better. Several teams therefore fine-tune or distil drafters on their production logs, subject to privacy constraints.
Interactions with the rest of the stack
Speculation does not live in isolation. In disaggregated or gateway-routed deployments, such as those built with the Gateway API Inference Extension described in our Istio ambient multicluster LLM serving guide, the router sees only queue depth and cache hits, not per-request acceptance, so load balancing can concentrate low-acceptance traffic on some replicas and create uneven latency. On consumer or workstation hardware, memory bandwidth is the binding constraint and speculation helps most, which is relevant to the setups discussed in our AMD Ryzen AI Max local LLM workstation guide. Unified-memory machines have lower bandwidth than data-center HBM, so the memory-bound argument is even stronger there, though VRAM left for the drafter is also tighter.
Quantization and long contexts
Two further interactions are worth knowing. Quantizing the target model reduces bytes per token and therefore the baseline step time, which lowers the headroom speculation can exploit; a 4-bit target is already faster per step, so the relative gain from speculation is smaller than for the same model in BF16, though absolute latency is still better. The drafter must also be matched to the quantized target’s behaviour, since quantization shifts logits slightly and acceptance can drop if the drafter was trained against the full-precision model.
Long contexts shift the balance the other way. As the KV cache grows, attention reads become a larger share of each step, and those reads are paid once per pass regardless of how many positions are verified. That makes verification of extra positions even cheaper relative to the baseline step, at least until the tree is wide enough that attention over many query positions becomes costly. In short-answer workloads over long prompts, such as document question answering, speculation tends to be attractive; n-gram drafting is especially strong there because answers quote the document.
Trade-offs, Gotchas, and What Goes Wrong
Speculation can make a busy server slower. Past the compute-bound crossover, rejected tokens consume FLOPs that would have served other requests. Teams that benchmark at batch 1, see a 3x gain and enable it fleet-wide sometimes watch aggregate throughput fall at peak. Always benchmark at your real concurrency and compare tokens per second per GPU, not just time to finish a single request.
Acceptance is workload-specific. A drafter that shines on code and structured output can disappoint on high-temperature creative generation. If your traffic mixes both, a single global setting hides the loss. Log accepted tokens per verification pass in production and alert on a drop, because a model or prompt-template change can quietly collapse alpha.
The lossless guarantee is distributional and numerically soft. Exact rejection sampling preserves the target distribution, but batch-dependent kernels mean outputs can differ between runs, which matters for regression tests that compare strings. Methods using relaxed or typical acceptance, as Medusa describes, are not strictly lossless at all. Know which mode you are in before you claim equivalence in an evaluation.
Drafters are tied to a target checkpoint. EAGLE-family and Medusa heads are trained against specific weights. If you fine-tune the target, swap in a new quantization or upgrade a minor version, the old drafter’s acceptance can drop sharply. Pin the pair together and re-validate on every target change.
Memory pressure is real. Draft models, draft KV caches, extra CUDA graphs and wide trees all take VRAM, shrinking the KV cache available for concurrency. That can lower the number of sequences you can batch, which is its own throughput cost. SGLang’s own OOM guidance, which starts by lowering the static memory fraction, shows how routinely this bites.
Structured decoding and tool use complicate verification. Grammar-constrained or JSON-schema decoding modifies the target distribution token by token. The drafter does not know the grammar, so it may propose tokens the constraint forbids, lowering acceptance. Check your engine’s documented compatibility between constrained decoding and the speculation method you choose.
Short outputs get little benefit. The fixed cost of drafting and tree building is amortized over generated length. For classification or one-word answers there is nothing to speculate on, and the overhead can only hurt.
Reported speedups are best cases. Every number in this post comes from the authors or engine maintainers’ chosen models, tasks and hardware. None of them is a promise for your stack. The honest framing is a range: roughly 1x (or below) in the worst conditions to a few times in the best, with the position inside that range set by batch size, acceptance and drafter cost.
Practical Recommendations
Start by classifying your workload on two axes: how loaded your GPUs are, and how predictable your outputs are. Latency-sensitive, lightly loaded interactive chat and agent loops are the sweet spot for a trained EAGLE-3 drafter or a model’s built-in MTP head. Heavily loaded batch endpoints, such as offline summarization at saturation, gain little from heavy speculation and may lose, while an n-gram method can still pay if outputs copy from inputs.
If a trained drafter exists for your exact target, use EAGLE-3 first. If you serve a single model, cannot train per-checkpoint drafters, and want something simple, Medusa-style heads or a matched draft model are reasonable. If you want no training and have spare compute on a strong GPU, lookahead decoding is worth a trial, especially for code. For retrieval-augmented generation and editing tasks, try n-gram lookup before anything heavier, because it is nearly free.
Treat configuration as a tuning problem, not a switch. Start with the defaults the engine auto-tunes, then sweep draft length or tree size against your own traffic replay. The formula in this post tells you what to expect: if measured speedup is far below the prediction from your observed alpha and drafter cost, something other than acceptance, such as scheduling overhead or memory pressure, is the culprit.
A short checklist before rollout:
- Replay production prompts at production sampling settings and record mean accepted tokens per pass.
- Benchmark at the concurrency you run at peak, and compare aggregate tokens per second per GPU, not single-stream latency.
- Confirm lossless versus relaxed acceptance for your engine and method.
- Pin the drafter to the exact target checkpoint and quantization, and re-test after any change.
- Budget VRAM for the drafter and wider tree, and check the effect on maximum concurrency.
- Enable dynamic or adaptive speculation where available, and keep a per-request kill switch.
- Alert on acceptance rate drift in production.
Frequently Asked Questions
What is speculative decoding in LLM inference?
Speculative decoding is a technique that speeds up LLM inference by letting a cheap drafter guess several upcoming tokens, then having the large target model verify them all in one parallel forward pass. A rejection-sampling rule accepts the correct prefix, so the output distribution matches ordinary decoding. It works because decoding is memory-bandwidth-bound, leaving compute idle that verification can use for free.
Does speculative decoding change the model’s output quality?
With the standard rejection-sampling acceptance rule, it does not change the output distribution, and the original papers prove this. In practice the guarantee holds up to hardware numerics, so tiny logprob differences can appear, as vLLM’s documentation notes. Methods that use relaxed or typical acceptance, such as an option described in the Medusa paper, trade strictness for higher acceptance and are not strictly lossless.
How much faster is EAGLE-3 than Medusa or lookahead decoding?
The papers report different models, tasks and hardware, so no clean head-to-head exists in the sources reviewed. EAGLE-3’s authors report up to 6.5x and about 1.4x over EAGLE-2, Medusa-2 reports 2.3 to 3.6x, and lookahead reports about 1.5x to 2.3x in its own write-up. Benchmark on your own model and traffic; these are each method’s best-case claims.
Why does speculative decoding help less at high batch sizes?
At small batch, decoding is memory-bound, so verifying extra draft tokens uses idle compute. At large batch the tensor cores are already busy, and each rejected draft token wastes compute that could serve another request. Per-request latency may still improve, but aggregate throughput can fall below baseline. The EAGLE-3 paper reports only a 1.38x throughput gain at batch size 64, versus much larger single-stream gains.
Which is better, a draft model or head-based speculation?
A separate draft model needs no training if a compatible smaller sibling exists, but costs extra memory, a second KV cache and sequential drafting steps. Head-based methods like Medusa and EAGLE-3 reuse the target’s hidden states, so they are lighter and usually reach higher acceptance, but need a drafter trained for your exact checkpoint. Choose based on whether a matched, trained drafter exists for your model.
Which serving engines support speculative decoding?
vLLM supports EAGLE, MTP, draft model, n-gram, suffix decoding and others. SGLang supports EAGLE-2, EAGLE-3, MTP, standalone draft models and n-gram. TensorRT-LLM documents draft-target, n-gram, Medusa, ReDrafter, EAGLE and lookahead decoding. Method lists and flags change between releases, so confirm against the documentation for the exact version you deploy.
Further Reading
- LLM inference benchmark of vLLM, TGI, SGLang and Triton for engine-level throughput context.
- Continuous batching architecture for LLM inference for how batching interacts with speculation.
- AMD Ryzen AI Max local LLM inference workstation for bandwidth-limited local hardware.
- Istio ambient multicluster and Gateway API Inference Extension for LLM serving for routing speculating replicas.
- Leviathan et al., Fast Inference from Transformers via Speculative Decoding and Li et al., EAGLE-3, the primary sources.
- vLLM speculative decoding documentation for current methods and configuration.
By Riju — about
