SEED: Self-Speculative Decoding via Implicit Encoder–Decoder
Hankun Lin ⋅ Patrick Pynadath ⋅ Ruqi Zhang
Abstract
Self-speculative decoding accelerates large language model (LLM) inference by drafting tokens from the target model itself, but faces a sharp tradeoff between the quality and cost of the draft. Early-exit methods produce drafts cheaply by terminating computation at intermediate layers, but forgo the deeper representations that later layers provide and thus suffer in draft quality. Multi-token prediction preserves draft quality by emitting from the model's final hidden states, but pays for a full forward pass to produce those states at every drafting step. We propose ***s****elf-sp****e****culative ****e****ncoder-****d****ecoder* (**SEED**), a self-speculative method that obtains high-quality drafts cheaply by reusing the contextual representations already computed during verification. We reinterpret the standard decoder-only transformer as an implicit encoder–decoder: the first layers (encoder) build deep contextual representations, and the last few layers (decoder) emit tokens from them. Encoding and verification are merged into a single step: verification is performed by the full encoder–decoder, and the contextual representations of the verified prefix are cached for reuse during drafting. Drafting is therefore very fast: between verifications, the decoder drafts multiple tokens autoregressively, each conditioned on the cached representations and on preceding drafts. Experiments across multiple benchmarks show that SEED achieves up to 2.5$\times$ average speedup on 4B-scale models, outperforming both early-exit and MTP-style self-speculative baselines and running 17\% faster than the state-of-the-art EAGLE-3, while preserving or even improving the generation quality of standard autoregressive fine-tuning.
Chat is not available.
Successful Page Load