Gumbo: Gumbel Optimized High-Temperature Speculative Sampling
Jonah Yi ⋅ Dan Fu ⋅ Yu-Xiang Wang
Abstract
Speculative sampling is a widely used method for losslessly accelerating Large Language Model (LLM) inference. State-of-the-art speculative sampling methods (e.g. EAGLE series) use a ``dynamic draft tree'' with drafter temperature implicitly set to $T=0$. Greedy drafts, however, perform poorly as temperature increases. Moreover, we show that naively instantiating the dynamic draft-tree with drafter temperature $T>0$ leads to an incorrect distribution. Fixes exist, but are computationally prohibitive as they require $O(|\mathcal{V}|^k)$ computation, where $\mathcal{V}$ and $k$ denote vocab and number of child nodes. We address these limitations by introducing Gumbo, a new speculative sampling algorithm designed for high-temperature settings that uses only $O(|\mathcal{V}|)$ compute. Gumbo is drafter-invariant in that it does not require the draft distribution during verification. Additionally, Gumbo increases the average acceptance length and speedup by expanding and reranking the dynamic draft tree based on Gumbel scores rather than draft probabilities. At the heart of Gumbo is a multi-draft generalization of the communication-free coupling principles and Gumbel sampling developed by Daliri et al., which is of independent interest. Replacing EAGLE-3's speculative sampler with Gumbo results in up to a +26.7\% increase in speedup at a temperature of 1.0 across four models and five datasets, without altering the draft model or requiring any other code changes.
Chat is not available.
Successful Page Load