Why Speculative Decoding Works Better Than Predicted on Sparse MoE Models
Abstract
Speculative decoding (SD) accelerates autoregressive inference by leveraging the memory bound nature of decoding, verifying multiple draft tokens in a single forward pass. For mixture-of-experts (MoE) models, prior work predicts a concave speedup-batch size relationship, with gains peaking at intermediate batch sizes due to expert saturation. We evaluate this across 10 sparse MoE models (30B-1T parameters, 8-512 experts) and find recent fine-grained models break this trend, showing higher speedup at low batch size. We find verification cost is governed not only by activated expert weight, but also by temporal overlap across draft tokens and GPU execution dynamics. At low batch size, fine-grained MoE layers run faster than expected. Small per-expert weights and small per-token activated parameter footprint leave bandwidth underutilized and yield higher L2 cache hit rates. Across draft tokens, routing overlap further reduces the unique experts activated during verification relative to independence assumptions. We model this effect using a maximum-entropy Iterative Proportional Fitting (IPF) framework that predicts unique expert counts from routing-overlap statistics with less than 4\% error, and reproduces observed speedup curves from synthetic routing.