Exploiting Per-Sequence Token Sparsity in Batched Block-Diffusion LLM Inference
Mike Mao ⋅ Fanjiang Ye ⋅ Edward Sun ⋅ Xinze Feng ⋅ Yuke Wang
Abstract
Diffusion large language models (dLLMs) are emerging as an alternative to autoregressive models, capturing bidirectional context and generating in parallel. Inference remains expensive because every position in the active block is recomputed at every denoising step. We find that most positions converge well before a block finishes decoding, and propose a training-free staleness gate that freezes a position once its hidden-state drift falls below a threshold, reusing its cached state until a periodic recheck. Because each sequence freezes different positions, a dense batched forward skips a position only when it is stale in every sequence. We call this gap the batch-union tax, and show that a ragged variable-length execution of the gate recovers it. On Fast-dLLM v2 (7B) on GSM8K, the ragged kernel is near-lossless and reaches 1.47$\times$ over dense at batch 512, with accuracy preserved to roughly 1.5$\times$ across various staleness thresholds.
Chat is not available.
Successful Page Load