Temporal Expert Caching for Accelerating Inference in MoE Diffusion Language Models
Vikhyath Kothamasu ⋅ Sandeep Kumar ⋅ Deepika Palagani
Abstract
Diffusion language models (DLMs) offer a parallel-decoding alternative to autoregressive generation, and sparse mixture-of-experts (MoE) variants such as the LLaDA series scale quality while keeping per-token compute low. On commodity GPUs, however, masked diffusion repeats a full-sequence forward pass across dozens of denoising steps and each step independently reloads expert weights from host memory, making generation prohibitively slow. We show that this cost is largely avoidable since adjacent denoising steps exhibit strong *temporal locality*--per-token hidden-state cosine similarity of 0.97 and layer-level routed-expert Jaccard overlap of 0.94 - 0.97 -- which we exploit through three coordinated mechanisms: (i) per-expert *output caching* with cosine-similarity gating that skips expert work for stable tokens; (ii) *temporal weight prefetching* that uses previous-step routing to predict the next step's expert set; and (iii) a CPU SwiGLU *fallback* that absorbs misses by sending small activations to the host rather than pulling large weights back. Evaluated across two LLaDA-mini models, two GPU architectures (L40S and A6000), and seven benchmarks spanning code, mathematical, and general reasoning, our combined caching delivers a median **9$\times$ end-to-end speedup** over an HF~Accelerate baseline at matched memory budgets, with task accuracy preserved within measured run-to-run variance.
Chat is not available.
Successful Page Load