Efficient Diffusion Language Model Serving beyond Autoregressive Assumptions
Guan-Ming Chiu ⋅ Jeng-Yue Liu
Abstract
Diffusion language models (dLLMs) are now served by an inference stack built for autoregressive models, and this inheritance is unsound: the assumptions that stack makes about what memory costs and what latency depends on both fail under diffusion. We argue the failures are not incidental but structural, and that the diffusion paradigm itself supplies the repairs at little cost. Building on one abstraction, a separation of holding capacity from receiving compute, we realize both repairs in a production engine on a state-of-the-art dLLM, moving the concurrency wall from 32 past 128 and raising tight-SLO attainment up to $2.5\times$. Experiments across three serving systems show which of these findings are properties of the paradigm and which belong to the engine.
Chat is not available.
Successful Page Load