From Autoregression to Block Diffusion: Adapting Language Models for Efficient Parallel Decoding
Abstract
Small Language Models (SLMs) are now capable enough for real workloads on consumer and embedded hardware, where decoding cost increasingly dominates. Autoregressive generation requires one forward pass per emitted token, making long agentic and multi-turn workloads particularly expensive. Diffusion language models tackle this issue by replacing token-by-token causal decoding with bidirectional parallel refinement, reducing the number of forward passes needed for long outputs. However, they still suffer from a quality gap relative to autoregressive models, especially at the small scale. In this paper, we propose a compact recipe for adapting a pre-trained autoregressive model into a block-causal diffusion model with an improved quality-efficiency trade-off and reduced generation cost. We apply our methodology to LFM2.5-350M as a case study and provide a fused kernel implementation of the resulting diffusion model for efficient serving.