From Next-Token to Next-Block: A Principled Adaptation Path for Diffusion LLMs
Abstract
Diffusion language models (DLMs) can generate multiple tokens in parallel, but training large DLMs from scratch remains expensive. A practical alternative is to adapt off-the-shelf autoregressive (AR) checkpoints into diffusion models, reusing their linguistic and reasoning capabilities. Existing adaptation recipes either modify logits and grow attention masks toward full-sequence diffusion, or directly fine-tune ARweights under a block-diffusion objective, leaving two questions underexplored: what diffusion paradigm should AR-to-DLM adaptation target, and what transition path preserves AR knowledge most effectively? We argue that Block-Diffusion is a natural destination because AR decoding corresponds to block size one at the level of attention and generation order, while larger blocks introduce controlled intra-block bidirectionality and parallel generation. Based on this view, we propose a context-causal adaptation path that keeps committed context strictly causal, a one-pass parallel training formulation with auxiliary AR guidance, and a gradual block-size curriculum. Across several AR initializations and model scales, these components improve average adaptation performance over random mask annealing and direct fine-tuning. Scaling the recipe yields NBDIFF-7B, which supports 32K-token contexts and achieves the strongest average performance among the compared diffusion LLM baselines on general, math, and code benchmarks. Code and checkpoints will be released upon publication.