On the Trainability of Masked Diffusion Language Models via Blockwise Locality
Abstract
Masked diffusion language models (MDMs) can train less reliably than autoregressive (AR) models on tasks with local left-to-right dependencies, yet can outperform AR models on tasks requiring global refinement and constraint satisfaction. We study this tradeoff through generation order. We propose Scatter, a synchronized order across blocks that denoises tokens at the same offset in every block in parallel while conditioning on earlier offsets. On structured diagnostics, Scatter preserves reverse planning on star-graph path-finding, improves over standard MDMs on in-context linear regression, and shows a tradeoff between locality and consistency on Sudoku. On LM1B, Scatter improves test perplexity from 28.16 to 26.35 under the BD3-LMs setup and lowers generation perplexity across sampling budgets. We further analyze the Jigsaw control, an entropy-guided order with sequential locality that is better on local binding and Sudoku, but fails on reverse planning and performs poorly on language modeling. Our results suggest that, beyond block size, generation order matters for diffusion LMs: locality should preserve global propagation rather than force premature block commitment.