When Does Decoding Order Cost Coverage in Diffusion Large Language Models?
Sean Wu ⋅ Jaewon Lee ⋅ Hongjian Zhou ⋅ Fenglin Liu ⋅ Zachary Shinnick ⋅ David Clifton ⋅ Junchi Yu
Abstract
Diffusion large language models (dLLMs) offer flexible token-generation orders, but it remains unclear how decoding freedom affects downstream metrics like Pass@$k$. We compare eight selection policies across three open dLLMs (LLaDA-8B, Dream-7B, Nemotron-3B) over 450,000 generations. At one token per step, left-to-right decoding outperforms confidence-based ordering by 3.0--7.9 Pass@32 points on four of eight model--benchmark pairs. At two tokens per step, this advantage disappears, and we trace this effect to adjacent co-commits rather than to decoding order itself. We use commit probes and find that confidence-based ordering postpones tokens like *Or* and *First* until surrounding text is fixed, by which point their influence has collapsed. Committing these tokens early recovers $+7.3 \pm 3.3$ pp Pass@32 on Dream and $+3.0 \pm 2.9$ pp on LLaDA without enforcing left-to-right order. Raising the sampling temperature yields a comparable $+5.7 \pm 2.6$ pp gain, yet the reordering benefit survives only at low temperature. Across all policies, regimes, and temperatures, we find that sample diversity alone predicts Pass@32 with held-out $R^2 = 0.94$, so diversity drives coverage rather than ordering. AR order is efficient and delivers the same coverage at half the single-sample accuracy cost. Ultimately, the coverage benefit that motivates left-to-right constraints holds only at one token per step and disappears in the parallel regime where dLLMs are fastest.
Chat is not available.
Successful Page Load