Measuring Parallel Decoding in Diffusion Language Models: Decoder Settings, Generated Text, and Commit Order
Parsa Rahimi ⋅ Azim Dehghani Amirabad
Abstract
Masked diffusion language models can commit several tokens in one network call, but tokens per forward ($\tpf$) is not a checkpoint constant: it depends on the decoder, request, generated text, stopping rule, and accounting. We ask whether a reported reasoning-minus-knowledge gap recurs across checkpoints and what remains when the question or target text is held fixed. On the same 288 MMLU-Pro item keys and outer decoding settings, the chain-of-thought contrast is positive for all five checkpoints; four rows are post hoc and show descriptive recurrence, not model-population replication. Requesting a direct answer changes the gap non-uniformly, while a separate vocabulary-constrained Nemotron intervention also changes the decoder itself. Forced-target replay on Nemotron-8B gives a non-native donor-group contrast of $+0.0005$ $\tpf$ (crossed-bootstrap 95\% interval $[-0.0070,+0.0091]$), and decoder sweeps move $\tpf$ more than the benchmark-group gap. Across the aligned step traces, reasoning-labelled outputs have more simultaneous commits in all five rows. With observed-pair order measured by mean-block Somers' $D$, the contrast is distinguishable from zero only for the LLaDA-8B reference and LLaDA-1.5. Parallelism and order therefore describe a specified checkpoint--decoder--request--output run, not model weights alone.
Chat is not available.
Successful Page Load