Low-Confidence Remasking Traps Flexibility: Realizing Arbitrary-Order Potential for Diverse Rollouts in Diffusion LLMs
Moongyu Jeon ⋅ Dongjae Jeon ⋅ Bumjun Kim ⋅ Mingyu Kim ⋅ Albert No
Abstract
Masked diffusion language models support arbitrary-order generation, suggesting a natural fit for diverse rollouts. However, recent work argues that this order flexibility reduces diversity by decoding high-uncertainty positions later. We trace much of this loss to *low-confidence remasking* (LCR), which samples a token at every masked position but commits only the highest-probability proposal. We show that this cross-position filtering can exponentially suppress lower-probability tokens as more positions compete, and observe the same suppression in LLaDA. In contrast, *top-probability position selection* (TPP), often conflated with LCR under *confidence-based decoding*, selects a position before sampling and directly follows its tempered token distribution. Replacing LCR with TPP restores diversity and yields Pass@$k$ comparable to left-to-right decoding, indicating that the reported diversity loss stems largely from LCR rather than confidence-prioritized ordering itself. Finally, we introduce **Entropy-Guided Initialization** (EGI), which samples the first token at the highest-entropy position and then follows TPP. EGI further improves rollout diversity and solution coverage beyond left-to-right decoding, with gains extending to downstream policy optimization.
Chat is not available.
Successful Page Load