Fork-dLLM: Avoiding the Flexibility Trap in Diffusion Language Models
Abstract
Masked diffusion language models (dLLMs) have shown strong potential for faster inference through parallel token generation when combined with confidence-based samplers. However, recent work has shown that such methods can defer unmasking high-entropy fork positions at which multiple plausible continuations exist. This results in reduced generation diversity, as shown by worse pass@k scaling, and limits gains obtainable from post-training and test-time compute. To avoid this flexibility trap, prior work therefore advocated for autoregressive (AR) sampling, which directly confronts uncertain positions. In this work, we show that discarding confidence-based sampling is unnecessary and, once inference cost is taken into account, wasteful. Instead, we propose a simple, hybrid strategy that uses AR-style sampling only when fork positions are encountered. Across our experiments, the resulting Fork-dLLM sampler maintains the pass@k performance of AR sampling as well as the efficiency gains of confidence-based sampling, thereby effectively resolving the flexibility trap in dLLMs.