Test-time Scaling for Diffusion Language Models with Frequency-Aware Remasking
Abstract
Diffusion large language models (DLLMs) generate text through iterative denoising, where the choice of which tokens to commit and which to refine drives output quality. Existing certainty-aware strategies rely on instantaneous confidence or entropy to guide this choice, but these local signals can be misleading on hard reasoning problems. This raises a natural allocation question for DLLM test-time scaling: rather than applying one fixed decoder to all inputs, can we allocate not just more compute but a different decoding strategy to questions that remain unresolved? We propose a two-stage adaptive DLLM sampling method. The first stage uses a fast base sampler to solve easy questions; the second applies frequency-aware remasking to the remaining hard questions. Our frequency-aware strategy tracks how often token predictions change across denoising steps and keeps frequently-changing tokens available for further refinement, providing a trajectory-level signal that complements confidence- and entropy-based remasking. Experiments with LLaDA-8B-Instruct and Dream-7B-Instruct show consistent gains over a strong adaptive allocation baseline: up to 3.8% (LLaDA) and 7.4% (Dream) absolute accuracy improvements on MATH-500, and up to 11.18% (LLaDA) and 10.77% (Dream) absolute coverage improvements on HumanEval. On questions that remain unsolved by the base certainty-aware strategy even after 100 samples, frequency-aware remasking recovers solutions for 6.25% to 18.75% of examples.