AnchorRank: Training-Free Anchor based Plan-and-Execute Decoding in DLMs
Abstract
Diffusion language models enable multi-token decoding and any-order generation, providing the potential for faster inference and improved generation quality. However, existing inference-time samplers primarily select tokens using marginal confidence or entropy, producing locally expanding, approximately sequential decoding trajectories. Recent Plan-and-Execute methods instead unmask informative anchor tokens before filling in the remaining details, but require model retraining or auxiliary selection modules. In this work, we propose AnchorRank, a training-free Plan-and-Execute decoding algorithm that identifies anchors using inference time signals. AnchorRank models the masked sequence as an entropy-weighted attention graph and uses finite-horizon PageRank to identify confident tokens that are globally important to uncertain parts of the response. AnchorRank achieves up to a 5.5 percentage points accuracy improvement over baselines at matched NFE budgets on math and coding tasks across two open source DLMs.