Latent Refinement Decoding: Enhancing Diffusion Language Models by Refining Belief States
Abstract
Autoregressive (AR) models remain the dominant paradigm for natural language generation, but their strictly sequential decoding process leads to high inference latency. Recent diffusion-inspired language models, such as LLaDA and Dream, alleviate this issue through parallel generation, yet they still face two key limitations: information loss, since predictive distributions over non-finalised tokens are discarded at each step, and unstable commitment dynamics, where local token decisions are not sufficiently coordinated at the global level. We propose Latent Refinement Decoding (LRD), a two-stage decoding framework consisting of Latent Refinement and a Predictive Feedback Loop. In the first stage, LRD keeps masked positions as distributional mixtures of predicted tokens and the mask embedding, enabling the model to form more globally consistent beliefs before committing tokens. In the second stage, LRD progressively finalises confident tokens while preserving uncertain ones for further iterative refinement, with KL-divergence dynamics serving as a stable criterion for convergence and early stopping. Experiments show that LRD improves performance on both coding tasks, including HumanEval by 6.3 and MBPP by 2.6, and reasoning tasks, including GSM8K by 2.9 and MATH500 by 3.8, while achieving up to 10.6× overall speedup. Moreover, LRD is compatible with single-token decoding, multi-token threshold commitment, and system-level accelerators; when combined with Fast-dLLM, it reaches up to 18.4× speedup over vanilla decoding while further improving accuracy.