Reinforcement Learning for Diffusion LLMs with Entropy-Guided Step Selection and Stepwise Advantages
Abstract
Reinforcement learning (RL) has been effective for post-training autoregressive (AR) language models, but extending these methods to diffusion language models (DLMs) is challenging due to intractable sequence-level likelihoods. Existing approaches therefore rely on surrogate likelihoods or heuristic approximations, which can introduce bias and obscure the sequential structure of denoising. We formulate diffusion-based sequence generation as a finite-horizon Markov decision process over the denoising trajectory and derive the policy gradient that decomposes over denoising steps in terms of stepwise advantages, without requiring explicit evaluation of the sequence likelihood. Grounded in this theorem, we develop tractable approximations for large-scale training: (i) denoising steps are selected for policy updates via an entropy-guided approximation bound, and (ii) stepwise advantages are estimated using a one-step denoising completion naturally provided by the diffusion model, avoiding costly multi-step rollouts or auxiliary value networks. Experiments demonstrate state-of-the-art results on nearly all benchmarks spanning coding, logical reasoning, and mathematical reasoning, outperforming existing RL post-training approaches for DLMs.