Distribution-Adaptive Policy Optimization
Abstract
Reinforcement Learning with Verifiable Rewards empowers Large Language Models to enhance their reasoning capabilities. Most algorithms rely on on-policy training, which is stable but inefficient due to its strict dependency on fresh samples. To overcome this inefficiency, off-policy RL decouples experience generation from policy optimization via Importance Sampling (IS). However, this introduces a policy discrepancy between the behavior policy and the target policy. We observe that mainstream RLVR algorithms struggle to maintain a balance between performance and stability in off-policy scenarios, suffering from either severe performance degradation or training collapse. We identify the root cause as the failure of static trust regions to maintain the bias--variance trade-off associated with IS ratios in off-policy settings. To address this, we propose DIstribution-adaptive Policy Optimization (DIPO), a novel hyperparameter-free RLVR algorithm employing an adaptive trust region. DIPO leverages the intrinsic statistics of the IS ratio distribution to achieve a robust bias--variance trade-off. Experimental results demonstrate that DIPO enables stable training in extreme off-policy scenarios, achieving performance comparable to or even surpassing the on-policy baseline, and outperforming other off-policy baselines by 13.9%.