Boosting Off-Policy RLVR with Data-Centric Replay
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has become an important approach for improving the reasoning capabilities of large language models (LLMs). On-policy training remains the dominant RLVR paradigm, but its reliance on fresh rollouts leads to poor data efficiency. To address this bottleneck, off-policy paradigm has attracted increasing attention by reusing historical trajectories through replay buffers. Despite its efficiency advantages, off-policy training often struggles to match the performance of on-policy training. Our analysis reveals that this limitation primarily stems from data staleness introduced by replay buffers: to maintain training stability, mainstream algorithms discard a large fraction of high-staleness optimization signals, leading to substantial performance degradation. To address this issue, we propose STAR (Staleness-Aware Replay), a data-centric replay framework decoupled from specific optimization algorithms. STAR improves off-policy RLVR by constructing low-staleness and high-quality training data. Experiments on a range of mathematical reasoning and code generation tasks show that STAR consistently improves off-policy RLVR with up to 19% relative performance gains, and enables it to surpass the on-policy baseline using only around 13% of the rollout data while achieving 5x-6x wall-clock speedup.