DUDS: Dual-stage Data Selection for Efficient Reinforcement Learning with Verifiable Rewards
Abstract
Recent advances in large language models (LLMs) have leveraged reinforcement learning with verifiable rewards (RLVR) to enhance reasoning capabilities. However, RLVR typically relies on massive training data and extensive rollouts, posing substantial challenges to computational resources. Existing data selection approaches address this challenge in isolation: offline methods perform static, one-time dataset pruning that cannot adapt to the model’s evolving learning needs throughout training, while online methods conduct per-iteration filtering at the expense of significant additional computation. In this paper, we propose DUDS, a Dual-stage Data Selection framework for RLVR, which organically integrates the strengths of both paradigms to improve training efficiency while maintaining competitive performance. Specifically, the offline stage curates a candidate pool from the full dataset by jointly considering quality score and sample diversity via Determinantal Point Processes, providing an informative warm start for subsequent training. The online stage employs Bayesian posterior updates to estimate sample pass rates without introducing additional rollouts, and combines them with a freshness metric for data sampling, then applies asymmetric filtering to phase out mastered or intractable problems. Extensive experiments across four reasoning benchmarks demonstrate that DUDS consistently outperforms existing methods in both offline and online data-selection scenarios, achieving competitive performance with significantly less data and computation. Code is available at here.