DUET: Optimize Token-Budget Allocation for Reinforcement Learning with Verifiable Rewards
Haoyu Hu ⋅ Xuandong Zhao ⋅ Nori Jacoby ⋅ Xuhai "Orson" Xu
Abstract
Reinforcement learning with verifiable rewards (RLVR) generates thousands of tokens per training step, with rollout generation dominating the computational cost. The overall token budget can be controlled along two main dimensions: (i) deciding which prompts to allocate rollouts to, and (ii) deciding how long each rollout should be. Prior work has generally controlled only one of these dimensions at a time. We show that jointly tuning both decisions under a shared compute budget improves both reasoning quality and wall-clock training time. We instantiate this view as \textbf{DU}al-controlled tok\textbf{E}n alloca\textbf{T}ion (DUET), a computationally efficient layer over GRPO, in which we have a lightweight pre-rollout surrogate of prompt informativeness to set how many rollouts each prompt receives and a marker-gated abort rule with importance reweighting set when to stop them. On Qwen3-1.7B trained on MATH, DUET outperforms full-budget GRPO and the other three budget-aware baseline methods. DUET's advantage was further generalized to other benchmarks across math and coding, and was on par with the best baseline on the scientific Q\&A domain, while also achieving a $1.62\times$ wall-clock speedup. More notably, using only 50\% of the token budget, DUET still outperforms all baseline methods, and performance often increases rather than reduces with the more limited budget, achieving an even higher $2.51\times$ speedup. We verified the high performance of DUET on other backbone LLMs, including Qwen-4B and Llama-3.2-3B-Instruct. Notably, the gap between DUET and the strongest baseline \emph{widens} as the budget tightens, contrary to the usual pattern in which efficient methods trade off quality as compute decreases. More broadly, these results suggest that DUET budget-aware control strategies are valuable not only for accelerating training, but also for improving the quality of the learning signal.
Chat is not available.
Successful Page Load