Nudging Beyond the Comfort Zone: Efficient Strategy-Guided Exploration for RLVR
Chanuk Lee ⋅ Sangwoo Park ⋅ Minki Kang ⋅ Sung Ju Hwang
Abstract
Reinforcement learning with verifiable rewards (RLVR) is a scalable paradigm for improving the mathematical reasoning of large language models, but it is fundamentally limited by exploration: the policy can only improve on trajectories it has already sampled. Sampling more rollouts alleviates this at prohibitive compute cost, while objective-level modifications offer little control over what is explored. We propose NudgeRL, a framework for structured, diversity-driven exploration in RLVR whose core component, Strategy Nudging, conditions each rollout on a lightweight strategy-level context, without requiring the context generator to solve the problem itself. To learn from such exploration, we decompose the advantage into inter- and intra-context terms and add a policy correction term that transfers discovered behaviors back to the base policy. Across five mathematical reasoning benchmarks, NudgeRL with 8 rollouts matches the strongest GRPO baseline using 32 rollouts with roughly $5\times$ fewer total tokens and $2.9\times$ less training compute.Our code is available at https://github.com/tally0818/NudgeRL.
Chat is not available.
Successful Page Load