The Reasoning Boundary Paradox: How Reinforcement Learning Constrains Language Models
Nguyen Phuc ⋅ Chinh D La ⋅ Duy M. H. Nguyen ⋅ Nitesh Chawla ⋅ Binh T. Nguyen ⋅ Khoa D Doan
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a key method for improving Large Language Models' reasoning capabilities, yet recent evidence suggests it may paradoxically shrink the reasoning boundary rather than expand it. This paper explains this shrinkage issue of RLVR by analyzing its learning dynamics and reveals two critical phenomena that account for this failure. First, we expose negative interference in RLVR, where learning to solve certain training problems actively reduces the likelihood of correct solutions for others, leading to the decline of Pass@$k$ performance, or the probability of generating a correct solution within $k$ attempts. Second, we uncover the winner-take-all phenomenon: RLVR disproportionately reinforces problems with high likelihood, correct solutions under the base model, while suppressing other initially low-likelihood ones. Through extensive theoretical and empirical analysis on multiple mathematical reasoning benchmarks, we show that this effect arises from the inherent on-policy sampling in standard RL objectives, causing the model to converge toward narrow solution strategies. These insights motivate us to design a \textit{simple yet effective} data curation algorithm that focuses RLVR learning on low-likelihood problems (the non-winners) and achieves notable improvement in Pass@$k$ performance, illustrating the potential of a broader class of data-curation mitigation strategies.
Chat is not available.
Successful Page Load