Reward Budgeting Reduces Premature Convergence in Reinforcement Learning for LLM Reasoning
Mengni Jia ⋅ Mengyu Zhou ⋅ xiaoxi jiang ⋅ Guanjun Jiang
Abstract
RL often plateaus early: policies sharpen quickly, and further training yields little gain despite remaining model- and data-capacity. Entropy is recently used to combat this plateau, but its promotion alone is not a sufficient intervention target in the settings we study. By analyzing RL's dynamics through the optimal policy, we instead identify a sharpness factor $\kappa$ as a more effective alternative. We then introduce a plug-in reward-budget mechanism, which allocates a fixed total reward budget to each prompt and distributes it among correct rollouts. Theoretically, this induces a more regulated $\kappa$-trajectory that deviates from many RL methods, and its updates are equivalent to optimizing a concave objective, approximately $\mathbb{E}(\log R)$. Empirically, our method achieves significant gains on Qwen3-8B and Llama3.2-3B-Instruct (up to +12.50 on AMC23 and +10.00 on AIME24 \& AIME25). Together, these results suggest that regulating algorithm-induced sharpening is more effective for mitigating premature convergence than directly preserving entropy.
Chat is not available.
Successful Page Load