Mitigating Reward Hacking in Mathematical Reasoning Without Ground-Truth Labels
Abstract
Reinforcement learning (RL) methods have demonstrated remarkable effectiveness in improving the mathematical reasoning capabilities of large language models (LLMs). However, existing approaches often rely on substantial amounts of labeled data or are restricted to domains where generated outputs can be automatically verified. Reward models (RMs) offer an alternative source of supervision by assigning rewards during training, but they are susceptible to reward hacking, where the policy exploits spurious correlations in the RM without improving its reasoning capabilities. To address this limitation, we introduce a reinforcement learning method that uses RM supervision for mathematical reasoning without requiring task-specific ground-truth labels. Our approach is based on a causal framework that removes the influence of spurious attributes on the reward by estimating the contribution of causal attributes, thereby limiting the policy's ability to hack the RM. Experiments on Countdown show that our method mitigates reward hacking while recovering 79% of supervised RLVR performance without requiring ground-truth labels. Our findings suggest that RM-guided reinforcement learning may provide a scalable path toward improving language models in domains where conventional sources of supervision are limited or unavailable.