Sparse Reward Auditing for RLVR under Policy-Dependent Verification Error
Abstract
Policy optimization can amplify an imperfect verifier's errors by increasing the prevalence of responses it incorrectly rewards. We show how trusted feedback on just 1\% of generated responses can support effective reinforcement learning with verifiable rewards (RLVR) through two routes: preserving rare corrections and learning reusable reward predictions. In a controlled arithmetic testbed with an injected false-positive channel, linear inverse-probability correction reaches 66.4--77.3\% final accuracy, compared with 2.3--2.7\% when corrected rewards are standardized within each current group; with dense oracle feedback, both updates learn. A fixed-group bound explains a coefficient-level mechanism behind this contrast: bounded normalized coefficients suppress the inverse-probability amplitude needed to compensate for rare audits. Historical audits provide a second source of improvement. At the same 1\% label budget, policies trained using a simple text-based reward predictor reach 79.54\% and 80.71\% accuracy on a common 1,024-question evaluation, compared with 64.36\% and 75.05\% for linear correction. Together, these findings connect reward estimation, update construction, and feedback reuse, providing concrete guidance for converting scarce trusted labels into effective policy-learning signals.