When Group Normalization Undermines Sparse Reward Auditing in RLVR
Abstract
Reinforcement learning with verifiable rewards (RLVR) can reward a verifier's false positives instead of correct final answers. We examine how trusted labels should enter policy updates when the learner can audit only one percent of its responses. Our arithmetic experiments use Qwen3-1.7B-Base, an injected false-positive rule, and matched audit budgets with uniform sampling. Inverse-probability residual correction with linear leave-one-out advantages avoids collapse and improves on starting accuracy in all three continuations. Dividing these advantages by the current group's reward standard deviation instead collapses both 120-update continuations. Full-oracle training improves accuracy with or without scaling in one long continuation, whereas direct replacement of sparse labels collapses. Linear residual correction and two sparse-oracle baselines all give unbiased fixed-batch advantages, yet their outcomes differ: residual correction finishes 12.1 and 19.9 percentage points above oracle-only in the long runs; the sparse-oracle baselines lead in the 40-update pilot. A coefficient bound identifies one channel through which normalization removes the inverse-probability weight of rare corrections. A control that chooses the scale before auditing preserves that weight and recovers part of the accuracy gap. These results characterize how current-group normalization interacts with sparse residual correction in a controlled arithmetic setting.