Own Exposure, Advantage Reweighting, and Behavioral Improvement in RLVR
Abstract
Which training examples explain behavioral improvement during reinforcement learning with verifiable rewards (RLVR)? We examine two readily available training records: a question's own exposure and the advantage weights assigned to its responses. Using matched GRPO and practical MaxRL runs of a 0.6B model on GSM8K, we track 256 training questions across 20 policy snapshots. In the low nonzero-reward stratum, substantial correctness gains precede own exposure. Persistent reweighting toward the two lowest-reward strata is not accompanied by a persistent same-stratum absolute correctness advantage. In the GRPO run, questions in the low nonzero-reward stratum gain 10.0–12.8 percentage points before supplying their own RL training responses; post-outcome question-resampling ranges remain positive. The MaxRL intervention increases cumulative absolute advantage mass in the two lowest reward bins at every snapshot, reaching 1.72–2.21 times GRPO at the endpoint, yet their absolute correctness advantages change sign across training. Together, these observations caution against treating a question's own exposure or scalar advantage mass as a direct measure of improvement on that question. The exposed-versus-unexposed contrast remains uncertain under resampling, and a whole-panel-centered comparison gives mixed evidence about relative behavioral reallocation. The study uses one matched training seed and measures exposure and scalar weights, without identifying causal source examples or the mechanism of cross-question transfer.