The RLVR Policy Shift Is Sparse and Predictable from the Base Model
Jose Luis Cantarero ⋅ Jezabel R Garcia ⋅ Marcelo Marín ⋅ Antonio Tiene ⋅ Román Orús
Abstract
After RLVR, the disagreement between a trained model and its base is concentrated at only a small fraction of the positions they generate. We can identify these positions in advance using a single scalar: the base model's next-token entropy. We set the threshold at a quantile fixed before examining the trained checkpoint and measure how much of the policy change occurs above it. Across ten panels on two Qwen2.5 zero-RL pairs, the per-position Jensen--Shannon divergence between the two policies is 18.4--57.1$\times$ higher inside this fork set than outside. Much of that factor follows by construction because uncertain positions allow larger divergences. We therefore adjust it by the entropy ratio between the two zones. After this adjustment, an excess of 1.06--1.70$\times$ remains on all ten panels. Rate-matched random placement shows no such excess, indicating that the concentration depends on placement rather than rate. We introduce Fork-Patched Decoding (FPD), a decoder that splices base logits into the trained policy where the gate fires. The same decoder supports firing at the gated positions, at their complement, at rate-matched random positions, or without restrictions. Model behaviour follows the same partition. On a pooled AIME panel, the two parent models differ by 6.91 points in $\text{pass@1}$. Mixing the base back in outside the fork set changes $\text{pass@1}$ by at most $0.07$ points, only one percent of that difference. In contrast, applying the same operation at the forks costs about one point and moves the number of distinct answers 32.8\% of the way back toward the base. This recovery measures distinct-answer diversity, with accuracy evaluated separately. It is replicated on disjoint seeds, while random placement recovers only 4.7\%. At 7B, an entropy-matched control under the same gate reproduces about 60\% of this recovery; the remainder requires the next-token predictions of the base model rather than the entropy added by mixing alone. All measurements use one model family, one RLVR recipe, mathematics benchmarks and one sampling protocol. Our results provide a partition that is available before training and inexpensive to compute from a single model. Interventions inside and outside this partition have different effects, which makes the partition an instrument for locating where RLVR rewrites a policy.
Chat is not available.
Successful Page Load