The Path Not Taken: RLVR Learns Off the Principals
Abstract
RLVR reliably improves reasoning, yet it appears to update only a tiny fraction of weights. We resolve this paradox by revealing a persistent, model-conditioned optimization bias: independent runs concentrate updates in similar parameter regions, invariant to dataset or RL recipe, while finite-precision storage (e.g., bf16) obscures widespread micro-updates as a visual sparsity artifact. To characterize this unique bias, we show that RLVR preferentially learns along off-principal directions via a cascaded Three-Gate mechanism: updates are first bounded by a empirical KL constraint (Gate I), then steered by the model’s anisotropic geometry(Gate II) toward spectrum-preserving off-principal subspaces, and finally filtered by finite precision (Gate III), effectively masking minor updates. Empirically, we validate the off-principal dynamics: RLVR exhibits minimal spectral drift, reduced principal-subspace rotation, and strong off-principal alignment that set RLVR strikingly apart from SFT. Together, our results provide the first parameter-space account of RLVR and uncover consistent regularities in weight evolution, advancing a more white-box understanding of RLVR. Moreover, we show that RLVR follows an optimization regime distinct from SFT, showing directly that transferring SFT-era PEFT can be flawed and motivating geometry-aware, RLVR-native methods.