When the Judge Drifts: Separating Safety Behavior from Its Measurement Along a GRPO Trajectory
Abstract
Verifiable-reward post-training (RLVR) optimizes models against checkable rewards. Safety evaluations then test whether safety-relevant behavior changed. We separate two explanations for a moving safety score, changed behavior and changed measurement. We construct a fully counterbalanced safety instrument with 24 authored scenarios, four wordings, and all six orderings of three semantic response classes, then evaluate it on one open GRPO lineage (Tülu 3.1 8B, DPO base, 12 pinned checkpoints; 6,912 structured responses) under frozen claim gates with a prespecified practical-equivalence margin of ±0.10. The marginalized behavioral score never left the band (largest increase +0.057). Late in training, permutation-invariance fell by 0.146 and answer-order range rose by 0.188 between steps 1,920 and 2,240. A triggered dense follow-up localized this change to steps 2,080→2,120 (−0.135 and +0.146), with only +0.028 behavioral change. An exact-match capability panel improved from 36.7% to 63.3%, and an option-free safety panel was too imprecise to replicate a safety change. The final label is measurement drift, not safety drift.