What Happens to a Monitor’s Accuracy When You Train Against It
Abstract
A linear probe on a language model's hidden states predicts whether the model's next answer will be correct, and it does this well. On held-out data it ranks a correct answer above an incorrect one 98% of the time. We make a probe of this kind the reward for reinforcement learning, and true accuracy falls by 31 percentage points. The probe's own accuracy does not show this. Scored at every checkpoint on that checkpoint's fresh answers, it is unchanged for about 40 training steps after the gaming begins. An LLM judge scoring the same run shows the same thing. The judge is verifiably a fixed function, so what changed was the population of answers it was handed, not the judge. Accuracy statistics belong to a classifier and a population together. Deploying the artifact is what moves the population, so the number measured before deployment no longer describes it. One statistic does register the change, about 40 steps earlier and without ground truth: the share of answers scoring above a threshold fixed at the start. It has not been calibrated against benign training. Held-out accuracy also does not say which monitors are safe to train against. A monitor built from 39 features of the answer text, with no access to activations, scores higher than the probe as a monitor. As a reward it drives accuracy to exactly zero while 99.9% of answers remain well-formed. Nothing in the objective rewards evading the monitor. The policy is doing what it was asked.