Probes as Training Signals: Iterative Retraining Overcomes Evasion and Reverses Reward Hacking
Abstract
Feature-based supervision is proving increasingly useful for shaping how models learn to solve tasks. For example, white-box monitors like linear probes have been proposed as penalties in reinforcement learning to prevent undesirable behaviors such as reward hacking. However, optimizing against a probe can compromise it as both a training signal and a monitor: models can learn to evade the probe while maintaining the target behavior. In exploratory experiments, we show that iteratively retraining probes on on-policy data can reverse learned reward hacking and overcome probe evasion, reducing reward hacking from as high as 86\% to 3\% or below. We propose retraining adaptively, whenever the probe shows signs of being evaded, which suppresses reward hacking more reliably than retraining on a fixed schedule. Although the model can learn to evade individual probes, each retrained probe recovers detection ability, showing that optimizing against a probe does not remove the underlying signal. We demonstrate this in a code-generation setting where reward hacking is incentivized on a portion of inputs. Our results suggest that iterative retraining can make probe-based penalties an effective intervention against reward hacking and maintain probes' validity as monitors under optimization pressure.