Sensitive but Not Specific: Log-Side Screening for Reward Hacking in Continually Updated Agents
Jules Roussel
Abstract
Deployed agents may be updated continually-through feedback-driven retraining, distillation, or re-tuning---and operators must decide whether each candidate can safely replace the trusted incumbent, often using only trajectories the incumbent has logged. Off-policy evaluation (OPE) is a natural offline verification tool, but importance-sampling estimates become unreliable precisely when a candidate departs strongly from logged behaviour. We ask whether this failure is itself informative: can log-side support diagnostics distinguish reward hacking from benign policy evolution? In a controlled Wordle agent testbed with exact action probabilities, we construct and gate-certify faithful, benign-drift, benign-improving, degraded, and reward-hacked policies-the hackers distilled from scripted proxy-farming teachers after hacking failed to emerge across seven escalating GRPO runs. Effective sample size (ESS) is sensitive but not specific: it collapses for every OPE-evaluated certified hacker, yet also falls below the reliability threshold for our strongest benign improver, while a degraded policy passes. A second diagnostic, \%floor, orders the certified categories without overlap: benign $\leq 0.05$, degraded $0.12$, hacked $0.15$-$0.16$. Yet under a benign control matched to a hacker on importance-weight variance, 0 of 6 diagnostics pass a pre-registered hacking-specificity criterion. Trusted-policy logs can therefore triage which agent updates deserve direct evaluation first, but cannot by themselves verify a reward-hacking verdict.
Chat is not available.
Successful Page Load