The Verifier Gap: A Model-Free Audit Predicts Reward Hacking Before Training
Abstract
Reinforcement learning from verifiable rewards is bounded by its verifier. Outside mathematics and code, rewards come from rubrics, learned judges, and hand-built checkers whose distance from the true objective is unknown, so when a policy's reward rises, no one can say how much of that rise is progress and how much is hacking. We construct a setting where the true reward is exactly computable at both the outcome level and the process level: executable authorization policies over an enumerable context space of 124,416 contexts, where an oracle enumerates every minimal sufficient reason behind every decision, which makes the causal fidelity of a model's cited reasons an exact quantity rather than an estimate. We compile policy-as-code into exact process rewards, degrade the verifier along six realistic axes (outcome-only checking, the literature's standard sufficiency proxy, extraction noise, coarsened taxonomies, stale policies, and an LLM judge), and train against each degraded verifier while the exact one scores every rollout. Four failure phenotypes separate cleanly, each with its own behavioral signature: hacked, starved, diluted, and harmless. Among them is the predicted signature of sufficiency rewards teaching padding, where the proxy reports 1.000 while true faithfulness falls from 0.731 to 0.667. The phenotypes reduce to an empirical gap law. The verifier's selection error, measured before training and without any policy model, weighted by the fraction of updates that flow, predicts post-training damage (R2 = 0.99 optimistic, 0.92 pessimistic; n = 9 cells, leave-one-out error 0.031), and verifiers that overestimate quality do 2.4x more damage per unit of error than verifiers that underestimate it (95% bootstrap CI [1.2, 7.2]). A pre-registered replication on an independent implementation reproduces the phenotypes and their ordering but not the constants, and shows why: the verifier's error is a property of the verifier alone, while the damage it causes also depends on how much headroom the policy has left. We release the testbed, 14 real public policies (Cedar, OPA/Gatekeeper) ported to executable oracles, and the training instrumentation.