VerifierContractBench: Interpreting Meta-Agent Changes to Agent Evaluators
Abstract
Automated verifiers compress web-agent trajectories into rewards, benchmark scores, or deployment decisions. Coding meta-agents can rewrite them, but a final score hides how the evaluator's decision boundary moved. We introduce VerifierContractBench, a trace-rich benchmark of this process. From one frozen verifier artifact and 102 expert-labeled trajectories, each meta-agent receives six edit--evaluate--rollback rounds; the selected verifier is then audited on 492 task-disjoint trajectories. Five hosted model--provider configurations, two evidence modes, three gates, and three repetitions yield 90 episodes and 540 rounds. A trajectory is human-unsafe to certify unless jointly labeled successful and free of unintended side effects. The new unsafe-certification rate (NUCR) measures cases that were initially rejected but were newly certified after optimization. All 78 edited finals regress; 77/78 also regress on a side-effect-labeled case, and none improves protected balanced accuracy without one. Safety-gate policies lower NUCR from 8.81% under utility to 5.52% and 5.63%, while exposed-case identities vary across repetitions. Traces show that 274/383 executable proposals withdraw a correct rejection, and every executable proposal edits the specification. VerifierContractBench exposes proposal, promotion, rollback, artifact, and protected-transition behavior that final scores omit.