VerifierContractBench: Auditing Safety Regressions in Meta-Agent-Optimized Verifiers
Lakshya Narula ⋅ Saurabh Tripathi ⋅ Lily Sharma ⋅ Nikhil Rao
Abstract
Automated verifiers supply rewards, benchmark scores, and deployment approval for agents. Coding meta-agents can rewrite them, but aggregate accuracy can conceal a withdrawn protection: an expert-rejected trajectory may become certified. We introduce VerifierContractBench, which evaluates the complete verifier-improvement episode. From one frozen verifier artifact and 102 expert-labeled trajectories, each meta-agent receives six edit--evaluate--rollback rounds; the selected verifier is then audited on 492 task-disjoint trajectories. Five hosted model--provider configurations, two evidence modes, three gates, and three repetitions yield 90 episodes and 540 rounds. A trajectory is human-unsafe to certify unless labeled both successful and free of unintended side effects. Our primary new unsafe-certification rate (NUCR) measures cases initially rejected but newly certified after optimization. All 78 edited finals have an observed regression; 77/78 also regress on a side-effect-labeled case, and none improves protected balanced accuracy without one. Removing cases whose baseline decision changes across five fresh calls leaves all 78 positive; in 77/78, the regression recurs in the original audit and both rechecks. Aggregate-minus-utility and instance-minus-utility contrasts are $-3.29$ points ($[-5.67,-1.10]$) and $-3.19$ ($[-5.44,-1.08]$); the safety gates do not differ detectably. VerifierContractBench makes meta-agent optimization an auditable safety intervention rather than a scalar gain.
Chat is not available.
Successful Page Load