VerifierContractBench: Testing Behavioral Preservation in Agent-Rewritten Verifiers
Lakshya Narula ⋅ Saurabh Tripathi ⋅ Lily Sharma ⋅ Nikhil Rao
Abstract
Coding agents can rewrite verifiers that reward or approve other agents. Such a revision changes an executable verification artifact under an operational contract: improve utility without withdrawing a correct rejection. Aggregate validation does not test that relation. We introduce VerifierContractBench, an empirical behavioral-preservation benchmark. Each episode starts from the same frozen multimodal verifier, exposes $102$ expert-labeled web-agent trajectories, permits six edit--evaluate--rollback rounds, and freezes the incumbent before evaluation on $492$ task-disjoint trajectories. Five hosted model--provider configurations, two evidence modes, three gates, and three repetitions yield $90$ episodes and $540$ rounds. The new unsafe-certification rate (NUCR) counts human-unsafe trajectories rejected by the starting verifier but newly certified after optimization. Every one of $78$ edited final verifiers has a protected counterexample; none improves protected balanced accuracy without one. In this finite panel, aggregate-safety and instance-preserving policies have mean NUCR $3.29$ and $3.19$ percentage points below utility, with $95%$ design-resampling intervals excluding zero, but neither eliminates regressions. All $383$ executable proposals edit the natural-language specification; only nine also edit formatting and none parsing, so realized search is primarily semantic. Baseline controls and two final-verifier rechecks leave a strict persistent regression in $77/78$ edited episodes. Immutable baselines, finite-set checks, whole-artifact rollback, protected transitions, and provenance make updates auditable, not formally proved. We will release the code and the data upon acceptance.
Chat is not available.
Successful Page Load