VerifierContractBench: Auditing Safety Regressions in Evaluators for Enterprise Web Agents
Abstract
Enterprise web agents increasingly depend on automated verifiers to assign rewards, admit training data, and approve deployment outcomes. Coding meta-agents can now rewrite those verifiers, but aggregate accuracy can conceal a newly approved unsafe trajectory. We introduce VerifierContractBench, a benchmark that measures the complete verifier-improvement episode. Starting from one frozen multimodal verifier, a meta-agent receives 102 expert-labeled trajectories and six edit–evaluate–rollback rounds; the selected verifier is then frozen and tested on 492 task-disjoint trajectories. We compare utility-only, aggregate-safety, and instance-preserving promotion gates across five coding meta-agent endpoints and text versus multimodal evidence, producing 30 episodes and 180 rounds. All 26 genuinely edited final verifiers newly certify at least one human-unsafe protected trajectory while also repairing old errors. Instance preservation reduces the mean New Unsafe-Certification Rate from 8.19% to 5.19%, but cannot eliminate protected regressions. Exposure is structured: the WorkArena enterprise subset has a normalized exposure of 16.67%, versus 1.67% for AssistantBench. VerifierContractBench turns evaluator improvement into an auditable enterprise-agent safety process based on paired decision transitions rather than a single leaderboard score.