VerifierContractBench: Auditable Verifier Evolution for Agentic Systems
Abstract
Agentic systems increasingly treat evaluators as mutable system components: coding meta-agents can rewrite the prompts and programs that approve other agents' work. Yet aggregate validation scores can conceal newly introduced unsafe approvals. We introduce VerifierContractBench, a systems-oriented benchmark for the evolution of auditable verifiers. Each episode starts from the same frozen multimodal verifier, exposes 102 human-labeled web-agent trajectories, permits six edit–evaluate–rollback rounds, and freezes the final incumbent. We compare utility-only, aggregate-safety, and instance-preserving acceptance gates across five coding meta-agent endpoints and text-only versus multimodal evidence, yielding 30 episodes. Final verifiers are evaluated once on 492 task-disjoint protected trajectories. All 26 genuinely edited verifiers introduce new unsafe certifications while also repairing old ones. Instance-preserving rollback lowers the mean new unsafe-certification rate from 8.19% to 5.19%, but does not eliminate protected regressions. The benchmark contributes a reproducible execution substrate with immutable baselines, bounded edit surfaces, provenance-rich traces, rollback decisions, and protected post-evolution scoring. It shows why an agentic operating layer needs auditability and fail-closed baseline handling in addition to aggregate quality gates.