VerifierContractBench: Auditing the Safety Validity of Optimized Agent Evaluation
Abstract
Evaluation protocols increasingly include mutable LLM-based verifiers whose outputs supply rewards, leaderboard scores, or deployment approval. Optimizing such a verifier can raise aggregate accuracy while silently moving errors onto cases that were previously handled safely. We introduce VerifierContractBench, a benchmark that treats evaluator optimization itself as the object of measurement. Each episode begins from the same frozen multimodal verifier, exposes 102 expert-labeled web-agent trajectories, permits six edit–evaluate–rollback rounds, and freezes the final incumbent. We compare utility-only, aggregate-safety, and instance-preserving promotion rules across five coding meta-agent endpoints and text-only versus multimodal evidence, yielding 30 episodes. Final verifiers are evaluated once on 492 task-disjoint protected trajectories. Our primary paired metric asks: among human-unsafe trajectories that the starting verifier correctly rejects, how many are newly certified after optimization? All 26 genuinely edited final verifiers create at least one such regression while also repairing old errors. Instance preservation reduces the mean regression rate from 8.19% to 5.19%, but no edited episode combines higher protected balanced accuracy with zero new unsafe certification. The benchmark separates aggregate improvement, error swapping, rollback dependence, and protected validity in a reproducible evaluation protocol.