Verifying Agents in Rubric-Graded Environments
Markus Dücker ⋅ Vaibhav Kumar ⋅ Yi Liu ⋅ Ronak Chaudhary ⋅ Andreas Plesner ⋅ Francisco Guzmán ⋅ Anish Athalye
Abstract
As AI agents take on increasingly open-ended and complex tasks, verifying their outputs has proven to be difficult. Rubric-graded agent environments—where a verifier judges multiple criteria against the deliverables and environment state changes produced by the agent—have emerged as a popular paradigm. We conduct the first systematic study of verifiers for rubric-graded agent environments. First, we create BankerVerifierBench (BVB), a meta-evaluation dataset of 3,204 human- judged rubric criteria across 21 investment-banking tasks. Next, we develop a methodology that derives from any rubric corpus the capabilities a verifier should possess, yielding a nine-capability taxonomy that we distill into three general design principles—reactive verification, environment alignment, and domain guidance— which we implement in Gandalf, an open-source verifier. Finally, we evaluate Gandalf against existing verifiers on BVB. We find that Gandalf outperforms all baselines on seven of nine capabilities and is Pareto optimal: its cheapest configuration (F1 0.633, \\$42) exceeds the most expensive baseline (F1 0.538, \\$414). Ablations show that environment alignment and domain guidance primarily reduce cost (up to 4× fewer tokens) rather than error, while architecture choice dominates model choice. Gandalf generalizes to an OpenClaw benchmark, a structurally different personal-productivity environment, achieving F1 score 0.951 and outperforming the next-best verifier by 7.5 F1 points.
Chat is not available.
Successful Page Load