Metrics That Write Themselves: Evolving an Evaluator from Its Own Blind Spots
Xing Zhang ⋅ Yanwei CUI ⋅ Guanghui Wang ⋅ Zhihao Lin ⋅ Peiyang He
Abstract
Agents improve quickly against a reliable automatic metric and stall without one, and the applications that need them most, report generation among them, are the ones nobody knows how to score. Can the metric write itself? Saying what makes an answer good is hard; pointing at something wrong with one is easier, so the metric we evolve is a pool of small Python \emph{operators} that each flag a candidate for one named defect, or abstain, and vote. Asking a model for operators directly does not work: 183 candidates realise only 96 distinct behaviours, from one narrow region of an enormous space. EvalCEGAR instead borrows counterexample-guided abstraction refinement from program verification. It reads the pool as an abstraction and searches for a \emph{collision}, two answers the operators score identically, one correct and one not. That pair, not a prompt, is the authoring request, and when a collision defeats every attempt the loop widens what an operator may read rather than resampling. On MBPP+ and HumanEval+, whose hidden unit tests give exact ground truth, six of eight runs admit an operator and all six help on 428 unseen tasks, 4 of them significantly, at a median $+0.0029$ and a best $+0.0065$: 55 lines of Python closing $15.4\%$ of the gap between flagging nothing and a perfect filter, at a quarter of our best hand-written operator's flags. On the benchmark the loop never saw, the best of them matches that operator's effect exactly at a third of its flags. Our 15 hand-written operators applied together as one filter \emph{lose} accuracy. An LLM judge on the same information ties that delta on a nearly disjoint set of candidates, and charges a model call per candidate forever where the operator charges none.
Chat is not available.
Successful Page Load