When the Metric Lies: A Full-Stack Audit of a Multi-Agent Evaluation Protocol
Jan Safrata
Abstract
When an evaluation is doubted, we usually blame the data: contamination, leakage, benchmark reuse. Interactive multi-agent evaluation has a deeper failure mode: the test distribution is endogenous—generated by the very training process under evaluation. We audit standard multi-agent evaluation practice end to end in CalHunger, a four-player market-survival game built on inescapable interdependence: provably, no agent survives without repeatedly trading with its competitors, so the co-player population can never be factored out of what is measured. Every layer of the standard protocol fails. The instruments are invalid: self-play survival—the number every training loop watches—rises monotonically to 90% while survival against a fixed population peaks mid-training and then halves, and terminal cash rises $15→$174 in those same games, so ranking checkpoints by cash selects the most degraded ones. The verdicts are properties of the protocol, not the policies: on a 10×10 peer matrix at 1,000 paired games per cell, one policy ranks anywhere from 1st to 9th depending on the evaluating population, and 15 of 28 pairwise verdicts invert significantly under an alternative defensible metric. The sampling is inadequate: ten disjoint-seed replications of the standard 100-game protocol crown three different champions, and certifying the full ranking would cost 190× that budget. The proxies do not proxy: an opponent ladder engineered for this exact game predicts peer competence at ρ=0.00. The decisions are costly: every checkpoint-selection rule we test—final, self-play, cash, held-out portfolio—forfeits 18–27 survival points against an oracle. A stress test with an adversary generated from the champion's own weights localizes the failure: the self-play certificate breaks under distribution shift, not adversarial attack. None of this is an RL artifact: four production LLM agents, audited as black boxes under the identical protocol, reproduce the population-relative inversions. We condense the audit into a portfolio protocol whose components each catch a failure the others mistake for a healthy result, and argue that a trustworthy evaluation reports every component's verdict—each indexed by the budget that decides it—rather than collapsing them into a single number.
Chat is not available.
Successful Page Load