Can an Injected-Failure Benchmark Pick Your Monitor? A Four-Part Validity Audit for Enterprise Agent Deployment
Abstract
An enterprise that runs agents in production needs to detect when one fails, and it usually picks the monitor that scores best on a benchmark of labelled failures. Those labels are cheap to obtain by injection: change a step of a successful trajectory and the benchmark knows which step it changed. A correct label, however, does not show that the injected failure stands in for the failures the monitor will meet in service. We treat the injector as part of the evaluation and ask four questions. Can a detector read data written only by the injector? When two methods create the same error, can it tell which method was used? On the same task, can it tell injected and natural failures apart? Would injected and natural data lead us to pick the same monitor? On 24,000 agent traces the four tests give different answers. Removing injector-only data erases some apparent detection ability. One injection method stays identifiable after we control measured error severity. Matching on the query removes an apparent realism trend, and the monitor chosen from injected data changes with the detector set. The gap matters at deployment: a monitor selected on injected traces can lose measurable accuracy under the shift to natural failures. The audit gives enterprise benchmark builders a practical way to state which claims their injected failures support before a monitor goes into production.