The Benchmark Was the Bug: Seven Instrument Defects Behind Apparent Failures of Data-Analysis Agents
Abstract
Agentic systems are increasingly asked to run the statistical analysis that biological findings rest on, and benchmarks report how often they get the answer right. We built one such benchmark, in which every task outside a trap-free control family carries both a correct value and a trap, the exact number a specific named analytical mistake produces on that data, so that "wrong" can be separated from "wrong in the predicted way". Over five versions of the benchmark, apparent model failures led us to 7 defects. All 7 were in our own tasks or harness, not in the models. The fifth reached a paper we had already submitted to a workshop. Its headline (trap-taking that appeared model-specific, a trapped share of 0.308 against 0.077) is contributed entirely by the one task family whose question let a solution read the label column. Across the other families that version records 0 trap events. On the corrected benchmark, the two models behind that headline take 0 traps in 720 rollouts across three independent runs, while in the third run Phi-3-mini takes 10 traps on the corrected task, so the zero is a measurement rather than a broken scorer. No aggregate metric pointed at any of the 7 defects, except when a trap-free control scoring near zero flagged a broken run; all 7 were located by reading, or trying to read, the agent's output and code. We argue that for agent benchmarks the instrument is the dominant source of variance, that a positive control belongs in every null result, and that accuracy tables should not be believed without transcript-level evidence that the task measures what it claims.