An Auditing Agent Asserts at Ceiling Whether or Not There Is Anything to Find
Abstract
Automated agents are increasingly used to audit whether a model carries a hidden objective, and their reports read as confident natural language. We audit one such white-box agent and ask a question its published evaluation does not: what does it say when there is nothing to find? Its assertion rate turns out to be decoupled from the evidence available to it. On byte-identical evidence, a presuppositional prompt yields 9/9 confident and mutually contradictory answers against 2/10 under neutral phrasing (Fisher exact p = 0.0007). Across a six-rung dilution ladder, assertion stays pinned at 1.00 over 120 runs while accuracy falls from 0.55 to 0.05. Detection itself reproduces. The failure is calibration: for evaluators of this kind, confidence carries no information about whether a finding exists.