A Broken Verifier Emits a Pass: 238 Audits from a Deployed Review Panel Run by a Scientist Who Cannot Read the Code
Abstract
This workshop asks how to trust, judge, and act on AI-generated science when verifiers are imperfect, scarce, or absent. We report a verifier deployed in that regime: a bounded two-vendor review panel that has guarded an environmental metagenomics programme for two months, whose principal investigator cannot read the analysis code and therefore has no option but to verify by mechanism. We classify every audit it was asked to perform (n = 238) from its own logs, and we report its failure modes at more length than its successes, because those are the part a reader can act on. (i) The failure that matters is not a wrong answer but a verifier that has silently stopped verifying. A hand-off state was, at the caller's interface, indistinguishable from the second reviewer never having answered; our own first measurements read one for the other and put second-side non-delivery at 92%. Measured by evidence it is 2 of 10 hand-carried audits, 2 of 228 under the orchestrator and 0 of the 52 under the current release; our record cannot attribute that decline to any single repair, and we say so rather than take the credit. What instrumenting the interface did establish is that the quantity is measurable at all. The same shape then recurred three times after release, in three separate layers, each shipped through a passing test suite: an approval that covered a finding which had quietly dropped out of the record; a single-reader run returning the terminal state of a two-vendor audit; a statement of work accepted by the schema and delivered to nobody. (ii) A check written by the author of the work inherits its premise. One validator passed twelve of twelve assertions while being structurally incapable of testing whether its premise held; a worse variant pins the currently measured numbers as its expected values, stays green, and would reject the correction at the moment it arrived. The measurement reported here was itself wrong five times -- in its unit, twice in its population, once in the meaning it assigned to a status field, and once because the log corpus it reads expires on a retention schedule and a third of it was deleted between freezes. Not one of its six adversarial tests could have failed in any of those cases (Section 4). (iii) Adversarial review reliably produces plausible criticism, which is not defect detection: roughly 2% of review-driven additions traced to a real defect, while a 27% resource miscalculation survived two review rounds and was awarded three further decimal places. (iv) Of the 154 audits in which both reviewers delivered, 4 ended in agreement (2.6%, [1.0,6.5]); we read non-convergence as a question about the problem statement rather than a shortfall of compute. The fail-closed defaults cost about one refused audit in four, which we report rather than net out. The system is an installable, domain-neutral skill, released under Apache-2.0 at https://github.com/ttomasyoung/dual-audit.