Auditing Executed Work in Computational Biology Agents
Eva Ge
Abstract
Closed-loop scientific agents need a gate that decides whether an intermediate result is sound enough to build on, and the natural way to make that gate trustworthy is to show the reviewer the executed work. We test this on real computational-biology trajectories from BixBench: a Claude Sonnet agent executes analyses in the benchmark data environment, and three judges review the 116 answers carrying deterministic ground truth. Showing the executed notebook raises acceptance of wrong answers from 0.30 to 0.90 for Sonnet and from 0.06 to 1.00 for GPT-5 mini --- while acceptance of correct answers rises by the same margin, leaving balanced accuracy within 0.02 of chance. Read from the verdict alone, the evidence appears useless. It is not. The same judgements carry a stated confidence, and its AUROC against ground truth is 0.73--0.77 in the notebook conditions against 0.48--0.51 when only the answer is shown: the notebook is precisely what makes the confidence informative. The clearest case is a condition that accepted all 116 answers, so its verdict carries exactly zero information, yet ranks correct above incorrect 66\% of the time by confidence alone. The failure is therefore the decision threshold, not the evidence, and the remedy is to gate on the score rather than the token. On this more sensitive channel we also find no self-review leniency: reviewing its own prior turn, the producing model is $2.55$ points less accept-prone than when the identical work is quoted as another agent's (95\% CI $[-3.43,-1.67]$).
Chat is not available.
Successful Page Load