When The Assay Lies: Evaluating Artifact-Aware Agents in Biological Discovery
Shivum Telang
Abstract
A closed-loop AI scientist reasons about which biological hypothesis explains a result. It rarely reasons about whether the measurement itself produced the result. We formalise artifact-aware discovery, in which a technical artifact is a first-class competing explanation, and release \textsc{Mirage-Bio}, an interactive benchmark of $\nEp$ episodes built from two documented artifact classes: copy-number-driven toxicity in CRISPR screens and direct compound interference with firefly luciferase. Every episode carries exact ground-truth likelihoods, so the optimal policy is computable rather than annotated. Two consequences follow and we verify both. Repeating the assay that produced the result leaves the log-odds between the biological and artifact explanations \emph{exactly} unchanged --- measured drift $\drift$ over $\nRepeats$ repeats on all $\nEp$ episodes --- so an agent without an artifact hypothesis is already at $p(\text{biology})=\confNaive$ after the first readout, including on the episodes where no biology is present. And standard expected information gain is an actively misleading objective here: it ranks an experiment that provably cannot separate the hypotheses above every diagnostic one in $\priorTrap$ of $\nEp$ episodes, while an objective restricted to the competing contrast separates the two classes perfectly. Among reference policies, planning on that contrast reaches $\accMir\%$ accuracy $\accCIMir$ at cost $\costMir$ and spends $\wasteMir\%$ of its actions on experiments that cannot decide, against $\accRep\%$ and $\wasteRep\%$ for a policy that repeats the original measurement. The benchmark, the exact references and the agent harness are released.
Chat is not available.
Successful Page Load