Repeated-run reliability of biomedical agents: measurement pitfalls and selective deferral
Abstract
Biomedical agents often answer differently when run twice on the same task. We ran four agent, benchmark and decoding configurations (Biomni, GenoMAS, Au- toBA, Open-Rosalind; 13–118 tasks ×4 runs), marking results as pre-registered, pre-specified or exploratory. Configurations differ mainly in how they fail: AutoBA answers 46% of runs but is right on 88% of those; GenoMAS answers 93% and is right on 40%. When answers count as the same if and only if the grader would score them identically, and failed runs never agree, agreement among Biomni’s runs predicts correctness (run-level AUROC 0.84, task-level 0.734) but weakly within a task (0.65 [0.51, 0.79]): the signal mostly reflects which tasks are hard. Failure handling reversed AutoBA’s conclusion, and answer-key granularity changed GenoMAS’s AUROC from uninformative by construction to 0.734 (post hoc key); we give a reporting checklist. In an exploratory replay of archived Biomni runs, accepting an answer only when two runs agreed cut error from 31% to 13.3% at 41.5% coverage, for 2.0×the agent tokens (12.3% at 36.7% with the first-launched run drawn first). Post hoc, in an earlier pre-registered controller test whose hypotheses failed, the 65 of 150 fresh tasks where the first two runs agreed had 12.3% error [6.4, 22.5], under that study’s earlier failure conventions. Deferral did not clearly help GenoMAS. Three correction tests did not robustly beat plurality voting.