Reproducible by Construction, Wrong by Construction: The Price of Catching an AI Scientist’s Error
Abstract
Autonomous research agents emit seed-pinned, fully packaged artifacts as a by-product of running at all: every result they produce, right or wrong, arrives at the top of the computational-reproducibility ladder by construction. We present a naturally-occurring case in which an agent's headline medical-imaging result ("vessel segmentation solved, Dice 0.997, +15.6% over the published baseline") regenerates bit-for-bit and is completely invalid: a mask-polarity convention meant the model was scored on segmenting the 83% background, where a constant predictor achieves Dice 0.906. We then measure, on this single fully-audited artifact, a verification yield curve: what each class of check catches, at what cost. Re-execution (≈125 s) catches only tampering, and its one catch here is counterfactual because the file has since been fixed. Five cheap, generator-independent assertions (≈7.3 s total; authored post hoc, with transfer and prospective evidence bounding the circularity of the check class) caught all five invalidating defects. Blinded text-only reimplementation (≈2 min agent time) confirmed the corrected claim. Discovering which assertions to write consumed the 10-20 hour audit campaign that everything else replays. We argue that the automatable share of verification is precisely the share agents saturate, a counterfeit parallel fraction in Amdahl-type accounts of AI-accelerated science. We release the harness, the artifact, and six checklist items, each traceable to a defect that reproducibility checking cannot see.