How Autonomous Science Fails in Biomedicine: A Pre-Registered Audit of AI Scientist v2 on 11 Research Tasks
Abstract
Autonomous AI systems are increasingly claiming end-to-end scientific research capability, but tooling to verify the science those systems produce lags behind system capability. We contribute a pre-registered evaluation protocol and replication package for independently auditing autonomous scientific-research systems in biomedical settings, and apply it to Sakana AI's AI Scientist v2 (v2). Eleven biomedical research prompts spanning epidemiology, pharmacovigilance, medical imaging, respiratory acoustics, drug repurposing, canceromics, chem-informatics, variant pathogenicity, and adolescent health were pre-registered on OSF prior to running v2's experimentation phase (registration ID withheld for double-blind review; available to program chairs on request) and evaluated under a two-track scoring workflow: an AI-drafted first pass (Track A) and single-reviewer human adjudication (Track B), with both tracks preserved as audit artifacts. Our scoring corpus is 24 v2 manuscript PDFs, 20 v2 ideation records, 3 peer-reviewed high-school-first-author Journal of Emerging Investigators (JEI) papers as positive-control comparators, and 3 v1 template outputs. We independently confirm and biomedically extend observations from Lu et al. (2026, Nature) with three findings. 1. Under our custom rubric and single-reviewer scoring, none of 24 v2 outputs met our predefined peer-review-readiness criterion, while the 3 JEI positive-control comparators — peer-reviewed by construction — met 3/3 at peer review and 2/3 at honors thesis. 2. On quantitative comparisons with bootstrap 95% confidence intervals (n=24 v2 vs n=3 JEI), v2 shows measurably higher citation-hallucination rates (mean 1.12 per manuscript [0.79, 1.50] vs 0 for comparators), lower figure quality (2.38/5 [2.17, 2.58] vs 4.00/5), and lower methods reproducibility (2.67/5 [2.42, 2.96] vs 3.67/5 [3.00, 4.00]); reading-level dimensions (topic understanding, methods soundness, writing fluency) are descriptive-similar. 3. Single-reviewer human adjudication classified all 20 v2 ideation outputs as partially rather than fully novel. Two v2 code bugs (F11 VLM whitelist rejection, F13 UnboundLocalError on pdf_path) fire in all 24 runs at the pinned SHA (96bd516...); both were disclosed to Sakana AI via a public issue on the v2 repository (number withheld for double-blind review; available to program chairs on request) before submission. The 3 v1 template outputs exhibited a silent write-up failure. All quantitative conclusions apply to the pinned configuration and single-reviewer scoring framework evaluated here.