Nobody Checks the Instrument: Methodological Critique and the Instrumental Use of Foundation Models in Biology
Abstract
Between 2023 and 2026, at least twelve independent evaluations reported that single-cell foundation models (scFMs) fail to beat simple classical baselines such as logistic regression, PCA on highly variable genes, scVI or Harmony. We ask what that body of criticism actually changed. We assembled every work citing six anchor scFMs (2,099 records), retrieved full text for 998, and hand-coded the 311 that name an scFM in their Methods or Results. Each paper was coded first for how it relates to the model, and only then for whether it validates the model against a simpler alternative. The relation turns out to determine the answer. Papers that develop methods — introducing a rival method, a new foundation model, or a benchmark — report a classical baseline 69-75% of the time. Papers that use an scFM as an instrument, to obtain a biological result about tumours, immune cells or development, do so in 2 of 39 cases (5.1%, 95% CI 1.4-16.9%; Fisher p < 1e-13). Citing a critique is strongly associated with heeding it (78% vs 56%, p < 1e-3), so the criticism works on the papers it reaches. Reach is the bottleneck: only 10% of instrumental users cite any critique, against 50% of benchmark papers, and 59% of all citing papers cite a critique only in an Introduction or Discussion, never where it could alter a design. We also report that the dispute is not settled — three counter-critiques argue the opposite — and that instrumental validation shows no upward trend through 2026. The pathology is not that scientists ignore methodological criticism. It is that criticism circulates inside the community that already practises what it preaches.