Convergence is Not Discovery
Abstract
While Artificial Intelligence (AI) has already led to various major discoveries in biology, its translation to real world impact has been limited. In this paper, we show one example of why: a model can converge on its objective while the biological conclusion drawn from it does not. Using an open single-nucleus RNA sequencing dataset of lung cancer progression from 23 patients, we optimize a family of flow matching models to predict transcriptional regulators of disease progression across a thirty-twofold optimisation budget: 220 fits spanning four objectives, ten seeds, and 100 to 800 epochs (3,200 for one). Training loss falls 30% and loss on withheld transport couplings 26%. The regulator ranking does not follow: two seeds of the same objective share 6 to 10 of twenty regulators at every budget, with full-list Spearman correlation averaging 0.51. Against a uniform chance line this looks like 6–12×enrichment, but that is the wrong comparison. Holding the fitted model fixed and changing only the biology bounds that floor at 0.125, against which the agreement is at least 1.5–2.6×. Thirteen read-outs applied to the same fitted fields leave the conclusion unchanged: none becomes more reproducible with optimisation, initialisation alone accounts for 88–92% of the per-regulator variance, and no regulator survives a donor bootstrap. Reproducibility trades off against how much of the learned field enters the score. Therefore, a falling loss is not evidence that the mechanism someone reports is reproducible.