Bigger Isn’t Better: Why the Indiscriminate Scaling of Foundation Models Can’t Solve Biology
Abstract
In the wake of AlphaFold’s spectacular achievements, a new generation of biological foundation models have promised to bring about similarly transformative advances in other life science domains. However, we argue that the availability of prefabricated datasets, architectures, and benchmarking metrics—coupled with the pressure to publish novel results—has led computational biology to prioritize scale, visibility, and convenience over progress. Without confronting the inbuilt limitations of our existing data and components, and without realigning the institutional and professional incentives that drive research decision-making at the programmatic level, we risk misallocating effort at great cost to our ability to achieve meaningful biological insight. This paper focuses on foundation modelling in proteomics, genomics, and single-cell biology, offering a critical synthesis of common failure modes. We highlight empirical challenges that have been brought against a number of high-profile models as case studies, and—using AlphaFold as a countercase—explore what it would take to build a research ecosystem that would facilitate the development of theory-driven, fit-for-purpose models.