How Learning Decides Outcome: Comparing Frontier LLMs to Specialists on Biology Tasks
Abstract
Frontier large language models (LLMs) are increasingly used as the planner at the centre of biological pipelines, which raises a practical question: which specialist models do they make redundant? We argue that the answer depends on what the specialist learned from. Specialists trained on knowledge that the literature expresses in natural language, which we call language-substrate models, can be replaced by a capable enough LLM; non-language-substrate models, trained on measurements or on the statistics of protein sequences and structures, keep their own value. We test this on two benchmarks, one for each kind of specialist. To predict whether a missense variant causes gain or loss of function, Claude Opus 5, with no task-specific training, beats the dedicated predictors (by 0.09 AUROC over the best one, LoGoFunc) and a strong feature-based baseline (by 0.17) on variants first reported after its knowledge frontier. Adding these predictors to the LLM does not help, the gene name alone accounts for almost all of the LLM's advantage, and an older LLM only matches the best predictor. To rank variants within a protein by measured fitness, Opus 5 falls behind the best of 27 sequence- and structure-based specialists in every assay category, and describing structure and conservation to it does not close the gap. Where the specialists keep their value, a small stacking model combines them with the LLM better and more cheaply than the LLM can: asked to integrate the specialists, the LLM loses to stacking and mostly follows their average, whereas as one input to the stacking model it improves accuracy on both benchmarks.