Tools, Not Talk: The Knowledge Ceiling of Biology Agents
Abstract
Agentic systems for biology are increasingly built around orchestration: critics that review answers, teams of role-playing scientists that debate, and models that reason for longer. We show that on objectively scored biology research tasks this machinery adds little beyond what matched independent sampling provides, and we identify why. Across 150 pre-registered LAB-Bench and LABBench2 questions, compared at a matched budget of three model calls, self-critique never beats simple majority voting, multi-agent debate helps only a model whose reasoning is switched off, and longer reasoning does not improve accuracy even at 25x the generated tokens. The reason is a ceiling set by the questions themselves: which question is asked explains 56% of the variance in correctness, while the choice of model, strategy and reasoning effort explains 0.6%. That ceiling is largely missing knowledge rather than missing deliberation. A single model call with web search and code execution solves 47% of the questions that all eight systems in our main study failed. For life-science agents, the budget belongs in tools, not talk.