Position: Agentic AI scientists should halt for the expert’s insight and quality control
Abstract
AI scientists generate, evaluate and rank scientific hypotheses, and the systems published in recent years differ considerably in how much of a workflow run proceeds without human checks and intervention. Here, we argue that, for interdisciplinary science R&D, quality control should sit directly inside the AI scientist workflow: the model must search outside the expert's own field, where it adds the breadth the human expert lacks, and the run must halt at fixed checkpoints that only the expert's decision can pass, because the knowledge that decides whether a scientific concept is worth pursuing is held by people and is not in model training data. The position is based on an asymmetry: frontier models reason across every field at the level of a graduate student, but a domain researcher knows one field far more deeply. As an implementation of a human-in-the-loop AI scientist, our Breadth Engine agentic workflow instantiates the position in a six-phase workflow to invent new research directions; it decides which science to do but executes none of it. In a two-hour run on a chemical-sensing problem, the system returned 10 research opportunities from 10 mechanism families outside conventional chemical-sensor development, and stress-tested selected concepts down to a gating experiment. A pilot with 20 researchers is running at our R&D institute, with quantitative results on invention expected in October 2026.