Approval Is Not Discovery: When Can a Model of the Scientist Steer a Discovery Agent?
Nicholas Chen
Abstract
Discovery agents are increasingly steered by scientists' approval of the experiments they propose, and by models trained to predict that approval. We model a scientist who scrutinizes each proposal with probability $\alpha$ and otherwise approves exactly the proposals that look plausible. Approval then rewards plausible misses and penalizes hits that do not look plausible, and when $\alpha < 1/2$ a model that clones the scientist's approvals agrees with them more often than they agree with themselves while being less valid than a single review, so matching a scientist's consistency need not imply validity. Under this exact-scrutiny model, repeated reviews can audit validity without assays, from four reviews when only approval counts are recorded and $\alpha$ is known to be below one half, or from two when the scientist's quick-check verdict is also recorded. In retrospective simulations on four genome-wide CRISPR screens, approval-trained agents cut hits outside a predefined plausibility annotation set by 79% to 83% while raising their overall hit rate. An agreement gate passes the learned approval predictor in 79 of 80 runs on outcome-enriched candidate lists (26 of 80 on the full gene pool), and neither audit gate grants a clone autonomy. Direct assays at ten reviews each are more precise than the count-only audit in every tested setting, so the audits suit settings where assays are unavailable or slow. A real LLM judge approves plausible misses more often than hits outside the set on every screen, and its own first impression passes the agreement gate against its reviews, although its careful reviews find too little of the truth for the audits to apply.
Chat is not available.
Successful Page Load