Position: Automated Scientific Verification Needs Standards Before Scale
Abstract
AI agents are beginning to reproduce computational results, inspect scientific code, and test results under alternative analytical choices. This could extend scientific screening beyond what human reviewers can sustain. Yet these assessment tasks answer different questions. Computational reproduction asks whether specified data and code produce reported outputs; code auditing evaluates whether an imple- mentation is consistent with the stated method and relevant constraints; robustness analysis studies how conclusions change across defensible choices. None alone establishes that a scientific claim is true. We argue that the range of possible assessments within and across these tasks makes community-governed standards necessary before automated outputs acquire broad institutional or reputational force. Today, heterogeneous tasks, rubrics, tolerances, and labels can make superficially similar outputs not directly comparable. A shared core could make the judgments embedded in those outputs explicit, testable, and contestable, while domain and use-specific standards remain free to differ and evolve. We explain why ad hoc practice will not scale, identify the epistemic and institutional requirements for legitimate standards, and outline a staged process for developing them. Standard- ization should remain empirically calibrated, pluralistic across domains, and alert to the danger of rewarding what is easy to validate rather than what is scientifically meaningful.