Stop Mechanizing Reform Heuristics as Scientific Quality Filters in AI Review
Abstract
AI-driven review is poised to structure which science gets attention. By implementing checks for reproducibility, robustness, preregistration, claim scope, and other proxies, machine learning research is institutionalizing metascientific filters on scientific production. This position paper argues that mechanizing reform heuristics whose theoretical grounding remains contested or incomplete is counterproductive to scientific progress. AI review papers should be evaluated as proposed decision policies that shape future research. However, currently the emerging literature blurs the line between integrity filtering, based on necessary but insufficient signals of validity like reproducibility of stated results or lack of fake citations, and epistemic filtering, which uses machine-detectable signals to judge scientific quality. Drawing on debates in metascience, we show that proposed filters--including replicability, multiverse robustness, and preregistration of analysis surfaces--are insufficiently justified as general indicators of scientific value. We argue that human-in-the-loop review fails to resolve the problem, because automated signals shape attention and create incentives upstream. Instead, the field must move toward more rigorous motivation of signal-to-decision pipelines, including explicit specification of target constructs, linking assumptions, decision uses, failure modes, and incentive effects.