Large language models can not and should not be banned from peer review
Abstract
Scientific peer review continues to treat large language models (LLMs) primarily as integrity risks, often prohibiting their use outright. On the position paper track at this conference, for example, reviewers must commit to not using AI tools to help write their reviews. In light of the difficulties of administering high-quality peer review at conferences that now draw tens of thousands of submissions, as well as the infeasibility of enforcing total bans of this kind, we set out to lower-bound how well state-of-the-art LLMs equipped with web search can approximate the initial stage of paper review. We find that these models produce scores broadly aligned with aggregate human judgment and, for high-variance papers, provide a more stable consensus signal than a held-out human review. They reproduce most human-identified strengths and help identify previously undiscovered weaknesses. Upon further inspection, these LLM-identified weaknesses are roughly on par with human-written weaknesses. Our position is therefore that—while humans should remain involved in every part of the review process—major ML conferences should officially permit LLM assistance during peer review.