Automated Review of Mathematical Papers: Errors Found and Repaired in the Top Four Journals
Abstract
AI can now solve research-level problems in mathematics. We ask what it can do for the review process at a journal, and test it on the published record of the Annals of Mathematics, Inventiones Mathematicae, the Journal of the American Mathematical Society and Acta Mathematica. We gave GPT-5.6 Sol 2,272 papers, most of the research articles these journals printed in the past twenty years, in their published form, and asked it to report every error it found and the step at which the proof fails. It found that in 30\% of the papers the proof of a main theorem is not rigorous as written, and that in 4\% the main theorem is false as stated. Only 4\% of these errors have been noted in print. We then gave each affected theorem to GPT-6 Astra to prove or refute without human assistance. It restored 87\% of the theorems with their contribution unchanged, two thirds of these by a local correction and a third by replacing a substantial part of the proof, and 12\% in a narrower form that keeps the core of the result. It found an explicit counterexample to 7 of the remaining 10 theorems, and only 3 remain open. The results suggest that automated review is already useful to authors checking a paper before submission and to reviewers and editors evaluating it.