Auditing AI peer reviewers: a dose-response and false-positive benchmark on real scientific papers
Abstract
LLM-based peer review is now used at production scale, but rigorous evaluation remains limited. We introduce a benchmark for evaluating AI peer reviewers’ detection of injected errors and quantifying false-positive rates across error type, severity, prompting strategy, and reviewer architecture. We also present skepthical, a multi-agent AI reviewer that verifies citations by retrieving and reading cited papers, checks mathematics with computer algebra, audits numerical claims with code, and uses a cross-model ensemble for general scientific issues. On 15 astrophysics preprints containing 177 injected errors, we benchmark 12 reviewer systems: skepthical; six single-LLM reviewers from three providers, each run with free-form and structured prompts; two Claude Code reviewers using the same Claude model and tools, one with a free-form prompt and one with a structured prompt; and three external reviewers. Detection performance varies sharply by model, prompt, tool use, and error type. Under the same free-form prompt, three frontier LLMs differ by 48.6 pp in total detection rate; a structured prompt halves this spread by improving the weaker models. Within Claude Code, changing only the prompt from free-form to structured raises detection by 32.2 pp. skepthical achieves the highest overall detection rate, 70.6±6.2%, with the best performance on general issues, 82.2%, and citation errors, 46.7%. On numerical errors, it matches both GPT-5.5 configurations at 80.0%, behind structured Claude Code at 91.1%; on mathematics, its 73.8% rate overlaps with both GPT-5.5 configurations within confidence intervals. False-positive rates also vary substantially: free-form Opus 4.7 has the lowest overall rate at 1.0%, while skepthical is the only reviewer with no false-positive failure mode jointly across mathematics, numerics, and citations, yielding a 2.8% overall rate. skepthical is deployed publicly at skepthical-ai.org, and we release the evaluation set as a benchmark.