Verifying the Verifier: Agreement with Human Reviewers Does Not Measure the Validity of LLM-Generated Reviews
Abstract
Automated reviewers are entering the scientific pipeline: autonomous research systems gate their own output on an LLM reviewer, and review-generating products are already deployed. The standard way to evaluate such a reviewer is overlap between its criticisms and human reviewers’ criticisms. Overlap measures agreement, not whether the criticisms are true, and we show that this distinction decides rankings in practice. We compare a deployed multi-agent reviewer with the published single-agent baseline on 50 ICLR 2026 submissions with real reviews and decisions, holding the underlying model fixed at every stage so architecture is the only variable, and we add a validity layer: three judges from three model families rate every criticism for whether an area chair would weigh it. The two measures give opposite verdicts on the same systems. The multi-agent reviewer scores 0.106 lower in F1 against human reviewers, yet its extra criticisms are confirmed as decision-material 1.6 to 2.9 times as often as the baseline’s, yielding 4.26 confirmed defects per paper that no human raised, against 1.26. Both differences are significant (p≤0.002), while recall is statistically identical, so the entire difference lies where overlap can only count errors. An evaluation that measures agreement alone ranks the reviewer that imitates humans above the one that finds more real defects; measuring validity directly corrects the ranking.