When Concern Matching Collapses to Counting: Structural Stress Tests for AI Peer-Review Benchmarks
Abstract
AI reviewers are increasingly evaluated by the article-specific concerns they identify. With article-wise one-to-one matching, the pooled precision-recall score (micro-F1) can become almost entirely a list-length statistic when eligible matches approach the maximum allowed by the two list lengths, the count ceiling. We introduce four stress tests that ask whether rankings depend on list length, concern identity, reference construction, or related manuscript versions across data splits. In BioReview-Bench v4.1, six fixed system outputs evaluated against three reference sets constructed with large language models (LLMs) reach 99.991-99.993% of their count ceilings, and a count-only version of F1 reproduces every ranking. The targets derive from formal journal-review records attributed to human reviewer roles and have not been independently adjudicated. At the historical SPECTER2 cutoff of 0.65, wrong-article reassignment and duplicate padding each retain 99.999% of matches. BM25 retrieves a training-set version of the same manuscript at rank one for all 42 DOI-linked test articles; purging 51 linked rows moves it from first to sixth, whereas none of 100 size-matched control removals changes its first-place rank. PeerReview Bench provides a labeled comparison: GPT-5.4 attains 0.909 balanced accuracy on 163 released expert-labeled pairs, and a sentence-embedding matcher calibrated from those labels reduces wrong-article retention to 8.2% on a non-overlapping 45-paper cohort, although duplicate padding retains 78.1%. Semantic leaderboard claims therefore require evidence that scores respond to concern identity, target construction, and manuscript lineage. Independent adjudication remains necessary for target fidelity, review quality, and scientific correctness.