Position: A Score Should Travel With Its Repair History
Abstract
In machine learning a capability claim becomes established when a model outperforms its predecessors on benchmarks the community already accepts, and peer review confines itself to checking that those scores are reported and compared correctly. We argue that benchmarks consequently perform the adjudicating role that peer review performs in other sciences, and that the field has organised no stewardship to match, since a benchmark is reviewed once, when the paper introducing it is accepted, and never again during the years in which it serves as evidence. Our position is that this arrangement, which we call review once and use forever, is the first bottleneck a meta-science of AI research should address. We support the position with the public lifecycle of MMLU, the benchmark its successors describe as the de facto standard for general capability. Between 2024 and 2025 five independent projects repaired four distinct defects in MMLU, none of them by the original authors and none as part of any joint programme. A count across the benchmark paper, model reports, leaderboards and governance syntheses then shows the qualifications attached to an MMLU score falling from 2.4 per reporting event in the repair papers to 1.0 in frontier model reports, three in ten of which carry none. On this basis we propose three norms for venues, funders and governance bodies, and we answer the objection that unstewarded benchmarks have served the field well.