Answer-Only Accuracy Measures the Wrong Construct: An Anti-Guessing Verification-Gap Protocol for Trustworthy Olympiad-Math Evaluation
Abstract
Answer-only accuracy is the default protocol for LLM math evaluation, but on olympiad problems it measures the wrong construct: a model can produce the correct final answer without any valid derivation, so answer-match scores reasoning it never performed. We make this measurement failure explicit through the per-problem verification gap (answer-only accuracy minus graded solution correctness) and introduce a simple, low-cost multi-stage filtering pipeline that, applied to any candidate pool of olympiad-style problems, removes problems susceptible to answer-only guessing and retains only those for which it is hard to guess the answer without completing the full proof. The result is a reproducible, human-calibrated protocol that reports what answer-only accuracy cannot: whether a correct answer is accompanied by a rigorously valid derivation. Applying our pipeline to 4,000 NuminaMath-1.5 olympiad problems across two disjoint runs (1,000 and 3,000 candidates), we create Numina-HARD-guess, a benchmark of N=881 problems (200 + 681). We adopt the relative verification gap as the headline metric: the per-problem answer-only accuracy minus solution-correctness accuracy (solution correctness defined as the mean of three judges above 6.5/7 and the answer being correct), normalized by answer-only accuracy and averaged per model. On 8 held-out comparable-tier models (8B-72B, held out from the three filter stages), the relative gap on Numina-HARD-guess is 21.4% (bootstrap 95% CI [18.2%, 24.6%]), roughly half the 41.74% AIME 1983-2024 baseline; the two independent runs yield 20.8% and 21.0%, replicating within 0.21 pp at 3.4x scale. The triple-judge consensus aggregator is validated against 200 IMO-GradingBench human-graded examples (aggregate Pearson r=0.820, 95% CI [0.77, 0.87]; Cohen kappa at threshold 5 = 0.677, substantial agreement), with high binary inter-judge agreement at the 6.5 cutoff (Fleiss kappa = 0.87). We release Numina-HARD-guess (N=881, anti-guessing benchmark from two disjoint runs), the per-stage easy-to-guess pools as transparency artifacts, and the full reproducible filtering pipeline applicable to olympiad-style pools, at a total compute cost of roughly $1,100. Our work provides a fast, low-cost diagnostic of demonstrated solution rigor that complements answer-only metrics.