The Verification Gap: Auditing Proof Progress and Repair on IMO 2026
Abstract
In the IMO 2026 campaigns we examined, model graders decided which proofs counted as correct. We took 8 model-accepted proofs, none checked by an expert, and AI agents made 24 single edits: a false step a later line uses (fatal), an argument replaced by an announced omission (gap), or a false line nothing uses (harmless). GPT-5.6-Sol and Claude Opus 5 graded all 32 texts on a rubric that scores “a small error that is straightforwardly repairable” 5–6 of 7. Required of both, a pass mark of 5 accepted 13 of the 16 fatal or gap texts in each of two runs (8 of 11 without the five disputed labels; the other labels are unconfirmed), though a grader quoted the edited passage in each. Requiring 7 instead, or the prompt’s undefined “complete” field (both read post hoc), rejects every fatal edit in both runs but also most harmless ones. With a pass mark of 7, which grader seems more accurate flips in both runs on one choice: whether a harmless false line counts as a mistake. These are counts on hand-picked texts under one rubric. The study grew out of an audit of solving campaigns, where two of our campaign gates accepted a false monotonicity claim that the sequence 2, 1 refutes. We release edits, checks, prompts and responses.