Diagnosing Math-Reasoning Failure Structure with Milestone Oracles
Abstract
Aggregate math benchmark scores tell whether a model solved a problem, but not where reasoning broke down. We introduce milestone-oracle probing, a symbolically verified diagnostic that turns math evaluation into a reasoning-gap readout. For each parent problem, a teacher writes a fixed milestone roadmap. We evaluate a student under three levels of help: no help, the roadmap, and the roadmap plus gold milestone answers. We also run a separate milestone-only test, asking whether the student can solve each milestone in isolation. Deterministic symbolic verification grades all parent and milestone answers. Crossing these two measurements yields five gap categories: roadmap, milestone-execution, composition, missing-milestone, and capability. Applied to 354 NuminaMath problems and a six-model panel from 8B to 671B parameters, the largest residual category across all six models is families that pass the milestone test but are not recovered under any of the three probes (33--48\% of families per model). A follow-up audit on a broader sample of unrecovered families shows the residual is mixed: it contains real composition failures together with decomposition gaps and verifier artifacts. The protocol surfaces the residual; the audit reads it into its component parts. We release the diagnostic set, prompt templates, symbolic verifier, audit annotations, and the leak-safe repair logs and bookkeeping scripts.