We Need to Improve Benchmarks in AI for Mathematics
Abstract
Benchmarks used to evaluate AI systems for mathematics (both in natural and formal language) exhibit critical shortcomings that limit progress toward genuinely useful mathematical assistants. These limitations range from restricted mathematical complexity to insufficient fidelity in capturing aspects of formal languages such as Lean. Compounding some of these issues is a dynamic reminiscent of Goodhart's law: as benchmark performance becomes the primary optimization target, benchmarks become less reliable indicators of mathematical capability. We explore these limitations and argue that progress requires a course correction in benchmark design and an upgrade to evaluation standards. Additionally, we provide a live, community-extensible website to track the landscape of mathematics benchmarks and encourage the creation of novel benchmarks.