MathAtlas: A Benchmark for Autoformalization in the Wild
Abstract
Current autoformalization benchmarks are largely focused on olympiad or undergraduate mathematics, while graduate and research-level mathematics remains underexplored. In this paper, we introduce MathAtlas, the first large-scale autoformalization benchmark of in the wild graduate-level mathematics, containing ∼52k theorems, definitions, exercises, examples, and proofs extracted from 103 graduate mathematics textbooks. MathAtlas is enriched with a mathematical dependency graph containing ∼178k relations, and is the first autoformalization benchmark to include such relations, facilitating evaluation and development of dependency-aware autoformalization systems. Our experiments show that MathAtlas is high quality but extremely challenging: single-pass baselines achieve at most 9.8% correctness on theorem statements and 16.7% on definitions. We find that dependency depth is a strong indicator of difficulty, with a frontier agentic system (Claude Code) solving just 29.2% of the deepest items, and doing so inefficiently. We release MathAtlas to the community as a benchmark set for large-scale autoformalization of graduate-level mathematics in the wild.