Ground False: Uncovering Errors in Formal Mathematics Benchmarks
Abstract
The ground truths in formal mathematics benchmarks are taken on faith. For theorem proving, released formalizations come without formal proofs that certify their correctness. For autoformalization, faithfulness to the informal source is judged only by the benchmark authors, yet such judgments are inherently subjective and lack collective consensus. We show these assumptions fail at scale: our audit of 367 formal statements in ProofNet finds that 204 (56\%) are unfaithful, and over half of those are mathematically false and cannot be proved. We dissect these errors along three axes (provability, logical strength, root cause) and show that they bias evaluation in predictable, asymmetric directions: false ground truths cap theorem-proving signal and compress model gaps, while equivalence-based autoformalization metrics penalize faithful models more than unfaithful ones. We proceed to fix these ``Ground Falses'' and release the corrected benchmark \textbf{ProofNet-Verified}, produced by a semi-automated pipeline combining Lean provability checks, an LLM faithfulness judge, and human review. Applying the same pipeline to six other benchmarks reveals that unfaithful-GT rates span more than an order of magnitude (4.8\% to 60.0\%), confirming that the failure modes generalize while absolute quality is dominated by curatorial process. Our data and code are available at \url{https://github.com/anonymousauthor567/Ground_False}