ITPEval: Benchmarking Formal Translation Across Interactive Theorem Provers
Abstract
Formal theorem proving has emerged as a frontier challenge for machine learning, yet the ecosystem is fragmented: proofs remain siloed across incompatible systems, limiting both training data for learning-based provers and the portability of verified results. We present ITPEval, the first benchmark for evaluating automated formal proof translation across four major ITPs (Lean 4, Rocq, Isabelle, and HOL Light), spanning two distinct logical foundations (the Calculus of Inductive Constructions and higher-order logic). Our benchmark comprises 1,560 source files spanning 6,848 theorems and lemmas across four systems (390 aligned items per ITP), organized into two tiers: a controlled tier of self-contained, axiomatized files (64 files, 660 lemmas), and an ecosystem tier of 1,496 files drawn from existing libraries and community formalizations. We release itpeval, a unified multi-ITP verification infrastructure with state-isolated warm backends that amortize prover startup while preserving per-artifact native checking semantics. We evaluate both statement and proof translation across five frontier and open-weight LLMs on 12 directed translation pairs: statement translation peaks at 29.1% pass@1 and proof translation at 10.5%; controlled theorems reach 29.7% proof pass@1 versus 5.2% for ecosystem-level translations, confirming that library mismatch is the dominant bottleneck. In an autoformalization/auto-informalization round-trip study, multi-ITP context substantially improves Lean 4 formalization success (4.8% to 10.6%), showing that aligned cross-ITP corpora serve not only as evaluation benchmarks but also as generation-time evidence. Our benchmark, verification infrastructure, and evaluation pipelines are publicly released.