Beyond Compilation: Human-Calibrated Evaluation of Semantic Faithfulness in Lean Autoformalization
Abstract
Lean verifies that a generated declaration is well typed, but not that it expresses the statement a user intended. We study three questions for autoformalization without canonical Lean targets: how LLM judges compare with human semantic review, how much compilation overstates faithfulness across systems, and where judges and reviewers disagree. Our criterion combines Lean compilation with strict semantic consensus between GPT-5.2 and Gemini-2.5-Pro. On an independently audited random sample, it agrees with human majority on 91.5% of cases (Wilson 95% CI: 81.6–96.3%). Across eight systems evaluated on 227 graduate-level statements, every system has a nonzero compile–faithfulness gap, whose observed magnitude ranges from 1.3 to 29.5 percentage points. The full GPT-5.2 tool-augmented agent shows the largest gap, compiling 87.2% while satisfying the semantic criterion on 57.7%. Human audits confirm that most outputs in this gap are genuine semantic mismatches. Where all three reviewers agree, the criterion disagrees with them on only 6 of 73 items, mostly because one judge misreads Mathlib; further disagreements arise when the source textbook is itself imprecise. LLM judging is therefore useful as a human-calibrated, conservative aggregate measure, not as an equivalence oracle.