A Guard-Aware Junk-Value Linter for Lean 4: Auditing Formal Benchmarks and Improving Autoformalization
Abstract
Lean's kernel certifies proofs of formal statements, but not whether those statements faithfully express the intended mathematics. We present a Lean 4 linter built on a curated registry of 1,056 declarations with "junk" behavior. On a controlled bank of 149 natural-language statements, linter feedback raises the proportion of true guard-required formalizations produced by Haiku 4.5 from 48.0% to 70.1%, with smaller positive effects for stronger agents and no systematic degradation on control or intentional fallback items. We also test the linter on 13 Lean benchmarks and find 269 candidate unfaithful statements. Of these, we provide Lean certificates via their junk values that establish 96 false declarations, 16 complete junk-enabled or trivial proofs, and 11 declarations with inconsistent hypotheses.