FACBench: A Benchmark for Formal-Anchor Collisions in Multilingual Mathematical Grounding
Abstract
Mathematical autoformalization can fail before proof generation, when an informal concept is grounded to the wrong Lean/Mathlib object. Multilingual terminology makes this step non-monotonic: aliases can recover missing concepts, but under collision stress they can also pull retrieval toward nearby, non-interchangeable anchors. We introduce FACBench, a targeted dependency-light benchmark and audit protocol for multilingual formal-anchor grounding before proof generation. FACBench combines a five-language concept inventory with controlled collision queries, Mathlib-graph stress tests, no-label ablations, theorem-like probes, and Lean-facing anchor checks. In controlled collision settings, naive multilingual alias union can place the expected anchor in the candidate list while ranking a nearby distractor first. A simple rule-based source/context guard improves top-1 anchor accuracy over naive multilingual retrieval by about 11 percentage points on controlled collisions and 28 percentage points on a Mathlib-graph-mined controlled holdout. Less scaffolded and theorem-like probes show where this guard stops helping and point to the need for role-aware grounding. FACBench provides reusable collision-focused evaluation and regression-test scenarios for retrieval-augmented autoformalization systems.