SpecBridge: Learning Natural-Language Formalization Plans for the Formal Specification Synthesis Task
Abstract
Software requirements are usually written in natural language, which is essential for human communication but insufficient as a target for machine-checked correctness. While natural-language descriptions can state intended functionality, verification requires formal specifications with explicit types, relations, quantifiers, guards, witnesses, and edge cases. To bridge this divide, we study the Formal Specification Synthesis task: given a natural-language programming requirement and fixed formal signatures, synthesize a proof-assistant specification for the intended input-output relation. This task is difficult for two primary reasons. First, a large representation gap separates informal natural language from the rigorous structures needed by a proof assistant. Second, generated formal specifications are notoriously hard to evaluate, as proving full equivalence between specifications is often too complex to serve as routine feedback. To address these intertwined challenges, we propose SpecBridge, a reconstruction-guided formalization framework. At its core, SpecBridge introduces a Natural-Language Formalization Plan (NLFP), a semi-structured, readable intermediate representation that captures key formalization choices before decoding to a target specification language. Through reconstruction-guided pattern mining, SpecBridge learns these NLFPs by reconstructing formal specifications to discover reusable patterns, ultimately transferring them to improve natural-language task generation. Furthermore, to tackle the evaluation bottleneck, we introduce a multi-layer evaluation protocol for formal-specification correctness. This protocol utilizes LLM Specification Alignment as an LLM-based judge, alongside Test-Driven Computable Validation and Test-Driven Proof Validation, which approximate "running" testcases on formal specifications by translating propositions into executable pre/post checks and proving testcase obligations in Lean. In our experiments, SpecBridge outperforms the few-shot + CoT baseline by 6.1% on LLM Specification Alignment, 15.2% on Computable Validation Pass Rate, and 28.4% on Proof Validation Pass Rate. These results demonstrate that the NLFP bridge and our multi-layer evaluation protocol make formal specification synthesis significantly more reliable.