LeanCompass: Direction Emerges from Distance-Only Contrastive Learning of Proof-State Embeddings
Abstract
Guiding proof search or RL for interactive theorem proving needs a dense progress signal. We train an embedding for Lean 4 proof states with a contrastive triplet loss that supervises only relative proof-distance ordering, never direction, and find that a genuine directional structure emerges anyway: A step's embedding displacement is reliably more aligned with the direction toward the finished-proof state for genuine progress than for a wrong-but-legal tactic step. This separation holds robustly across training-set size (5.7kâ91k triplets). The embedding's distance ordering, the property it was directly trained for, shows the same pattern. Including such wrong-but-legal steps explicitly as training negatives, rather than leaving them out, raises accuracy at separating genuine progress from these to 99.8% (up from 61.1%), and shrinks a wrong-but-legal step's directional alignment by roughly a factor of four, while a genuine-progress step's alignment stays comparable. This is an intrinsic evaluation. A downstream demonstration is future work.