Repair, Retry, and the Cost of Hindsight: Verifier Feedback in Formal Theorem Proving
Abstract
Verifier-guided self-repair conditions a model on feedback from failed attempts before it tries again, without updating model weights. Yet two findings in prior work are in tension: verifier feedback can improve repair of failed outputs, but self-repair itself can provide little-to-no advantage over independent resampling. Here, we provide an empirical reconciliation in Lean theorem proving through controlled one-step experiments that deconstruct repair by holding the preceding failure fixed while varying whether the failed proof is shown, whether the compiler error corresponds to that failure, and how the model is instructed to try again. We test frozen Qwen3.6-27B, a general-purpose model, and DeepSeek-Prover-V2-7B, a specialized prover model, on difficult MiniF2F problems. We find two opposing effects. Displaying the model's failed predecessor substantially reduces success, while on Qwen adding that attempt's compiler error recovers this loss. Replacing the predecessor with another theorem's failed proof largely removes the degradation on both models; on Qwen, replacing the diagnostic with another problem's error substantially reduces the recovery. Without the failed proof, correct diagnostics show no resolved advantage over wrong-diagnostic controls on either model. The failed-proof penalty persists with and without the retry instruction, whereas Qwen's diagnostic recovery is weaker without it. The predecessor penalty is accompanied by sharply increased copying of the model's own failed proof. Separate 11-problem sequential experiments on both models extend the comparison through 64 attempts: Qwen retains a correct-diagnostic benefit within failed-proof histories, whereas DeepSeek-Prover shows no resolved within-repair recovery. However, these recovery effects do not imply an advantage for self-repair as a test-time technique. No tested one-step feedback condition establishes an advantage over instruction-only resampling, nor does sequential repair with failed proofs and correct diagnostics at 64 attempts. These findings support interpreting the usefulness of verifier feedback separately from the net performance of self-repair, since a diagnostic can improve repair without overcoming the performance loss from conditioning on a failed attempt.