Language Models Under False Lean 4 Reports: Initial Decisions and Final Proofs
Abstract
A proof assistant can check a proof correctly while an interface displays the wrong result. We study how language models respond to such reports using 14 pairs of elementary Lean 4 statements and seven model deployments. Each model sees a statement, a candidate proof and a report, then chooses whether to keep, revise or reject the candidate. We measure the first decision separately from the correctness of the final submission. In the original Claude batch, fabricated errors reduced correct keeps of working proofs from 40/42 to 2/41. False statements were rejected in 37/42 responses without reports and 35/41 under fabricated successes; this does not establish equivalence. On true statements with insufficient proofs, fabricated successes reduced correct revisions from 18/26 to 5/26 for Opus and Haiku, whereas Sonnet revised 11/13 with a larger output budget. Initial decisions did not always predict final correctness: in the later Opus and Haiku batch, valid final proofs of general statements fell from 28/28 to 20/28 under fabricated errors, and six responses to closed statements recovered from an initial rejection. Four additional deployments showed differing baseline competence. These results support reporting both decisions and final submissions. They do not isolate checking difficulty or establish reliability in an interactive proof workflow.