ProofDiff: Auditable Argument-Level Comparison of AI-Generated Mathematical Proofs
Abstract
Mathematical understanding depends on the route from hypotheses to conclusion: how the parts of a proof function and which dependencies carry the argument forward. This structure can persist when wording or notation changes, while a small alteration to a hypothesis or inference can produce a different argument beneath similar prose. For formal proofs, a proof assistant settles whether a rewritten proof still establishes the same statement; informal proofs produced and revised by generative models have no comparable checker, yet the same question of argument preservation arises. Comparing such documents requires an account of their mathematical correspondence rather than a measure of textual similarity. This account must hold whether either document was written by a person or generated by a model. We present ProofDiff, a system that compares mathematical documents through independently constructed argument maps. Each map separates mathematical statements from their proofs and records typed relationships among its components. These components remain grounded in locations in the source documents and, when available, their rendered page regions. ProofDiff proposes correspondences between the maps, including matches across different levels of granularity, and records the confidence and provenance of each proposal. Components with no proposed correspondence remain unresolved rather than being classified as novel. The result is an auditable, source-grounded account that lets readers assess which parts of an argument persist and where the reasoning changes. It is designed for human review and does not claim to certify mathematical correctness. We demonstrate ProofDiff on three illustrative pairs of alternative proofs, using these cases to examine shared structure, divergent proof strategies, and granularity disagreements that a similarity score cannot capture.