Attention-Graph Curvature for the Verification of Mathematical Reasoning
Abstract
Language models have recently solved open Erdős problems and achieved gold-medal level performance on International Mathematical Olympiad problems. But their solutions, written in natural language, cannot yet be verified by proof assistants. An erroneous reasoning step can invalidate every step that comes after, so it is advantageous to detect errors early. Several methods have been proposed to detect erroneous reasoning quickly based on the features of the forward pass of the language model that performs reasoning. In this work, we study whether the directed Forman curvature of that model's attention graph, a local quantity computed at an edge from the attention weights around it, carries information about a step's correctness. To do so, we evaluate four per-step verifiers on ProcessBench with and without the curvature appended to their input. Appending curvature values improves the overall mean performance of all four verifiers. ReProbe, the verifier that takes every token of a step as input, shows the largest improvement. Adding the curvature increases the time of the language model's forward pass by 18% on average.