LLMs Under-Detect Circular Reasoning in Mathematical Handoffs
Abstract
Language models increasingly finish mathematical work that someone else started, but a correct final answer does not show that they checked what they inherited. We give five models partial solutions to 160 competition problems in which the last step shown contains one of three planted faults: an invalid algebraic manipulation, a false lemma, or a circular step that assumes what it should establish. The prompt asks them only to continue. A judge model, blind to model identity and answer correctness, counts a detection only when the continuation explicitly flags or repairs the planted step, and it must quote the passage that does so. Across 13,700 continuations, every model flags circular steps least often (pooled, 27.3%, against 58.2% for false lemmas and 60.8% for invalid algebra), with false alarms on correct steps near 1%. The gap persists at every depth, under a second judge, and after matching or adjusting for surface differences in how the faults were written. In a follow-up with two models, an instruction to check the work raises circular detection from 31% to 55%, against 92–95% for the other faults. Of circular misses, 77.6% still end in the correct answer, against 25–42% for the other faults, so answer checking catches under a quarter of them.