Geometric Instability of Hidden-State Trajectories Predicts Reasoning Failures in Large Language Models
Abstract
Large language models often produce fluent reasoning chains that arrive at incorrect answers, and these failures are hard to detect from output probabilities alone. We show that a strong signal of correctness is reflected in the geometry of internal computation: hidden-state trajectories across transformer layers differ between correct and incorrect reasoning, with correct solutions following smooth, efficient paths and failures exhibiting elevated curvature, abrupt directional changes, and increased geodesic deviation. We capture this signal through VANE, five parameter-free geometric features (velocity, acceleration/jerk, curvature, geodesic deviation, and token coherence) computed from a single forward pass over layer-wise hidden states. Across six models (1.5B to 72B parameters) and five benchmarks spanning mathematics, code, and verbal reasoning, VANE achieves 70.7 to 96.6 AUROC and exceeds token log-probability on all 30 model–benchmark pairs (mean +25 percentage points, up to +50 percentage points). The geometric signal is near-independent of output confidence (r = -0.26), and a base-model control shows it emerges with instruction tuning. As a downstream application, filtering geometrically unstable outputs raises accuracy to 96.1% at 50% coverage on a 4-bit quantized 72B model. These results show trajectory geometry is a reliable single-pass indicator of reasoning correctness, even when the model is confidently wrong.