LinAlg-Bench: A Benchmark Exposing Structural Failure Modes in LLM Linear Algebra - Where Models Stop Computing and Start Hallucinating
Shradha Agarwal ⋅ Tariq J. ⋅ Deepak Rajbhar
Abstract
We introduce LinAlg-Bench, a diagnostic benchmark evaluating 10 frontier large language models on structured linear algebra computation across a strict dimensional gradient of $3\times3$, $4\times4$, and $5\times5$ matrices. Spanning 9 task types and 660 SymPy-certified problems, the benchmark exhaustively evaluates 6,600 model outputs. Beyond binary accuracy, LinAlg-Bench introduces a three-stage automated forensic pipeline classifying 1,156 failures into ten primary error tags, revealing that LLM mathematical failure is not random but structurally constrained by algorithm type and matrix dimension. Our central finding is a sharp behavioral threshold at $4\times4$ scale: below it, models fail through execution errors tracking failures, arithmetic drift, and parity errors; above it, failure transitions to computational abandonment, with models fabricating responses through tool roleplay, constraint-consistent confabulation, and structured hallucination rather than attempting computation. This fabrication-to-abandonment transition is near-universal across all model tiers and architectures, suggesting a working memory limit rather than a knowledge gap. A forced-strategy ablation shows that enforcing algorithmically efficient methods does not recover accuracy—the bottleneck is autoregressive execution depth, not method selection. All data, model outputs, error labels, and judge pipeline will be publicly released.
Chat is not available.
Successful Page Load