When Spectral Attention Graphs Fail: A Confound-Aware Study of Attention Graphs in LLM Reasoning
Abstract
In recent interpretability research, the spectral analysis of a transformer's layer-wise attention graphs has been explored as a signal for the correctness of an LLM's reasoning. To our knowledge, existing work has not jointly tested whether this signal survives potential confounds that may influence attention structure and its spectral properties: response length, attention-sink mass, and effective rank. We construct symmetrized attention graphs across Qwen2.5-0.5B's transformer layers on the GSM8K and ARC-Challenge benchmarks and analyze their Laplacian spectra. We then build a leakage-safe classifier to test whether spectral features at individual layers (static features), their layer-to-layer changes (dynamic features), or their combination predict reasoning correctness beyond potential confounds. We further apply the Benjamini-Hochberg (BH) correction to account for multiple testing across layers and metrics. For GSM8K, statistically significant differences in spectral features between correct and incorrect reasoning groups are substantially explained by response length, attention-sink mass, or effective rank. For ARC-Challenge, raw spectral effects are minimal, with only one BH-significant layer-metric pair and no statistically significant predictive improvement beyond the confound-only baseline. Further, across both benchmarks, dynamic features do not outperform static features or the confound-only baseline in predicting reasoning correctness. Our results position confound auditing as a necessary step for spectral interpretability of LLM attention graphs.