Mechanistic Interpretability of Financial Time-Series Transformers via Sparse Autoencoders
Abstract
Large transformer-based time-series models achieve strong financial forecasting, yet their internal representations remain opaque. We apply mechanistic interpretability methods to Kronos (102.3M parameters, a financial time-series transformer), training sparse autoencoders (SAEs) on layer-11 activations for six assets over 2023-2026. The resulting features are near-monosemantic (cosine similarity 0.029-0.041, close to the random-unit-vector baseline of ~0.028), economically labeled, and organized by asset class: FOMC/OPEC-sensitive features dominate commodities; earnings-event features dominate equities. Causal filtering via PCMCI+ reduces 25 bivariate Granger links to just 2, leaving gold feature 935 (mean-reversion/momentum, lag 16) and Microsoft feature 526 (earnings-volume, lag 3) as the only direct predictors, a 92% elimination rate. Raw Kronos predictions underperform the majority-class baseline (skill = -0.020); substituting causally filtered SAE features recovers positive skill (+0.016) and outperforms raw Kronos in 64.3% of asset-horizon pairs, with gains concentrated on event-driven assets where the feature taxonomy predicts the advantage. These results show that SAE-based mechanistic interpretability transfers beyond language to financial transformers, and that the resulting feature taxonomy correctly predicts which assets and metrics benefit from causal selection.