SASA: Subspace-Aware Sparse Autoencoders for Effective Mechanistic Interpretability
Seyed Arshan Dalili ⋅ Mehrdad Mahdavi
Abstract
Sparse Autoencoders (SAEs) are widely used for mechanistic interpretability in large language models, yet their standard formulation assigns each latent feature a single decoder direction, implicitly constraining features to be one-dimensional. We identify a fundamental geometric mismatch between this assumption and the structure of many model features, and show that it can provably induce feature splitting. In particular, for any feature with intrinsic dimension $d_i \ge 2$, achieving reconstruction error $\varepsilon$ with a standard SAE requires $\Omega \left((\frac{1}{\varepsilon})^{d_i-1}\right)$ distinct decoder directions. As a result, the SAE is forced to represent a single coherent high-dimensional feature using many nearly collinear one-dimensional latents, leading to spurious feature multiplicity and a failure to preserve the intrinsic structure required for mechanistic interpretability. Motivated by this lower bound, we introduce Subspace-Aware Sparse Autoencoders (SASA), which replace single-vector decoders with learned decoder subspaces. SASA enforces block sparsity through Top-$s$ group gating and adapts each group's effective dimensionality using a trace-norm regularizer. We prove a complementary upper bound: when the block size $r \ge d_i$, a single group can represent the entire feature slice, alleviating feature splitting while also improving sample complexity. Finally, we empirically validate that SASA improves feature splitting, monosemanticity, and interpretability while also enabling more efficient training.
Chat is not available.
Successful Page Load