Bypassing PC1 Makes SAEs More Reproducible
Nathan Delisle ⋅ Chenhao Tan
Abstract
Sparse autoencoders (SAEs) are widely used to decompose language model activations into interpretable features, but recent work finds that many features are not reproducible across random seeds. Intuitively, the more important a feature is, the more reproducible it should be. We test this hypothesis and find a surprising inversion: the 50 most important features by zero-ablation are among the least reproducible across seeds, far below dictionary-wide baselines. We trace this to residual stream anisotropy, a well-studied phenomenon in NLP embeddings that has been documented in transformer representations but not connected to SAE training. In transformer middle layers, a single principal component (PC1) explains 70--99.9\% of activation variance, and its sparse decomposition is non-identifiable: different seeds learn different features that tile the same direction. These features dominate importance rankings, explaining the inversion. We propose an intervention: bypassing PC1 during SAE training, routing it around the SAE and restoring it during inference. Across 9 models from 5 families (124M--70B parameters), this intervention raises top-50 feature recovery to 74--97\%. The recovered features show greater causal consistency under steering ($p < 0.01$ in 8 of 9 models), and shared features produce $4\times$ more accurate logit-lens predictions than orphan features. Our results reconcile recent SAE instability findings with the existence of dense model features: the dominant PC1 direction is real and reproducible, but the individual sparse coordinates used to reconstruct it are not. Once this component is separated, a reproducible core of sparse, semantically coherent SAE features emerges.
Chat is not available.
Successful Page Load