Idempotency Exposes Consistency Problems in Sparse Autoencoders
Abstract
The Linear Representation Hypothesis (LRH) posits that Transformer activations decompose into sparse linear combinations of interpretable directions, with Sparse Autoencoders (SAEs) serving as the predominant method for recovering such decompositions. We propose that idempotency -- the requirement that re-encoding an SAE's output reproduces the same latent representation -- is a desirable property for any faithful dictionary learning method, and demonstrate empirically that most SAEs deviate very far from this ideal. We identify two underlying failure modes: exploding representation norms and unstable active feature sets. Further analysis reveals that even individual SAE latents cannot be reliably reconstructed under re-encoding, and that SAE outputs exhibit significant sensitivity to input scaling. We show that idempotency can be improved via an auxiliary loss term and demonstrate its potential to improve SAE interpretability.