When Interpretability Tools Silently Fail: Feature Absorption and the Reliability of SAE Safety Interventions
Nishitha A
Abstract
Sparse autoencoders (SAEs) decompose the activations of a language model into interpretable features and are increasingly used for safety interventions such as concept erasure, unlearning, and refusal steering, yet these interventions are unreliable and it is unclear why. A leading suspect is feature absorption: a latent that appears to track a concept (e.g. "starts with S") silently fails on arbitrary tokens (e.g. "short"), because child latents tied to specific tokens absorb the concept direction into their decoders $\mathbf{d}_c$. Absorption has so far been studied only on toy graphemic features, measured post hoc, and never connected to downstream safety failures. We ask whether a latent's absorption blind spots can be predicted a priori, before any intervention, and whether they explain where interventions break. On Gemma 2 2B, we first test, fully offline and using only trained weights, whether the set of absorbed tokens replicates across independently trained SAEs (varying seed, width, and architecture), quantified by Jaccard overlap $J(A,B)=\tfrac{|A\cap B|}{|A\cup B|}$ against a baseline matched on token frequency, and whether it is predictable from cheap decoder geometry by scoring $$p(\text{absorbed}\mid \cos(\mathbf{d}_p,\mathbf{d}c),, f{\text{tok}},, \text{depth})$$ with ROC/AUC. We then locate a refusal or toxicity latent and test whether steering and unlearning fail specifically on absorbed triggers, and whether Matryoshka SAEs, which cut absorption from $0.49$ to $0.05$ at matched sparsity, restore efficacy on exactly those triggers.
Chat is not available.
Successful Page Load