TWIN-SAE: Recovering Hierarchical Relations Among Sparse Autoencoder Features
Abstract
Sparse autoencoders (SAEs) decompose language-model activations into dictionaries of interpretable features, but those dictionaries are flat: they record which features exist, not how features relate to one another. Recovering hierarchical relations is hard because the quality of a recovered hierarchy is confounded with the quality of the underlying features, and because feature absorption distorts hierarchical structure in ordinary SAEs. We formulate SAE hierarchy recovery as a quantitative prediction problem, with immediate-parent and ancestor recovery scored against a known ground-truth hierarchy, and introduce TWIN-SAE, which turns absorption into a hierarchy signal by jointly training a Matryoshka branch (TWIN-M) and a BatchTopK branch (TWIN-B) under a soft linkage loss that encourages corresponding latents to remain aligned. On SynthSAEBench, TWIN-SAE attains parent and ancestor F1 of 0.451 and 0.680 against 0.373 and 0.331 for the strongest competing native hierarchy method, ranks first in all five synthetic regimes, and transfers to independently generated held-out worlds; disabling the explicit linkage loss reduces these scores to 0.109 and 0.043. These results establish hierarchical organization as a quantitatively recoverable property of learned sparse representations rather than one that must be imposed architecturally or assessed qualitatively.