Matryoshka Transcoders and Hierarchy Misalignment: When SAE Absorption Protection Does Not Transfer
Abstract
Sparse autoencoders decompose activations into sparse, interpretable features and have become a central tool for mechanistic interpretability. Transcoders extend this approach from representation reconstruction to computation prediction by mapping MLP inputs to MLP outputs, making them attractive for circuit-level analysis. We ask whether a prominent sparse-autoencoder reliability fix, Matryoshka nested-prefix training, transfers to this two-hook setting. In sparse autoencoders, Matryoshka training improves absorption behavior by encouraging broad features to appear in early dictionary prefixes. In transcoders, we find that this protection does not transfer automatically. Across Gemma-2-2B MLP transcoders trained under a matched recipe, Matryoshka models form genuine prefix hierarchies, with reconstruction improving as later groups are added, but they do not inherit the absorption advantage observed for sparse autoencoders. BatchTopK achieves lower full absorption, lower FVU, and higher first-letter F1@1 in the canonical 16k comparison. Bridge controls recover the expected Matryoshka advantage when the cross-space map is removed, showing that the failure is specific to predicting MLP outputs from MLP inputs. We identify the mechanism as hierarchy misalignment. Prefix losses reward covariance with the reconstruction target, so early transcoder groups prioritize target-predictive output geometry rather than broad source-space semantic parents. These results show that hierarchical sparse-dictionary objectives do not preserve interpretability properties by default. For dictionaries that map between activation spaces, the key question is not whether a hierarchy forms, but what geometry that hierarchy is trained to organize.