Learning a Unified Cross-Model Semantic Dictionary via Gated Bottleneck Sparse Autoencoders
Jinchi Zhu ⋅ Thomas Tie Luo
Abstract
Foundation models pretrained on similar data distributions tend to encode semantically aligned concepts; yet these representations remain entangled with model-specific inductive biases, which hinders both unified interpretability and cross-model feature reuse. Existing approaches address this by aligning concepts across model-specific representation spaces, but do not explicitly account for model-specific biases, resulting in limited cross-model reconstruction fidelity (low $R^2$) and poor feature transferability. We propose the *Gated Bottleneck Sparse Autoencoder* (GB-SAE), a framework that takes multiple—homogeneous or heterogeneous—foundation models as input and factorizes their representations into shared and private components. GB-SAE jointly learns a single model-agnostic semantic dictionary and, for each model, a private residual dictionary, via a learnable gating mechanism that encourages competition between the two. Conceptually, we cast unified multi-model interpretability as a representation factorization problem and solve it through bottlenecked sparse dictionary learning. Empirically, GB-SAE achieves unified and faithful interpretability across diverse architectures, high reusability of the learned shared dictionary, and strong generalization to downstream tasks—capabilities that alignment-based paradigms, by design, cannot fully support.
Chat is not available.
Successful Page Load