Local–Global Sparse Autoencoders for Multiscale Interpretability in Vision Models
Abstract
Visual understanding is inherently hierarchical as high-level semantic concepts are composed of fine-grained, spatially localized components. Yet existing interpretability methods typically operate at a single scale, either local or global, leaving a critical gap in understanding the compositional structure between them. Recovering this structure without supervision is particularly challenging, as the correspondence between local and global concepts is latent, varies across samples, and may appear only through inconsistent or partial patterns. To address this challenge, we introduce Local-Global Sparse Autoencoders (LG-SAE), an unsupervised framework for multiscale interpretability in vision models. The core innovation of LG-SAE is a learned compositional bridge that explicitly captures the relationships between local and global concepts. By jointly decomposing patch-level and image-level embeddings, LG-SAE induces a sparse interpretable graph that reveals how each high-level semantic concept is supported by a small set of spatially grounded components. Building on this representation, we present a unified framework for discovering local-global concept relationships, constructing multiscale Concept Bottleneck Models, spatially aware concept naming, and model steering through localized concept edits. Extensive evaluations show that LG-SAE: (i) captures meaningful cross-scale feature relations; (ii) enables high-quality spatial grounding without compromising downstream accuracy; and (iii) provides interpretable model explanations, further validated by a user study.