Concept-Localized Generative Representations
Abstract
Motivated by interpretable generative models for applications such as scientific inquiry, we aim to learn representations that localize human-interpretable concepts and separate them from residual variation in the data. Reconstruction and prediction objectives alone are insufficient for this purpose, as concept-relevant variation can be distributed across multiple latent factors while still achieving high performance. We address this limitation with a generative model that learns a balanced decomposition of data into concept-aligned and residual latent representations inferred purely from observations. The objective encourages localization through an upper bound on the concept information bottleneck while reducing dependence between the concept and residual latent. Across synthetic and real datasets, the learned representations localize concepts while preserving generative fidelity. A neural-data case study shows that localized representations align more strongly with structured variation in observed neural responses. This supports localization, rather than prediction accuracy alone, as a more faithful criterion for interpreting concept-relevant structure in generative models.