MOSAIC: Concept Bottlenecks via Text-Anchored Optimal Transport
Abstract
Concept Bottleneck Models (CBMs) provide interpretability but typically incur a penalty in accuracy. Recent CBMs rely on LLM-supplied concept vocabularies that are misaligned with the underlying image evidence and obscure dataset-specific discriminative signals. Augmenting these representations with task-specific linear probes recovers little of the lost accuracy. Methods that instead learn concepts directly from data also fall short of the supervised CLIP linear-probe ceiling. We introduce MOSAIC (Mixtures Of Sparse Anchored Interpretable Concepts), a CBM whose concepts are discovered directly from image representations via text-anchored optimal transport, a balanced-marginal pass that softly assigns each image to a small set of learned concepts while keeping them tethered to class-text embeddings, and then named post-hoc by a VLM. Across 5 datasets, MOSAIC matches the supervised CLIP linear-probe ceiling on most dataset-backbone settings and beats raw zero-shot CLIP by +5.2 to +24.4 top-1 and class-level retrieval R@1 by +5.7 to +27.0. Beyond classification, the discrete concept basis enables a sparse per-image decomposition that exposes which features drive each prediction and aggressive bottleneck compression to as little as 2.5% of dense CLIP.