Weakly Supervised Concept Learning for Interpreting and Attributing LVLM Predictions
Abstract
Large vision–language models (LVLMs) are inherently opaque, making it difficult to determine whether their outputs are grounded in visual evidence or driven by language-model priors. Existing concept-based interpretability methods are limited to classification settings, rely on proxy models, or require predefined tokens, and thus do not extend LVLMs. We propose Text-Guided Concept Learning (TGCL), a weakly supervised framework for extracting multimodal concept vectors in LVLMs without token supervision. TGCL builds concept-to-image mappings from data and extracts patch-level activations via concept-guided probing. It then formulates concept learning as a contrastive disentanglement problem, isolating concept-specific patches from background patches to produce sparse, stable, and semantically aligned concept vectors. We conduct experiments on four datasets—ImageNet, MSCOCO, CIFAR100, and DTD—and three recent LVLMs. TGCL outperforms recent interpretability methods, achieving up to 4\% higher sparsity, 11\% lower instability, 17\% lower overlap, and 20\% improvement on attribution faithfulness compared to state-of-the-art baselines.