PLACE: Patch-Level Agnostic Concept Extraction
Gabriele Onorato ⋅ Fabrizio Silvestri ⋅ Nathaniel D Bastian ⋅ Francesco Restuccia
Abstract
While While Vision Foundation Models (VFMs) are increasingly used in several computer vision tasks, their internal representations remain substantially opaque. This mostly stems from polysemanticity, i.e., individual neurons encode mixtures of unrelated visual patterns. In stark contrast, human reasoning is inherently monosemantic, i.e., it naturally relies on distinct, isolated features to process information. Therefore, monosemanticity needs to be implemented within the model's activation space to obtain representations that are human-understandable concepts. Existing methods mostly operate on convolutional networks, so when applied to VFMs their "crop-and-resize" approach distorts the model's representations by disrupting global self-attention mechanisms and discarding spatial geometry. Furthermore, prior work requires downstream classification labels and is based on KL-divergence, which requires to propagate the gradients back from the classifier head. Ultimately incurring in excessive concept extraction time, and making it hardly applicable to extract task-agnostic concepts. To overcome these issues, we propose PLACE, a fully unsupervised concept extraction framework that operates directly on the native patch-token activations of a VFM. By introducing a geometric Gram-matrix alignment loss, PLACE mechanistically controls downstream KL divergence, thus enforcing faithfulness -- i.e., the notion that extracted concepts, although obtained in an unsupervised fashion, are relevant to downstream classification tasks. Evaluations on image classification tasks (ImageNet and COCO) on DINOv2, ViT-MAE and ViT-B/16 show that PLACE (i) using Gram-alignment it extracts concepts about $20\times$ faster compared to KL-supervised extraction and $11\times$ faster than Sparse Autoencoders (SAEs). Moreover, PLACE yields highly sparse concepts that (ii) achieve up to 40 percentage points sparser compared to the state of the art (i.e., max 0.933 Gini sparsity score), and (iii) are up to 92.3% similar to those obtained under classifier supervision (i.e., 0.923 cosine similarity).
Chat is not available.
Successful Page Load