Language-Assisted Image Clustering Guided by Discriminative Relational Signals and Adaptive Semantic Centers
Abstract
Language-Assisted Image Clustering (LAIC) augments the input images with additional textual information with the help of vision-language models (VLMs) to improve clustering performance. Despite recent progress, existing LAIC methods mainly construct image-text pairs for text-modality utilization, which overlook two issues: (i) textual features constructed for images are highly similar, leading to weak inter-class discriminability; (ii) the training is restricted to pre-built image-text alignments, limiting the potential for better utilization of the text modality. To address these issues, we propose a new LAIC framework with two complementary components. First, we exploit cross-modal relations to generate more discriminative self-supervision signals for clustering, which is compatible with the pre-training mechanisms of VLMs such as CLIP. Second, we learn category-wise continuous semantic centers via prompt learning to produce the final clustering assignments. Extensive experiments on eight benchmark datasets demonstrate that our method achieves an average improvement of 2.7\% over state-of-the-art methods, and the learned semantic centers exhibit strong interpretability. Code is available in the supplementary material.