Memory-Driven Contrastive Embedding Enhancement for Fine-Grained Open-Set Semi-Supervised Learning
Abstract
Existing VLM-based open-set semi-supervised learning (OSSL) methods primarily rely on coarse-grained class-level semantics, thereby hampering the accurate identification of in-distribution (ID) versus out-of-distribution (OOD) samples. To overcome this issue, we propose a novel memory-driven contrastive embedding enhancement approach by integrating fine-grained visual contrastive cues into class-level textual side. Concretely, we design a memory bank to construct and maintain a diverse collection of the most distinctive fine-grained visual embeddings. Guided by the memory bank, visual information is incorporated into the textual side through cross-modal contrastive learning, enabling more discriminative fine-grained separation between ID and OOD samples. Consequently, the learned OSSL model exhibits improved contrastive discriminability across known and unknown categories. Extensive experiments demonstrate that our method achieves state-of-the-art (SOTA) performance on multiple fine-grained datasets. The code is available at https://anonymous.4open.science/r/CodeForPaper.