Text as Semantic Anchor: Aligning Vision-Language and Vision Foundation Models for Semi-Supervised Segmentation
Abstract
Weak semantic understanding remains a major challenge in Semi-supervised Semantic Segmentation (SSS) models due to limited pixel-level annotations. In SSS, models learn from limited labeled data alongside abundant unlabeled data to reduce annotation costs. While effective at capturing object boundaries, recent methods confuse visually similar classes. Vision-Language Models (VLMs) address this by leveraging language-derived semantic knowledge. Specifically, each class name is decorated with a description that is encoded into a textual embedding, capturing its semantic meaning. However, most VLM-based approaches rely on generic class descriptions, whose textual embeddings fail to capture image-specific class context, leading to vision-language misalignment. To mitigate this, learnable prompts are used to generate adaptive class descriptions. Yet, since they are trained on limited labeled data, these prompts are prone to overfitting, resulting in biased and context-insensitive textual embeddings. Additionally, VLMs suffer from limited spatial precision due to image-level pre-training, while Vision Foundation Models (VFMs) capture rich spatial details but lack semantic understanding. To address these limitations, we propose MVLFormer, a unified Multimodal Vision-Language Foundation transformer for SSS. First, we introduce an image-aware prompt-learning strategy that dynamically conditions per-class textual embeddings on image-specific visual context via an auxiliary network, generating image-specific prompts that improve robust vision-language semantic alignment with minimal supervision. Second, we design a novel textual semantic-driven module that performs stage-wise alignment between VLM and VFM encoder features, using textual representations as semantic anchors to integrate spatially rich yet text-unaligned VFM features with semantically grounded VLM features. With less than 1\% training data, MVLFormer outperforms state-of-the-art methods on benchmark datasets.