Learning Semantic Consistency for Open-Vocabulary Dense Perception
Li Ding ⋅ Junjie Wang ⋅ Jingjun Yang ⋅ Jiaze Wang ⋅ Libo Qin ⋅ Li Jiang ⋅ Zhuotao Tian
Abstract
Open-vocabulary dense perception requires grounding language-specified concepts in local regions for recognition, localization, and segmentation beyond closed category sets. CLIP learns image-level vision-language alignment, whereas dense prediction requires spatially precise and semantically reliable local features. Although recent dense CLIP adaptation methods reduce this gap through region-level semantic transfer and spatial correlation guidance, maintaining stable local semantics under contextual variation remains challenging. Specifically, the global contextual modeling in CLIP may introduce undesirable dependencies into local features. Changes in background, scale, viewpoint, or scene layout may cause identical object regions to yield inconsistent dense representations, leading to unstable correspondences and semantic drift. Moreover, spatial correlation guidance from visual foundation models captures visual relatedness between patches, but not whether coherent regions correspond to the queried textual concept, leaving concept-level dense alignment under-constrained. To address these issues, we introduce Semantic Spatial Consistency Learning $\textbf{(SSCL)}$, an adaptation method with two complementary objectives: 1) Context-Invariant Spatiotemporal Consistency (CISC) enforces semantic consistency for the same anchor region under varying contexts to mitigate context-induced feature drift; 2) Structure-Guided Semantic Refinement (SGSR) leverages spatial affinities from a frozen visual foundation model to refine CLIP text-patch responses, producing text-conditioned dense supervision with improved spatial coherence and concept-level alignment. Experiments on video instance segmentation, region classification, semantic segmentation, and object detection show consistent gains over the baseline methods, demonstrating the effectiveness and generalization capability of the proposed SSCL. Our code and models will be made publicly available.
Chat is not available.
Successful Page Load