LoCo: Selective Local Competition for Discriminative Open-Vocabulary Multi-Label Recognition
Abstract
Open-vocabulary multi-label recognition (OV-MLR) aims to identify all queried semantic concepts present in an image. Existing methods mainly improve global image-text matching or text-side representations, leaving local prediction largely under-explored. In this paper, we revisit local prediction and reveal that it can serve as a strong discriminative branch once two properties are properly handled: spatial consistency of local visual features and compatibility-aware local competition. This perspective is motivated by local semantic sparsity: each image region is related to only a small subset of queried concepts, while compatible concepts of different granularities or types may still coexist locally. We instantiate this perspective with LoCo, a selective Local Competition framework for OV-MLR. LoCo adopts a spatially consistent visual encoder to reduce local feature entanglement, and introduces Selective Softmax to impose competition only among locally incompatible categories while preserving compatible responses. It further combines multi-scale local aggregation with global prediction to handle objects at different scales and scene-level semantics. Experiments on four multi-label benchmarks show that LoCo achieves state-of-the-art performance, with average gains of +11.3\% F1-score and +5.2\% mAP. Our results suggest that local semantic sparsity offers a useful basis for developing more discriminative open-vocabulary multi-label recognition methods. Code is available in the supplementary material.