NegSep: Analyzing and Mitigating Background Shortcuts in Negative-Label Zero-Shot OOD Detection
Abstract
Vision-language models (VLMs), such as CLIP, have shown strong capabilities in zero-shot out-of-distribution (OOD) detection. However, negative-label-based approaches remain vulnerable to background shortcuts, where contextual cues spuriously attract OOD images toward in-distribution (ID) labels. We analyze this failure by examining the semantic composition of the negative label space and how heterogeneous negative evidence is jointly scored. Based on this analysis, we propose NegSep, a training-free framework that separates negative labels into background-related and non-background subsets and evaluates them through independently normalized scoring branches. Shortcut-prone subset and branch-level analyses show that background-related negative evidence provides complementary OOD signals for shortcut-induced false ID acceptances. Experiments across multiple OOD benchmarks demonstrate that NegSep substantially improves robustness under background shortcuts while remaining competitive on standard OOD benchmarks, without retraining, prompt tuning, or region-level supervision.