Test-Time Adaptation via Self-Reinforced Optimal Transport for Zero-Shot OOD Detection with Vision–Language Models
Abstract
Pre-trained vision-language models (VLMs) enable zero-shot classification by matching images with text prompts, yet their raw confidence scores are unreliable for out-of-distribution (OOD) detection. We identify a confidence bottleneck: zero-shot VLMs produce posteriors whose maximum confidence yields much weaker ID/OOD separation than supervised classifiers trained on the target label space. To theoretically explain this gap, we relate VLM confidence to a well-separated supervised reference through a Markov-kernel transformation and analyze single-threshold detection using the Youden/KS index. Our analysis shows why raw VLM confidence is unlikely to recover the clean separation achieved by supervised classifiers. Motivated by the observation that class-wise posterior distributions contain useful structure beyond maximum confidence, we propose Self-Reinforced Optimal Transport (SROT), a label-free test-time framework that reshapes VLM posteriors with optimal transport and adaptively trains a lightweight OOD detector from rectified pseudo-labels. The detector further feeds open-set evidence back into posterior alignment, forming a self-reinforcing loop between distribution-level rectification and feature-level OOD learning. Extensive experiments on standard zero-shot OOD detection benchmarks show that SROT achieves state-of-the-art performance, with ablations validating posterior alignment, adaptive detector training, and self-reinforcement.