Training-Free Open World Object Detection
Abstract
Open World Object Detection (OWOD) requires a detector to recognize known categories, detect unknown objects, and incrementally learn new classes. Recent work such as OW-OVD adapts a pre-trained Vision-Language Object Detector (VLOD) to OWOD through fine-tuning. However, this approach demands costly backpropagation and often degrades the detector's own zero-shot performance on known categories, calling into question whether training is necessary. In this paper, we propose CAVE (Clustering in Attribute-Visual Embedding space), a training-free framework that adapts a VLOD to OWOD without any parameter updates. From a single forward pass, CAVE collects per-class visual and attribute statistics via von Mises-Fisher-based clustering, along with attribute co-occurrence patterns, without any gradient computation. At inference, CAVE fuses the original VLOD's text-based scores with MAS (Mixture-Aggregate Score) derived from the collected statistics for known-class detection, while unknown objects are identified by AOS (Attribute Objectness Score) that combines superclass responses with attribute co-occurrence patterns. Experiments on M-OWODB show that CAVE outperforms OW-OVD by +16.2 known-class mAP (Task 4) and +28.0 average unknown recall, while surpassing the zero-shot VLOD across all incremental tasks.