CD$^2$VLOD: A Novel Benchmark for Continual Learning in Pretrained Vision-Language Object Detection
Koyo Imai ⋅ Tsubasa Hirakawa ⋅ Takayoshi Yamashita ⋅ Hironobu Fujiyoshi
Abstract
The advent of large-scale pretrained vision-language models has significantly advanced Vision-Language Object Detection (VLOD), enabling object detectors to generalize from seen to unseen categories. However, detectors deployed in real-world environments must also continually adapt to novel domains that are not well represented in the pretraining data. To evaluate this capability, we propose a new benchmark for VLOD continual learning (VLODCL), termed Cross-Domain Continual Detection for Vision-Language Object Detection (CD$^2$VLOD). The existing ODinW-13 benchmark incrementally introduces novel categories within domains similar to those represented in large-scale pretraining data. Therefore, the ODinW-13 benchmark primarily evaluates how effectively models leverage pretrained knowledge and does not adequately assess their ability to continually acquire knowledge from novel domains. In contrast, CD$^2$VLOD includes specialized domains such as Microscopic and requires models to continually acquire domain-specific knowledge that is not sufficiently represented in the pretraining data. We find that existing VLODCL methods adapt effectively and retain previously acquired knowledge in domains similar to the pretraining data, but suffer from insufficient adaptation and forgetting on CD$^2$VLOD. To address these challenges, we propose a simple baseline method combining replay and knowledge distillation. Our method outperforms existing methods on CD$^2$VLOD and further provides insights for future research by suggesting that the importance of individual VLOD modules may vary across domains. The code will be publicly released.
Chat is not available.
Successful Page Load