Class-wise LoRA Anchoring for Continual Post-Training of Open-Vocabulary Detectors
Abstract
Continual post-training of a vision-language foundation model on a task stream erodes the zero-shot ability acquired in pretraining, a phenomenon known as catastrophic forgetting. We study this in open-vocabulary object detection, whose detectors start from strong zero-shot ability. Parameter-efficient methods mitigate this by training lightweight modules such as LoRA over a frozen backbone, and recent systems organize them into a reusable library of class-specific modules. How the modules in such a library should be optimized as the task sequence progresses is largely unexamined. We propose Class-wise LoRA Anchoring (CLA), a library-aware L2 anchor that keeps each class module close to its task-start state, anchoring a reused module to its consolidated prior and a new module to its low-correlated random initialization. On full-shot ODinW-13, CLA raises post-adaptation zero-shot COCO above the state-of-the-art class-wise library, DitHub, and above the unadapted detector itself, so adaptation stops trading generality for task accuracy, which stays comparable. The gain concentrates on categories never seen during adaptation and comes with a markedly less correlated class library, which makes training-free class unlearning cleaner. CLA is most effective in class-incremental regimes that add new categories, and reveals that a retention-adaptation trade-off remains in domain-incremental settings. Source code will be released.