COPRA: Conditional Parameter Adaptation with Reinforcement Learning for Video Anomaly Detection
Abstract
Vision-language models (VLMs) have demonstrated impressive performance in video anomaly detection (VAD), while providing interpretable predictions. However, we identify a key yet underexplored VAD issue: the mismatch between training and inference in both data distribution and model configuration. First, many VLM-based VAD approaches rely on a static parameter paradigm after training (or post-adaptation), causing them to overfit the training distribution and limiting generalization to unseen domains. Consequently, these methods lack effective test-time adaptation under substantial distribution shifts, such as novel environments or previously unseen anomaly types. Second, they treat VLMs as reasoning-based classifiers trained on extremely sparse frames from long videos, but perform inference on densely sampled short segments via sliding windows, leading to inherent inconsistencies between training and testing. To address these limitations, we propose COPRA, a conditional parameter adaptation framework for VLM-based VAD. Rather than relying on static post-training adaptation (e.g., fixed prompts or shared parameter updates), COPRA introduces instance-conditioned parameter adaptation, where a lightweight generator predicts input-specific parameter updates to dynamically modulate a frozen VLM on the fly, enabling per-segment adaptation during both training and inference. Empirically, our method achieves strong performance on standard VAD benchmarks. More broadly, it consistently outperforms static baselines in both in-domain and cross-domain settings, and further generalizes beyond VAD to unseen tasks such as multiple-choice Video Question Answering and Dense Captioning. These results suggest that COPRA provides an effective weight space generation mechanism for foundation models, enabling more scalable, adaptive, and context-aware video understanding.