Improving Continual Video Instance Segmentation via Spatial-Temporal Balanced Mixture-of-Experts Adapters
Abstract
Effective Continual Video Instance Segmentation (CVIS) requires balancing catastrophic forgetting with the plasticity to learn new tasks. Existing methods typically adapt pre-trained VIS models using trainable visual prompts, but suffer from two key limitations: they overlook the distinct degradation of frame-level spatial knowledge (e.g., misclassification) and video-level temporal knowledge (e.g., trajectory disruption), and rely on rigid prompt retrieval that struggles under task ambiguity and distribution shifts. To address these issues, we propose STEAM, a CVIS framework with Spatial-Temporal Balanced Mixture-of-Experts Adapters. Instead of explicit retrieval, we introduce a Distribution-Aware spatial-temporal Scaling and sHifting (DASH) mechanism to align the pre-trained feature space with sequential tasks and reduce the domain gap. Furthermore, our method adaptively learns Task-Adaptive Spatial-Temporal mixture-of-Experts (TASTE) at both the frame and video levels, while updating them with Subspace-Constrained Orthogonal Adaptation (SCOA) to mitigate interference. Extensive experiments across diverse CVIS settings show that STEAM preserves temporal-spatial coherence and achieves state-of-the-art performance.