Continual Adaptation Across Modalities: Learning and Forgetting in Unified Multimodal Models
Abstract
Unified multimodal models (UMMs), which can process and generate diverse modalities within a single framework, have recently emerged as a powerful paradigm. However, when these models undergo continual learning on multiple tasks across different modalities, their learning and forgetting dynamics remain poorly understood. In this paper, we systematically investigate how architectural design and adaptation method shape intra-modal learning and cross-modal interference between the generation and understanding sides. Specifically, we report three findings. First, the severity of cross-modal interference decreases with the degree of modality-specific parameter isolation. Second, a within-model asymmetry: the generation side of a UMM is markedly more fragile than its understanding side. Third, what post-training adds is the first to be forgotten: purely post-training capabilities such as personalized generation collapse by 84\% to 90\%, while pretraining capabilities remain stable. Together, these results clarify how parameter isolation and post-training fragility shape UMMs' learning and forgetting dynamics, offering a starting point for developing continually evolving UMMs.