Concentrated Gradients Amplify Forgetting: Dominant-direction Projection for Continual Multimodal Learning
Abstract
Multimodal Large Language Models (MLLMs) trained on sequential tasks suffer from catastrophic forgetting. We study this problem in LoRA-based multimodal continual instruction tuning. By analyzing the diagonal Fisher information of LoRA parameters, we find that forgetting is amplified not merely by shared high-sensitivity parameters across tasks, but by the concentration of updates on those shared parameters, \ie, some tasks distribute their change uniformly across parameters, while others pack most of it into a narrow subset, causing disproportionate interference on the parameters that prior tasks depend on. We further observe that this concentration manifests in the dominant directions of recent gradients. Based on this insight, we propose \textbf{D}om\textbf{i}nant-direction \textbf{G}radient \textbf{Pro}jection (\textbf{DiGPro}), which maintains a short buffer of recent gradient directions, extracts their dominant subspace with adaptive rank selection, and attenuates the parallel component for each optimizer step. Crucially, DiGPro requires no prior-task data or stored subspaces, operating on the current gradient history. Experiments on cross-modal and vision-centric continual learning benchmarks show DiGPro consistently reduces forgetting and attains competitive final results. The code is available in the supplementary materials.