Spectral Rank Calibration for Continual LoRA Merging in Multimodal Large Language Models
Abstract
Model merging enables the integration of multiple expert models with different capabilities into a unified model. In practical deployments, new expert models are often continuously updated and arriving. Meanwhile, due to the training and communication overhead of large language models (LLMs) and multimodal large language models (MLLMs), Low-Rank Adaptation (LoRA) is widely used to adapt large models to specific domains due to its efficiency. This scenario brings a new challenge: continual LoRA merging of deployed MLLMs to load new task capabilities. Naive merging faces catastrophic forgetting and parameter conflicts. We first observe the distinct features of the LoRA direction and layers in MLLM LoRA tuning. Based on this, we propose our Spectral Rank Calibration Merging (SRC-Merging), a data-free and train-free continual LoRA merging method tailored for MLLMs. SRC-Merging regards the deployed LoRA as a compressed spectral memory and performs calibrated rank-level merging for each incoming adapter. It aims to preserve accumulated multimodal knowledge while avoiding over-suppression of new vision-language task updates during continual merging. Extensive experiments on multiple multimodal tasks have demonstrated that our method significantly outperforms state-of-the-art merging methods.