Dual-Pronged LoRA: Achieving Near-Zero Forgetting and High-Performance Adaptation for MLLMs
Abstract
Adapting Multimodal Large Language Models (MLLMs) to specialized and diverse downstream applications necessitates fine-tuning, typically through parameter-efficient methods such as Low-Rank Adaptation (LoRA). However, fine-tuning MLLMs severely degrades pre-trained upstream capabilities, a phenomenon known as catastrophic forgetting. Existing Methods primarily rely on parameter pruning to strike a balance between the downstream adaptation and upstream knowledge retention. Empirical analysis indicates that these methods are subject to parameter space entanglement, wherein the imposition of downstream adaptation and upstream retention within a unified parameter space results in suboptimal outcomes for both objectives. To address this, we introduce Dual-Pronged LoRA, a novel fine-tuning framework for MLLMs featuring two components. (1) Task-Decoupled Contextual Routing (TDCR) makes the first attempt to explicitly decouple tasks by modeling a downstream contextual support region. It adaptively activates the LoRA branch exclusively for downstream-relevant inputs and bypasses it otherwise, thereby preventing downstream adaptation from interfering with pre-trained knowledge. (2) Downstream-Performance-Oriented Rank Sculpting (DPO-RS) maximizes downstream adaptation performance by dynamically sculpting ranks via gradient energy. It effectively strengthens critical layers while compressing redundant ones and is theoretically grounded in the minimization of information loss. Extensive experimental results indicate that Dual-Pronged LoRA achieves near-zero forgetting while attaining new state-of-the-art downstream performance.