Similar Gains, Different Stability: Layer-Dependent Interference in Continual Multimodal Adaptation
Abstract
Parameter-efficient adaptation of multimodal LLMs is typically evaluated by target-task gain, while the location of adaptation within the decoder receives less attention. Using Qwen2.5-VL-3B-Instruct, we apply equal-budget LoRA updates to different three-layer decoder windows and evaluate a recognition target alongside TextVQA, counting, and referring-expression grounding as collateral capabilities, repeating the analysis with counting as the target. Windows with closely matched target gains can produce substantially different collateral effects, and this discrepancy sharpens under sequential adaptation. In the main six-window sweep, two windows with only a 0.35-pp gap in first-stage recognition gain differ by 9.70 pp in subsequent backward transfer (evaluation-sample bootstrap 95% CI [7.70, 11.70]). Second-stage reseeding shows the forward-order gap is trajectory-sensitive, while the reverse-order disadvantage persists across three reseeds for both tested pairs. A window that looks risky after one-shot evaluation can remain stable under continual adaptation, while one that looks safe can forget substantially—showing that target gain and one-shot collateral behavior can be insufficient proxies for sequential stability, and motivating capability-aware layer selection for continual multimodal adaptation.