Diagnosing and Mitigating Modality Interference in Multimodal Large Language Models
Abstract
Multimodal Large Language Models demonstrate strong performance on multimodal benchmarks, yet remain fragile when one modality carries spurious or misleading content. We trace this fragility to a deficiency in cross-modality competency, defined as the ability to fairly evaluate and integrate information across modalities, and identify its concrete, measurable manifestation as \emph{Modality Interference}, in which task-irrelevant modality signals improperly influence the model's predictions. To diagnose this phenomenon systematically, we design a perturbation-based evaluation grounded in causal intervention, in which controlled noise is injected into the task-irrelevant modality across image-heavy and text-heavy tasks. Across diverse MLLM families and scales, we observe consistent performance degradation on modality-heavy tasks under perturbations, indicating that modality interference is a pervasive and scale-resistant failure mode rather than an artifact of any specific architecture. To mitigate this, we propose a unified perturbation-aware fine-tuning framework that combines (i) heuristic and adversarial data augmentation targeting the task-irrelevant modality, and (ii) output-level consistency regularization between clean and perturbed inputs. Extensive experiments across diverse MLLM architectures, model scales, and benchmarks show that our approach simultaneously improves unimodal robustness and standard multimodal performance, achieving Pareto-optimal gains over existing baselines.