When Language Overrules: Revealing Text Dominance in Multimodal Large Language Models
Huyu Wu ⋅ Meng Tang ⋅ Xinhan Zheng ⋅ Duo Su ⋅ Haiyun Jiang
Abstract
Text dominance, the tendency of multimodal large language models (MLLMs) to over-attend to textual tokens while under-utilizing non-text inputs, has been observed in vision-language settings, but whether it extends to other modalities remains an open question. We present a systematic, attention-level study of this phenomenon across five modalities (image, video, audio, time-series, and graph), using two diagnostic metrics: the Modality Dominance Index (MDI), which compares per-token attention between text and non-text inputs, and the Attention Efficiency Index (AEI), which normalizes attention share by token share. Across ten models and six benchmarks, text dominance is pervasive in image, video, audio, and time-series settings and intensifies in deeper layers. Controlled token-replication experiments show that expanding non-text sequences without adding semantic content amplifies the imbalance, while graph tasks provide a boundary case where compact, information-dense tokens reverse the effect. Guided by these findings, we evaluate attention-based token compression as a proof-of-concept intervention on the vision modality. On LLaVA-1.5-7B with MMMU-Pro, removing 90\% of visual tokens via informed selection preserves accuracy (33.70\% vs.\ 33.23\% baseline), while random dropping under the same ratio degrades it to 30.29\%; informed selection also lowers the late-layer MDI from 15.63 to 3.46 and reduces latency by 3.5$\times$. These results suggest that text dominance is a cross-modal phenomenon associated with token redundancy, and that token compression can reduce it in the tested vision setting while preserving task performance.
Chat is not available.
Successful Page Load