AMPS: Adaptive Modality Preference Steering via Functional Entropy
Zihan Huang ⋅ Xintong Li ⋅ Rohan Surana ⋅ Tong Yu ⋅ Rui Wang ⋅ Julian McAuley ⋅ Jingbo Shang ⋅ Junda Wu
Abstract
Multimodal Large Language Models (MLLMs) often exhibit significant modality preference, which is a tendency to favor one modality over another. Depending on the input, they may over-rely on linguistic priors relative to visual evidence, or conversely over-attend to visually salient cues over textual facts. Prior work has applied a uniform steering intensity to adjust the modality preference of MLLMs. However, strong steering can impair standard inference and increase error rates, whereas weak steering is often ineffective. Since steering sensitivity varies substantially across instances, a single global strength is difficult to calibrate. To address this, we introduce an instance-aware diagnostic metric Modality Contribution Ratio (MCR) that quantifies each modality’s information contribution and reveals sample-specific susceptibility to steering. Building on this signal, we propose a scaling strategy and a learnable module for instance-aware control of modality preference. Experiments show that AMPS outperforms conventional steering on $MC^2$ by improving preference control with lower generation collapse, and improves MLLM performance on general multimodal QA dataset such as MM-Vet v2.
Chat is not available.
Successful Page Load