GEAR-Align: Grounding-Evidence-Aware Gradient Routing for Multimodal Alignment
Yu Yongkang ⋅ Haobo Wang ⋅ Meng Chen ⋅ Han Fang ⋅ Xin Wei ⋅ Zhiyu Lin ⋅ Ye Yuan ⋅ Chiming Duan ⋅ Hao Sun ⋅ Haiyang Zhang
Abstract
Multimodal Large Language Models (MLLMs) are typically aligned with caption-pair data, but not all answer-token supervision is equally visual. In LLaVA-style alignment training, many tokens are predictable from language context alone, yet their losses can still update vision-related parameters. Such language-dominant updates may add noisy supervision for tokens that require visual evidence. To diagnose this vulnerability, we probe the model against visual counterfactuals to reveal the intrinsic grounding requirements of each token, distilling them into two soft indicators: visual necessity and evidence specificity. Guided by this diagnostic insight, we present $\textbf{GEAR-Align}$, a controlled post-training framework that makes supervision destination explicit. During training, it converts token-level visual reliance into parameter-level gradient allocation: all answer tokens update language-side parameters, high-necessity tokens update route-flow parameters, and the high-specificity subset additionally updates visual-evidence parameters. Across two controlled Qwen3-backbone settings, GEAR-Align improves hallucination-sensitive and local-evidence-heavy benchmarks while remaining competitive on general multimodal benchmarks. These results provide evidence for token-specific visual gradients as a useful principle for controlled MLLM post-training.
Chat is not available.
Successful Page Load