From Squeezing to Grounding: Visual Guided DPO for Multimodal Hallucination Mitigation
Wenqi Liu ⋅ Hongxin He ⋅ Yunxiao Wang ⋅ Xuemeng Song ⋅ Yupeng Hu ⋅ Yinwei Wei
Abstract
Direct Preference Optimization (DPO) is widely used to reduce hallucination in Multimodal Large Language Models (MLLMs). Prior studies and our observations show that during DPO training, the log probabilities of chosen and rejected responses can decrease simultaneously, revealing the squeezing effect. For MLLMs, the redistributed probability mass may move toward responses driven by language priors rather than visual evidence, weakening the role of the image and amplifying hallucination. Existing squeezing-aware methods constrain rejected updates with $\mathcal{V}$-usable information or generation confidence, while methods with anchor terms regularize the DPO implicit reward of chosen responses. Neither explicitly grounds the control signal in visual evidence. We propose Visual Guided Direct Preference Optimization (VGDPO), which relaxes negative updates on rejected responses according to multimodal squeezing risk. VGDPO estimates this risk by combining token-level visual dependency with phrase-level hallucination localization, defining a hallucination ratio for each rejected response. We further apply a related reweighting principle to a visual contrastive objective, providing additional supervision for chosen responses under the original image. Experiments on four hallucination benchmarks and three MLLM backbones show that VGDPO improves DPO log probability dynamics and reduces multimodal hallucination while maintaining informative responses.
Chat is not available.
Successful Page Load