Evaluating and Mitigating Hallucinations in Audio-Visual Multimodal LLMs with Spoken Queries
Abstract
Hallucinations in vision-language models are typically evaluated with textual queries, even though emerging multimodal AI assistants increasingly accept speech as user input. In this paper, we investigate whether changing only the query modality, from text to semantically equivalent speech, affects visual hallucination. To isolate this effect, we develop a controlled data-generation pipeline that converts existing hallucination benchmarks into spoken-query counterparts while preserving the original images, query semantics, tasks, and labels. We instantiate the pipeline on RePOPE and CHAIR to construct RePOPE-Spk and CHAIR-Spk, which cover binary and open-ended hallucination evaluation, respectively. Across both proprietary and open-source multimodal LLMs, spoken queries consistently increase visual hallucination even under clean speech. Adverse acoustic conditions further amplify the degradation. We further examine whether common training-free mitigation strategies transfer to this setting. While speech-aware prompting and contrastive decoding (CD) with acoustic negatives reduce hallucination, CD with visual negatives provides little benefit. Based on these findings, we propose multi-negative CD that combines partial-silence and Gaussian-noise negatives, reducing CHAIR-S below the text-only baseline (16.4% vs. 18.0%). Overall, our study establishes spoken-query robustness as a new evaluation dimension for vision-language hallucination in voice-enabled multimodal AI systems.