ANCHOR: Audio-Visually Grounded Chain-of-Thought Reasoning Benchmark
Abstract
Multimodal large language models (MLLMs) achieve impressive performance on audio-visual question answering, but correct answers do not necessarily imply perceptually grounded reasoning. Models may exploit linguistic priors or dataset shortcuts without attending to the visual regions and audio events that actually justify their predictions. Existing benchmarks predominantly evaluate final answer accuracy, providing limited insight into whether models truly “see” or “hear” the evidence required for reasoning. We introduce ANCHOR, a benchmark and evaluation framework for auditing perceptual grounding in multimodal reasoning. ANCHOR requires models to produce grounded chain-of-thought explanations that explicitly reference audio-visual evidence, including representative frames, spatial bounding boxes, and salient audio cues. To build a reliable gold standard, we combine automated annotation generation with a human-in-the-loop refinement pipeline, producing 224 human-verified videos and 975 question–answer pairs with aligned reasoning and localization annotations from 2.3K videos and 8.2K QA pairs. We further propose a metric and evaluation protocol that jointly measure answer correctness and reasoning-grounding alignment, complemented by multi-judge scoring of hallucination, logical divergence, final answer, and a token-based conciseness measurement. Experiments on nine state-of-the-art MLLMs reveal a substantial faithfulness gap: models with strong answer performance often produce explanations that are only weakly grounded in the underlying perceptual evidence. Ablation results confirm the importance of explicit grounding cues: removing bounding boxes reduces accuracy by 7.7 points, and performance drops from 57.7% with full cues to 16.1% in the audio-only setting. These findings show that answer-only evaluation is insufficient for multimodal reasoning and position ANCHOR as a new testbed for developing systems that are accurate, interpretable, and trustworthy.