Seeing or Rationalizing? Scene-Evidence-Guided Chain-of-Thought for Faithful 3D Multimodal Reasoning
Abstract
3D multimodal language models are becoming a foundation for embodied agents, indoor robotics, and human–AI interaction, but their answers and explanations often remain weakly tied to the specific objects and spatial relations observed in a 3D scene. This paper aims to make 3D reasoning more faithful by reducing language-prior rationalization and forcing reasoning chains to depend on verifiable scene evidence. We propose Anti-Rationalization Chain-of-Thought (AR-CoT), a plug-in training and decoding framework that represents each reasoning step with an explicit evidence pointer and scores candidate answer-chain pairs by their scene-evidence gain. AR-CoT contrasts scene-conditioned rationales with a frozen text-only prior, verifies local object and relation claims, and uses decoy-twin scene pairs to encourage reasoning chains to change when the relevant spatial evidence changes. Experiments on standard 3D QA and grounding benchmarks, as well as MSQA, Beacon3D, and shortcut-sensitive stress tests, show that AR-CoT consistently improves multiple 3D MLLM backbones while strengthening chain grounding, decoy-twin contrast, and scene-swap sensitivity. These results suggest that anti-rationalization offers a practical and verifiable route toward more accurate, interpretable, and scene-faithful 3D multimodal reasoning.