CausalConflictBench: Can Multimodal Models Follow Local Mechanisms That Conflict with Commonsense?
Abstract
Large multimodal models often benefit from commonsense priors, but these priors can mislead reasoning when a task specifies a local mechanism that conflicts with real-world regularities. Existing scientific VQA and visual reasoning benchmarks tend to align problem mechanisms with commonsense, so a correct answer may reflect either mechanism following or prior-based answering. We introduce Causal-ConflictBench, a diagnostic benchmark that makes this ambiguity observable by constructing samples where the current cause-to-effect mechanism contradicts the default commonsense mechanism. CausalConflictBench contains two modules: Textual Rule Override (TRO), which provides explicit commonsense-conflicting textual rules in real science image questions, and Visual Counter-Commonsense Induction (VCI), which requires models to induce a conflicting mechanism from a three-frame visual sequence. Beyond sample-level accuracy, we report Group-Strict Accuracy, factual-prior fallback metrics, and output-level process diagnostics. Across 14 proprietary and open-source multimodal models, we find that high sample-level accuracy can mask unstable mechanism following; errors concentrate strongly on factual-prior answers; and failure modes differ across rule-delivery paths, with TRO revealing rule-application failures and VCI revealing visual rule-induction failures. CausalConflictBench therefore provides a controlled diagnostic setting for uncovering commonsense fallback hidden beneath aggregate accuracy.