Diagnosing and Correcting Bias in MLLM for Long Video Understanding
Xusheng Liang ⋅ Jianqiao Sun ⋅ Hao Zhang ⋅ Yulei Niu ⋅ Hengshuang Zhao ⋅ Jiawei Ma
Abstract
Long-video question answering is challenging since the answer often hinges on a few decisive moments scattered throughout the video, while memory constraints force Multimodal Large Language Models (MLLMs) to sample frames at extremely low rates, resulting in extreme sparsity that risks omitting key evidence. Though aggressively enlarging the context size is intuitive, we argue that performance remains fundamentally constrained by the bias, $\textit{i.e.}$, the misleading cues caused by spurious linguistic correlations and salient yet irrelevant visual observations. In this paper, we first introduce S-MME, a benchmark designed to systematically diagnose shortcut-biased behavior in MLLMs. To mitigate this failure, we further propose a training-free framework, Counterfactual Long-video Evidence-Aware Reasoning (CLEAR). By treating the full video-question pair as a complete view, CLEAR estimates bias effects via other views by only keeping the question or locally salient visual clips. Then, we apply hidden-state intervention to mitigate the discrepancy at inference such that the correct answer can be predicted confidently. With experimental studies across multiple benchmarks, we show that our method is generalizable and can be integrated in different MLLMs, yielding consistent gains in prediction and robustness. We hope CLEAR can offer a complementary direction in test-time scaling for long-video understanding and will publish code.
Chat is not available.
Successful Page Load