Does Inference-Time Reasoning Really Improve Video Understanding?
Abstract
Video understanding models achieve impressive 70%+ accuracy on recent benchmarks, but do these truly measure genuine grounded understanding? We systematically investigate this by controlling frame sampling from text-only to real-time across major benchmarks (VideoMME, Video-MMMU, LVBench). Our analysis reveals that 13–46% of questions exploit shortcuts through world knowledge, language priors, or single-frame cues, allowing models to answer correctly without engaging visual evidence. Furthermore, requiring temporal grounding, both correct answers and time localization, exposes a ~40% accuracy gap, showing models frequently answer correctly without locating relevant evidence in videos. To address these limitations, we introduce TGVMME, a temporally grounded benchmark with 1.5k questions requiring both answer correctness and temporal localization. Our benchmark reduces shortcuts to 3.5%, validating that temporal grounding effectively ensures genuine video understanding. Using TGVMME to analyze state-of-the-art reasoning models, we make a surprising discovery: inference-time reasoning provides minimal gains (+2.5% at best), contrasting sharply with substantial gains on existing benchmarks. Analysis on accuracy changes of shortcut questions confirms reasoning primarily exploits shortcuts rather than enhancing genuine understanding. Further experiments reveal input scaling (increasing frames) provides more benefit than output scaling (reasoning) given sparse inputs, highlighting the importance of frame-efficient architectures. Our work demonstrates that grounded evaluation is essential for assessing video understanding capabilities and provides practical tools: validated shortcut taxonomy, filtered splits, and TGVMME for rigorous evaluation.