What Does Multiple Choice Actually Measure? Disentangling Video Understanding from Option-Exploitation
Abstract
Video understanding evaluations are intended to measure a model’s proficiency of video perception and task solving. However, current VideoQA benchmarks still struggle to accurately represent these capabilities due to a variety of reasons including a reliance on multiple choice answer guidance. Disentangling these confounding factors helps us identify which parts of a model’s score are supplied by the evaluation context and construction, rather than earned by faithful perception. Across six video-QA benchmarks of varying length, we find that a large portion of accuracy survives with the video removed, caused by short-cutting and option-prior reliance, and that on one benchmark it survives the removal of the question as well. We recast these effects in the language of faithfulness. Option-exploitation and answer-conditioned reasoning traces are unfaithful even when they are correct. Ultimately, we show a VideoQA score is jointly determined by language priors, option-timing, and how outputs and labels are converted into correctness. We produce a novel Option-delayed Replay (ODR) mechanism to get objective MC-like scores from more faithful OE inference. ODR demonstrates that agentic methods with ODR & faithful inference can reach option-exploiting performance, thanks to their accumulated evidence and reasoning. We package the diagnostics and ODR as a transferable verification and evaluation toolkit to use before trusting a VLM in a real-world video setting. We also release a relabeled long video dataset to mitigate the effects of data errors, an otherwise confounding variable.