The Prestige: Benchmarking Cognitive Visual Reasoning using Magic Tricks
Abstract
With the rapid development of Vision-Language Models (VLMs), it is increasingly critical to evaluate their ability in accurate and robust reasoning based solely on visual information. However, existing video understanding benchmarks often overlook the deep cognitive processes inherent in visual reasoning and frequently rely on multiple-choice formats, implicitly encouraging reasoning shortcuts. To address these limitations, we introduce \emph{The Prestige}, a novel video understanding benchmark designed to isolate VLM reasoning under deliberate cognitive misdirection through professional magic performances. The dataset consists of a curated collection of high-quality magic videos paired with a diverse distribution of strictly open-ended questions targeting causal state tracking, paradoxical reasoning, and resistance to adversarial questions. These questions need the models to perform long-horizon, multi-step commonsense reasoning with cognition. Through extensive evaluations of leading open-weight and proprietary VLMs, using a combination of machine and human grading, we find a substantial gap between model and human performance. \emph{The Prestige} establishes a rigorous new frontier for evaluating temporal, causal, and cognitive-flow reasoning in multimodal foundation models.