PhysVista: Benchmarking Physical Intelligence in VLMs via a Perception-Reasoning-Assessment Loop
Abstract
Vision-Language Models (VLMs) have demonstrated impressive multimodal reasoning, yet their ability to truly internalize the underlying physical consistency of real-world dynamics remains an open question. Existing benchmark paradigms often suffer from fragmented evaluation, focusing on isolated cognitive stages while overlooking the inherent synergy between perception, reasoning, and physical judgment. The lack of a holistic perspective limits the ability to diagnose whether VLMs can reliably evaluate the physical authenticity of emerging generative models. To address these issues, we introduce \textbf{PhysVista}, a benchmark designed to evaluate physical intelligence in VLMs through a closed cognitive loop framework inspired by the human \textbf{seeing–reasoning–assessment} process. PhysVista restores this loop by jointly evaluating physical state perception, physical dynamics reasoning, and physical plausibility assessment. It further distinguishes \textbf{event-level} reasoning and \textbf{scale-level} reasoning to enable fine-grained analysis of physical understanding. In addition, PhysVista incorporates both real-world and AI-generated videos, allowing evaluation across diverse domains and emerging generative scenarios. Extensive experiments on a wide range of state-of-the-art VLMs reveal significant limitations in current models’ physical intelligence, particularly in fine-grained reasoning and plausibility assessment. Our findings highlight critical gaps between visual recognition and genuine physical understanding, and provide insights for developing future physically grounded VLMs.