GEOPHYS: The Geometry of Physical Plausibility
Abstract
Whether a video is physically plausible is a question current generative models answer poorly and current evaluators answer expensively, through physics-targeted training, billion-parameter video world models, or multimodal-language model judges. We ask whether frozen images vision encoders, with no video or physics supervision, already contain a useful signal for the same question. Across four frozen backbones, self-supervised vision transformers (DINOv2, DINOv3) and biologically-inspired models of the ventral stream (CORnet-S, VOneNet), a small set of geometric signals (curvature, speed variation, acceleration, prediction residual) on per-frame feature trajectories separates plausible from physics violated videos with no retraining. We propose GEOPHYS, a training-free framework that 11 reaches 97.7% on LikePhys and 93.3% on IntPhys2 accuracy for physics-violation detection, surpassing V-JEPA 2, GPT-4o, Gemini, and twelve modern video diffusion models. The same signals track human EEG responses to object-permanence violations and scale with object number. Deployed unchanged as a best-of-N reward during video generation, GEOPHYS lifts MAGI-1 4.5B from 53.8\% to 64.7\% PhysicsIQ score at 5.8× lower wall-clock and 7.8× lower memory than V-JEPA 2 world-model rewards. A useful proxy for physical plausibility is already encoded as a geometric regularity of natural-video statistics in vision features.