Assessing the Gameability of Vision Language Model Judges for World Model Video Evaluation
Abstract
Vision-language model (VLM) judges are increasingly used to evaluate the physical plausibility of world model videos and, in some settings, as optimization signals for video generators. We stress test three specialized judges: WorldModelBench VILA, VideoPhy-2-AutoEval, and PhyJudge-9B using paired clean and perturbed videos designed to test sensitivity to physics-relevant temporal corruption and invariance to physics-preserving superficial cues. Temporal effects are evaluated against blinded human judgments, while superficial perturbations are compared against a frozen V-JEPA representation with a human-supervised ordinal plausibility probe. In a blinded human audit, temporal shuffle, freeze, and reversal reduced Physical Commonsense scores by 1.53--3.62 points on a 1--5 scale, while all three judges decreased substantially less. Conversely, every text overlay family increased scores for at least two judges despite negligible changes in the reference plausibility score. These perturbations can also alter generator rankings. Together, our results show that specialized world model video judges can underreact to physically meaningful temporal corruption while remaining sensitive to superficial visual cues, raising concerns about their reliability for both benchmarking and optimization.