3DSPA: A 3D Semantic Point Autoencoder for Evaluating Video Realism
Abstract
AI video generation is evolving rapidly. For video generators to be useful for applications ranging from robotics to film-making, they must consistently produce realistic videos. However, evaluating the realism of generated videos remains a largely manual process -- requiring human annotation or bespoke evaluation datasets which have restricted scope. Here we develop an automated evaluation framework for video realism which captures both semantics and coherent 3D structure and which does not require access to a reference video. Our method, 3DSPA, is a 3D Semantic Point Autoencoder utoencoder which is trained to reconstruct held out 3D point trajectories which have been augmented with DINO semantic features. By combining both semantic and geometric representations, 3DSPA enables robust assessments of realism, temporal consistency, and physical plausibility in generated videos. Experiments show that 3DSPA outperforms leading VLMs by more than 10\% in identifying videos which violate physical laws, and more than doubles the correlations between human and auto-eval ratings in judging physical common sense and motion quality across multiple generative video datasets. Our results demonstrate that enriching trajectory-based representations with 3D semantics offers a strong foundation for benchmarking generative video models, and implicitly captures physical rule violations.