Video Forensic Self-Descriptions: Leveraging Temporally Distributed Forensic Microstructures for Zero-Shot Detection of AI-Generated Videos
Abstract
AI-generated videos are becoming increasingly realistic, yet most detection methods focus on spatial artifacts within individual frames and require synthetic training data from known generators, causing them to fail on unseen generators. We identify two previously unexplored temporally distributed forensic microstructures that reveal whether a video was AI-generated. These traces arise from how a video's visual field and motion evolve over time, capturing subtle deviations that generators introduce when approximating real-world temporal dynamics. To expose them, we train two simple predictors on real video only, modeling visual evolution (VE-FSD) and motion evolution (ME-FSD), and analyze their prediction residuals. We fit 3D spatio-temporal autoregressive models to these residuals, producing compact descriptors we call Video Forensic Self-Descriptions (VFSD). VFSDs of real videos naturally cluster together and separate from AI-generated ones, enabling zero-shot detection and source attribution without synthetic training data. Experiments across multiple datasets and 35 generators show VFSD achieves state-of-the-art performance on both tasks.