MixForensics: Blend Before You Encode for Generalizable AI-Generated Video Detection
Abstract
Detecting AI-generated videos (AIGV) in a generator-agnostic manner is increasingly important as generative models close the gap with real footage, and a video, with many frames, intuitively offers richer forensic evidence than a single image. Yet existing detectors fall short of this expectation: under the encode-then-fuse paradigm, increasing the frame budget from 1 to 8 improves a voting-based image detector by only 5.28%, and even temporal-modeling backbones such as TimeSformer barely improve on this. Per-frame predictions are also noisy: per-frame logits fluctuate substantially within a video and correlate weakly with any frame-level proxy. We trace the cause not to the aggregation step but to per-frame encoding severing multi-frame cues before they can interact, leaving any post-hoc selection, weighting, or aggregation strategy without a reliable basis. We therefore propose MixForensics, a blend-then-encode framework that combines multiple frames in the pixel domain into a single composite image (Stochastic Frame Blending, SFB) and regularizes the encoder for consistency across blended views (Blend Invariance Regularization, BIR). To stress-test this approach, we further introduce MixForensics-Bench, a test-only benchmark covering ten recent generators (including five closed-source commercial platforms) and two partial-forgery scenarios (temporal splicing and conditional continuation), with matched controls that isolate generative artifacts from editing artifacts. On AIGVDBench and MixForensics-Bench, MixForensics outperforms the strongest baseline by 1.27% and 8.11% on the fully-fake splits, leads on both partial-forgery scenarios by 11.49% on temporal splicing and 9.15% on conditional continuation, and uses fewer encoder forward passes than per-frame methods.