TFDBench: A Comprehensive Benchmark and Codebase for Talking-Face Deepfake Detection
Abstract
Evaluating talking-face deepfake detectors remains challenging because existing studies rely on different training datasets, detector-specific preprocessing pipelines, and inconsistent evaluation protocols. To resolve this empirical gap, we introduce TFDBench, a comprehensive benchmark and unified codebase for multimodal talking-face deepfake detection. TFDBench integrates six datasets spanning lip-sync, face-swap, single-image talking-head generation, and text-to-speech-based manipulations, and implements 14 state-of-the-art visual and audio-visual detectors within a shared preprocessing, training, and metric computation. It supports in-domain, cross-domain, specific-generator, multilingual, class-imbalance, visual-degradation, and audio-rate-mismatch evaluations. Our results reveal several important limitations of current detectors: strong in-domain performance often fails to generalize under cross-domain, multilingual, and post-processing shifts. Furthermore, while semantic audiovisual alignment exhibits the strongest zero-shot transferability, it remains uniquely susceptible to advanced lip-synchronization optimizations. Ultimately,TFDBench establishes a highly reproducible and rigorous foundation essential for analyzing and advancing robust deepfake defenses.