Adaptive Compression and Targeted Perturbation: A Unified Framework for Generalized Audio Deepfake Detection
Abstract
Existing audio deepfake detection (ADD) methods frequently struggle with limited generalization, as fixed bottleneck architectures fail to adapt to the heterogeneous information densities inherent in diverse spoofing artifacts. Moreover, standard augmentation often incurs a pathological robustness-accuracy trade-off. In this paper, we propose a unified ADD framework to address these challenges. Our approach integrates: (1) an Adaptive Audio Feature Integration (AAFI) module that dynamically selects optimal latent dimensions via a compression pool and a stochastic exploration mechanism, ensuring flexible representations across multi-type dynamic attacks; and (2) a Mel-Band Adversarial Perturbation (Mel-BAP) strategy that applies targeted regularization in the mel-spectrogram domain to encourage the learning of invariant decision boundaries. Extensive evaluations on benchmarks including ADD 2022, FakeOrReal, and In-the-Wild demonstrate that our Whisper-small based model achieves a state-of-the-art average Equal Error Rate (EER) of 1.51\%. Compared to existing SOTA models such as the 3-billion-parameter Resemble-Detect-3B-Omni, our framework reduces parameter count by up to 91.6\% while improving average EER by 60.7\%. Notably, the model exhibits exceptional zero-shot transferability, achieving 3.91--5.65\% EER on unseen languages and maintaining high resilience against real-world acoustic distortions. This work provides a scalable and efficient “plug-in” solution for enhancing trust and security in the audio Web ecosystem.