X-AVDD: Cross-Attentive Audio-Visual Dataset Distillation
Saumyaranjan Mohanty ⋅ Aravind Reddy ⋅ Konda Reddy Mopuri
Abstract
*Audio-visual dataset distillation* aims to synthesize a small dataset from a large paired audio-visual dataset, while nearly preserving the training utility for neural networks. Existing approaches such as DM and AVDD face two critical limitations: 1) they are memory-intensive, making them difficult to scale to complex architectures and higher samples-per-class (SPC) settings, and 2) they often fail to preserve strong cross-modal semantic correspondence. To tackle these issues, we propose **X-AVDD**, a **cross-attention-based** distillation method that explicitly captures inter-modal relationships during synthesis. By decoupling cross-modal alignment from joint optimization, X-AVDD is substantially more memory-efficient, reducing end-to-end synthesis cost by $\sim 8\times$ GPU-hours, while producing highly aligned synthetic data. Empirically, X-AVDD significantly outperforms other audio-visual distillation methods. On VGGSound, an audio-visual dataset of short YouTube clips, X-AVDD improves top-1 test accuracy from $8.2$% → $24.0$% at SPC=$10$ and $9.8$% → $26.5$% at SPC=$20$. Furthermore, X-AVDD improves cross-architecture generalization (e.g., $25.5$% → $35.3$% for a ViT target on AVE dataset at SPC=$50$) and successfully scales to complex architectures such as AudioCLIP, achieving $62.3$% accuracy using only $5.65$% of the full dataset, against the full dataset accuracy of $72.4$\% while completing in $9\times$ lower training time compared to the full dataset.
Chat is not available.
Successful Page Load