Efficient Dataset Distillation for Pre-Trained Self-Supervised Models via Statistical Flow Matching
Qianxin Xia ⋅ JIAWEI DU ⋅ Yuhan Zhang ⋅ xin zhang ⋅ Xuewan He ⋅ Wenbo Jiang ⋅ Jielei Wang ⋅ Tao Luo ⋅ Guoming Lu
Abstract
Dataset distillation seeks to synthesize a highly compact surrogate dataset that achieves performance comparable to the original dataset on downstream tasks. For the scenario where pre-trained self-supervised models serve as priors, traditional $\textit{Linear Gradient Matching}$ optimizes synthetic images by encouraging them to mimic the gradient updates induced by real images on the linear probe or classifier. However, this batch-level formulation requires loading thousands of real images and applying multiple differentiable augmentations to synthetic images at each distillation step, leading to substantial computational and memory overheads. In this paper, we revisit the linear gradient and theoretically derive that it is essentially a $\textit{local}$ relative distribution directed from target class centers toward non-target class centers, which we term ``flow”. This property causes instability and suboptimality, often necessitating expensive multiple augmentations to compensate. To address this, we introduce $\textit{Statistical Flow Matching}$, a stable and efficient supervised learning framework that optimizes synthetic images by aligning $\textit{global}$ statistical flows in the original data. Our approach loads raw statistics only once and performs a single augmentation pass on the synthetic data, achieving performance comparable to or better than the state-of-the-art method with 10× less GPU memory usage and 4× faster distillation time. Moreover, increasing the number of augmentations for our method yields further performance gains while incurring lower additional cost.
Chat is not available.
Successful Page Load