Temporal Prototype Alignment for Frozen-Feature Dataset Distillation
Abstract
Dataset distillation condenses large-scale training sets into compact synthetic data. Recent frozen-feature distillation methods, such as Linear Gradient Matching, make ImageNet-scale synthesis feasible by matching gradients in pretrained representation spaces, but their reliance on exhaustive spatial view ensembles incurs memory and I/O costs that grow linearly with the number of augmentations. We propose Temporal Prototype Alignment (TPA), an online temporal gradient estimation framework for frozen-feature dataset distillation. TPA replaces spatial Monte Carlo gradient averaging with a temporal prototype estimator that accumulates stable semantic signals across optimization steps. An adaptive momentum gate modulates the estimator according to local signal reliability, reducing gradient variance while keeping memory complexity independent of the number of spatial views. To improve cross-architecture transfer, we introduce a consensus alignment objective that distills shared structure from multiple pretrained teachers while mitigating model-specific artifacts. Across ImageNet-scale and fine-grained benchmarks, TPA substantially reduces memory overhead, enables distillation on consumer-grade GPUs, and improves transfer to heterogeneous CNN and Vision Transformer backbones. These results suggest that online temporal estimation provides an efficient alternative to spatial ensembling for frozen-feature dataset distillation.