High Performance Differentially Private Fine-Tuning using Dataset Distillation
Noel Loo ⋅ Hanshen Xiao ⋅ Alexander Amini ⋅ Mathias Lechner ⋅ Ramin Hasani ⋅ Daniela Rus
Abstract
Differentially Private Stochastic Gradient Descent (DP-SGD) is the dominant approach to private deep learning, but it pays compute and privacy budget \emph{per trained model}: each new architecture, ensemble member, or downstream fine-tuning round incurs additional privacy cost. We propose \textsc{SPS} (Summarize--Privatize--Synthesize) and its enhanced variant \textsc{SPS+}, dataset-distillation algorithms that release a private synthetic dataset by privatizing intermediate activation statistics from a public pretrained model. Once released, the dataset is post-processed freely: any downstream architecture, ensemble, federated aggregation, or continual update incurs \emph{zero} additional privacy cost. In particular, a single \textsc{SPS+} dataset distilled from a Wide ResNet-22-8 transfers zero-shot to architectures with different inductive biases (Vision Transformers, Swin Transformers, and ConvNeXt), all without re-incurring privacy. Empirically, \textsc{SPS+} is competitive with state-of-the-art DP-SGD across $\epsilon \in \{1,2,4,8\}$ on CIFAR-10 and CIFAR-100, with a notable advantage on CIFAR-100 under strict privacy, achieving a $5.8\%$ gain at $\epsilon{=}1$ over compute-matched DP-SGD. \textsc{SPS+} additionally outperforms prior generation-based DP methods by a large margin and supports private federated and continual learning out of the box.
Chat is not available.
Successful Page Load