Convergence Rates of Training Data Reconstruction Attacks
Jessica Ingrey
Abstract
The ability to reconstruct a neural network’s training data from its parameters is a privacy concern. A family of attacks can provably reconstruct the entire training sets of models that are effectively linear in their feature representations, such as random features models or infinite-width networks in the neural tangent kernel regime. These attacks exploit the fact that the model parameters lie in the span of features induced by the training data, and reconstruct inputs that generate this span. Although experiments have demonstrated some transfer of these attacks to other models, they are not sufficient to determine the true threat posed. We therefore extend the theoretical results to deep neural networks of finite width $w$ by proving that the Wasserstein distance between the reconstruction and training sets is $O(\frac{1}{\sqrt{w}})$. Our experiments on two-hidden-layer binary classifiers are consistent with this result.
Chat is not available.
Successful Page Load