The Denoising Wrapper: A Modular Post-Processing Framework for Noisy First-Order Optimizers
Meisam Razaviyayn ⋅ Weiwei Kong ⋅ Grigoris Velegkas ⋅ Vahab Mirrokni
Abstract
Noisy gradient estimates are ubiquitous in machine learning, arising from stochastic sampling, distributed computations, or privacy- preserving mechanisms like Differential Privacy (DP). This paper introduces a general, optimizer-agnostic framework for denoising these estimates. Our approach operates as a modular wrapper that intercepts noisy gradient observations and provides denoised estimates to the optimizer, requiring no internal modifications to algorithms like SGD or Adam. We start by studying the problem of optimal denoising gradients in quadratic and cubic optimization problems. We develop maximum likelihood estimation of gradient and Hessian of the objective. We also develop other computationally-efficient order optimal algorithms for denoising gradients in such a setting. Finally, we utilize the developments for the cubic and quadratic setting to develop a more general denoising mechanisms for general smooth nonconvex optimization problems. Theoretically, by leveraging higher-order smoothness, we establish an improved convergence rate of $O(T^{-12/19})$ for smooth non-convex optimization. While our recursive algorithm requires two gradient queries per iteration, we show that the improved convergence rate yields a lower total oracle complexity than the standard $O(T^{-1/2})$ rate of SGD. This implies the gain in convergence speed asymptotically outweighs the additional computational cost. We apply our algorithm to various denoising problems, particularly nonconvex DP-training. Our experiments, including training private classifiers on CIFAR-10, demonstrate significant improvements over baselines.
Chat is not available.
Successful Page Load