Minimum Stability with Noise-Injected Inputs
Yu-Han Wu ⋅ Pierre Marion
Abstract
Many learning objectives, such as denoising score matching (DSM) in diffusion, corrupt the inputs of the network with fresh noise at every optimization step. On finite training sets, such noisy-input risks admit a unique global minimizer (for instance, the empirical optimal score for DSM). In this paper, we show that SGD with a large learning rate provably cannot stably converge to this global minimizer. For two-layer ReLU networks with nonzero input weights trained by unconstrained SGD, we prove that DSM with distinct training inputs has no mean-square stable affine dynamics around any stationary point. Therefore, the downstream analyses are conducted with inner weights either fixed or trained withspherical SGD. Extending the mean-square stability analysis of SGD to noisy-input risks, we prove that every linearly stable minimum of SGD for two-layer ReLU networks of width $m$ has a mean squared gradient norm at most $2d/(pm\eta)$, where $\eta$ is the learning rate and $p\approx1/B$. In consequence, we show that stability also implicitly controls the variation norm of the function represented by the network. For DSM, this yields an excess risk lower bound of order $d/\sigma^2$ whenever $p\eta\gtrsim\sigma$ and noise level is small, showing that large learning rates prevent memorization.
Chat is not available.
Successful Page Load