Fast and slow gradient descent dynamics of logistic regression through weak alignment
Han Bao
Abstract
Gradient descent has been of particular interest in modern machine learning beyond sole focus on optimization because some specific structures emerge through the optimization dynamics. Such behaviors result from optimization even though the learning objective does not explicitly encode the target structure, collectively called implicit bias, often preventing overparametrized models from fitting to spurious patterns. A typical instance is the max-margin implicit bias of a linear classifier, widely established for exponentially tailed loss functions. Even after having a given dataset separated, the parameter vector continues to evolve towards the max-margin direction asymptotically along the gradient descent dynamics. This phenomenon corroborates a frequent empirical observation of ``train longer, generalize better.'' However, the max-margin convergence is an asymptotic phenomenon, and what is worse, this asymptotic convergence rate is significantly slower than convergence in optimization. Even so, the parameter vector along gradient descent dynamics commonly correlates with the max-margin direction positively (though not exactly) within considerably fewer iterations than the asymptotic rate. By shedding another light on this classical yet profound problem, this work aims to understand the mechanism of this early-stage alignment phenomenon. Our theoretical results demonstrate that the parameter vector weakly aligns with the max-margin direction within $O(\exp(\exp(-\delta)))$ iterations, where $\delta>0$ is the permissible alignment error, which is shown to be tight. By tracking the radial and tangential flows, our proof operates on the alignment dynamics directly with dataset geometry and gets rid of the asymptotic expansion, which is a key insight to enabling faster weak alignment.
Chat is not available.
Successful Page Load