A Boundary-Layer Mechanism for One-Third Scaling in Online Softmax Classification
Marcel Kühn ⋅ Yoon Thelge ⋅ Bernd Rosenow
Abstract
Hard-label classification is usually trained with smooth surrogate losses, most prominently softmax cross-entropy. We isolate an asymptotic mechanism by which this mismatch between smooth surrogate and discrete labels produces power-law learning curves in an online teacher-student model. After subtracting the mean logit, the thermodynamic-limit dynamics close in centered variables: a growing centered student-teacher alignment $D$ and the residual student variance $\Delta$. At late times only layers of width $O(D^{-1})$ around teacher decision boundaries stay active, while the noise of fixed-learning-rate online gradient descent maintains a nonzero $\Delta$. As a function of the training time $\alpha$ the late-time solution yields a $\alpha^{-1/3}$ power law not only for the test loss but also for the generalization error $\epsilon_g$, i.e., one minus test accuracy. Learning-rate schedules can improve the error exponent towards $1/2$. Simulations support the predicted order parameter dynamics and learning curves.
Chat is not available.
Successful Page Load