The Multiscale Single-Index Model: A Toy Model for Hierarchical Feature Learning
Gordon Dai ⋅ Joan Bruna
Abstract
We introduce the Multiscale Single-Index Model, a stylized model for deep hierarchical feature learning with scale separation. Each layer extracts a shared single-index feature at one physical scale and passes it to the next, giving a tractable setting in which to study how deep architectures learn local multiscale representations. Under non-degeneracy and delocalization assumptions on the link function and planted features respectively, we prove two complementary recovery guarantees. First, for any fixed depth $K$ and local scale $d$, corresponding to an input of size $d^K$, the first Wiener chaos expansion of the target behaves as a perturbed spiked tensor, where the perturbation comes from the non-linearity; we then leverage this fact to show that a spectral method based on tensor unfolding strongly recovers all planted directions with $O(d^{\lceil K/2\rceil}\log d)$ samples. Second, for the two hidden-layer case, we analyze joint online spherical SGD and show that it achieves weak recovery from random initialization in $O(d\log^2 d)$ samples, followed by strong recovery to accuracy $\varepsilon$ in $O(d\log(1/\varepsilon))$ additional samples. The main technical challenge is that the second layer observes non-Gaussian learned features; we overcome this through a quantitative Gaussian comparison using delocalization, combined with sharp trajectory-level control of the coupled SGD dynamics.
Chat is not available.
Successful Page Load