Convergence of Near-Linear Width ReLU Networks with Unbalanced Initialization
Zhenghao Xu ⋅ Rachel Ward ⋅ Molei Tao ⋅ Tuo Zhao
Abstract
Linear-width convergence for two-layer ReLU networks is known in random-feature or conjugate kernel (CK) regimes whose analyses exploit nearly fixed hidden features. However, the corresponding neural tangent kernel (NTK) regime, where hidden-layer training is active, still suffers from substantially larger width requirements. The obstacle is the ReLU-induced kernel shift: activation-pattern changes can induce empirical NTK drift that breaks the contraction argument. We show that an output-dominant unbalanced Gaussian initialization creates an NTK-dominated training regime in which this kernel shift remains controllable. This single mechanism yields near-linear width, Nesterov acceleration, and low-rank adaptivity for two-layer ReLU networks with vector-valued outputs and a shared first layer. We prove that gradient descent (GD) achieves linear convergence for networks with $\tilde{\Omega}(Nn/\lambda)$ neurons, where $N$ is the sample size, $n$ is the output dimension, and $\lambda$ is the standard smallest eigenvalue of the limiting NTK. Within the same framework, Nesterov's accelerated gradient (NAG) attains a provable speedup without sacrificing near-linear width, improving the iteration complexity from $O(n\kappa\log\frac{1}{\epsilon})$ to $O(\sqrt{n\kappa}\log\frac{1}{\epsilon})$, where $\kappa$ is the limiting NTK condition number. Finally, our analysis establishes low-rank adaptivity: by introducing a sketching step at initialization and a subspace analysis, the width requirement reduces to $\tilde{\Omega}(Nr\kappa^2(\mathbf{Y})/\lambda)$ for responses $\mathbf{Y}$ of rank $r \ll n$, matching the ambient-dimensional result in the leading polynomial dependence, up to polylogarithmic factors, with $n$ replaced by $r$ when $\kappa(\mathbf{Y})=O(1)$. In contrast to CK/random-feature analyses that exploit nearly fixed hidden features, our proof controls hidden-layer NTK contributions, ReLU activation-pattern changes, and kernel shift; this control enables both acceleration and low-rank adaptivity in the NTK-dominated regime.
Chat is not available.
Successful Page Load