Threshold Renormalization Explains Weak-to-Strong Generalization in Diagonal Linear Networks
Abstract
Weak-to-strong generalization (W2SG) asks whether a model trained on predictions from a weaker model can outperform it. We study this question in a two-stage diagonal linear network, using its exact equivalence to LASSO regression. A first-stage estimator is trained on noisy ground-truth labels; a second-stage estimator is then trained on independent inputs labeled only by the first stage. Under a power-law signal model and a beyond-proportional LASSO state-evolution characterization, we show that the second stage is asymptotically equivalent, up to a vanishing perturbation, to applying an additional soft threshold to the first stage’s output; the resulting scalar estimator therefore has a strictly larger effective cutoff. This renormalized threshold can improve population prediction risk by discarding noise-dominated coordinates retained by the first stage, but can also shrink genuine signal, and cannot recover a coordinate the first stage already removed. We characterize this tradeoff by deriving a sharp phase diagram, in terms of the two stages' sample-size scaling exponents and the signal decay rate, that determines when the second stage improves prediction risk away from the critical boundaries. The phase diagram identifies a critical Stage-II sample-size scaling above which W2SG occurs when the Stage-I sample-size scaling remains below the scale set by the signal decay rate; once the Stage-II exponent exceeds that of Stage I, further samples preserve the asymptotic improvement but make its magnitude vanish.