When Should Artificial Learners Narrow? Plasticity, Noise, and Developmental Curriculum Order
Abstract
Human perceptual systems are often broadly sensitive early in infancy and become more selectively tuned with experience, a developmental pattern known as perceptual narrowing. A tempting AI analogue is therefore to train broadly first and specialize later. We ask when that intuition is actually justified. In a controlled teacher-student model, each developmental stage supplies curvature along a different set of representational directions while the learner's effective plasticity changes over time. We derive an exact adjacent-swap identity: for commuting quadratic stages a,b, the difference between their two-step contraction factors in direction j is (etat-eta{t+1})(h_{b,j}-h_{a,j}). Thus curriculum order is irrelevant under constant plasticity, but consequential when plasticity declines; the useful early stage is the one that covers error-weighted directions before plasticity is lost, not necessarily the stage that is globally “easy” or “broad.” With additive gradient noise, an exact second-moment recurrence exposes a bias-variance tradeoff that can reverse the preferred order. In a 32-dimensional four-stage benchmark, broad-to-narrow is best among all 24 permutations for 100/100 random initializations in the noiseless decaying-plasticity condition, while constant plasticity makes the fixed orders numerically identical. A noise sweep produces a clear order crossover, and perturbing the stage spectra shows why fixed developmental slogans are brittle. These results turn perceptual narrowing into a falsifiable curriculum hypothesis for larger artificial learners: allocate early plasticity to unresolved representational coverage, then narrow only as the residual error itself narrows.