When a Learning Curve Bends, Who Bent It? A Calibrated Test for Children and Language Models
Abstract
Developmental psychologists and machine-learning researchers have been arguing the same question without reading each other. In children, the vocabulary spurt was read as a qualitative reorganization until McMurray showed that parallel learning over variable difficulty produces acceleration with no change of mechanism. In language models, emergent abilities were read as capability discontinuities until Schaeffer and colleagues traced many of them to nonlinear metrics over smooth progress. Both fields are asking when a bend in an aggregate learning curve implies a change in the learner. We build one test for that question and apply it to both learners. We score a scan statistic over shared acceleration against a parametric bootstrap of a smooth micro-learning model fitted to each curve. The test holds size near nominal, detects almost nothing of a transition that ramps over months, and loses power from 0.73 on a uniform two-month design to 0.29 on the grids our cohort supplies. Among 326 Norwegian children with six or more longitudinal CDI administrations, 175 (54%) reject smooth aggregation and 164 survive false-discovery control, and among 18 Pythia curve families, 12 reject with 5 surviving correction. Neither result identifies what moved the curve, and on the child side we show that no analysis of these grids that controls size can. A shared per-visit shift in how much a parent ticks raises size from 0.055 to 0.58, combining five independent reports per visit leaves size at 0.546, and widening the null to carry the shift restores nominal size while dropping power against a factor-three break to 0.052. The only null that controls size against the shift has no power against the alternative. A sweep of 70 designs finds ten that reach power 0.80 at valid size, and the best observes from 20 to 36 months at two-month spacing with a wave on the transition. Spacing has an interior optimum, so sampling below two months makes the statistic worse. On the machine side seven of nine seed replicas at 160M depart on syntax, the ablation releases show data order unsettling the verdict where initialization doesn't, and dropping a fifth of the item bank reproduces the full-bank verdict for only 0.55 of departing families. Both designs need the measurement process replicated as well as the subject, and only the machine side can do that today.