Stable Averages, Unstable Items: Individual Differences in Machine Language Development
Abstract
No two children acquire language on the same schedule, and developmental psychology is built around that variation, with norms, percentiles, stability coefficients, and screening rules as its instruments. We apply those instruments to language models, using the PolyPythias seed replicas and the original Pythia run, ten at each of 14M, 31M, 70M, 160M, and 410M parameters, evaluated at 22 checkpoints on a battery of 696 CDI words, twelve BLiMP subtasks, and a nonce-morphology test. Aggregate accuracy is reproducible across seeds, as prior work reports, and the items that compose it are not. Measured on one benchmark at both granularities, with identical items and runs and corrected for the sampling noise of a binary outcome, seed spread at item level exceeds aggregate spread by 3.8x at 14M and 13.7x at 410M on the 1020 items every run at every scale acquires. The correction is conservative, which makes these lower bounds, and the gap replicates on a second battery read from continuous surprisal. Per-word timing varies across runs at a coefficient of variation of 0.26 falling to 0.20, against 0.16 among children, and machines appear to vary more than children at every scale measured. That pooled figure holds in four of six regimes we resolve and reverses in two, the function words and the words the models learn most decisively, which both fall below the child figure. Decoupled arms at 160M show initialization and data order contributing equally (0.110 against 0.116) and interacting rather than adding, by an excess of 11% that holds in all 120 subsets. Early rank predicts final rank at no checkpoint before 56% of training at 70M or 77% at 14M, and at no checkpoint at all at 31M, 160M, or 410M, and no single-checkpoint rank-threshold screening rule follows. We release MachineCDI, per-word developmental norms that carry seed error bars, together with the evaluation cache that regenerates every number here.