One Clock, Three Learners: A Conserved Developmental Schedule in Mammals, Children, and Language Models
Abstract
Comparative neurobiology holds one quantitative law of developmental timing, which says that the order of neurodevelopmental events stays fixed across mammals while a single species rate parameter absorbs almost everything else. We know of no fit of that law to data outside neural tissue, and outside the tissue neither a genome nor a maturational clock can carry the explanation. We fit one estimator, unchanged, to 372 dated neurodevelopmental cells across ten mammals, to 5,120 word acquisition ages across fifteen languages from Wordbank, and to 234 acquisition times that we measured ourselves by scoring 67 BLiMP paradigms on six Pythia scales at seventeen checkpoints each. All three systems obey the law against a permutation null built to destroy the shared ordering, and their correlations come in at 0.989, 0.807, and 0.882, which stand 19.2, 46.0, and 10.2 standard deviations clear of that null. The variance then splits in opposite directions, because the species rate on its own carries an R^2 of 0.788 in mammals, while the language rate carries 0.078 and the model scale rate carries 0.039, rising to 0.233 when we restrict the machine table to the paradigms retained at every size. Biological development runs on a clock, while in both learning systems the order carries the variance and the clock carries under a tenth of it on the full tables. We then held each unit out in turn, anchored its rate on three of its own events, and used the shared schedule to place a mammal it never saw at a median relative error of 8.1% against the observed day. The mammalian rates reproduce the phylogeny in order, which we never fitted for. The primate cortical interaction that we did include returns cortical events 1.27 times late, with a bootstrap interval clear of one rather than the zero the data were free to give it. The child schedule orders words as nouns, then verbs, then adjectives, then function words, which is the canonical sequence of early word learning. Model scale turns out to be a clock that decelerates, since the rate falls from 5.73 at 14M to 1.20 at 160M and then moves by under a quarter out to 1B, where 160M and 410M sit in a near tie. Scale compresses the schedule as well, because a stretch term fitted for each size falls from 1.79 at 14M to 0.51 at 1B while buying under 0.01 of variance in either biological table. The machine result also survives a change of training corpus, since the deduplicated Pythia ladder reproduces both the law and the event ordering, with the two fitted schedules correlating at a Spearman of 0.960.