Reflections of Human Mathematical Cognition in NLP Corpora and Large Language Models
Abstract
We investigate the proposal that mathematical concepts are largely learned through experience in the symbolic environment. This proposal has previously been investigated using relatively small corpora. Here we use the Pile, an NLP corpus of 800GB. We explore its statistical structure and the correlation between this structure and the mathematical performance of the Pythia family of models trained on this corpus. We first find that frequency of numbers and arithmetic problems is a power function of their size, the same function that has been documented in cognitive science studies. We next look for more direct evidence that this statistical structure impacts what is learned, finding that Pythia-12B-deduped's performance on mathematical tasks mirrors this power function and thus also matches human performance. Finally, we show that the internal representation of numbers and arithmetic problems in Pythia-12B-deduped is similar to the compressed mental number line and memory fact network documented for humans. Our findings support the proposal that numerical and arithmetic abilities and their representations can be learned from the statistics of the symbolic environment. This study illustrates the potential of large NLP corpora for studies of mathematical cognition.