Data Diversity Drives the Emergence of Symbolic Mechanisms Supporting Abstract Reasoning
Abstract
Recent work has identified a set of emergent symbolic mechanisms that support abstract reasoning in large language models, but it remains unclear what factors drive the emergence of these mechanisms. To address this question, we trained neural networks from scratch on abstract sequence tasks, varying both architectural and data distributional factors, and investigated the effects of these factors on the emergence of symbolic mechanisms. Using a combination of representational, attentional, and casual mediation analyses, we first confirmed that transformer language models trained from scratch on our task developed symbolic mechanisms, partially capturing the set of mechanisms learned by large-scale pretrained models. We then investigated the effect of data diversity, operationalized as vocabulary size, on the emergence of these mechanisms, finding that mechanistic and behavioral signatures of emergent symbol processing scaled with data diversity, including downstream measures of systematic (i.e., out-of-distribution) generalization. Surprisingly, we found that the emergence of symbolic mechanisms was not driven by architectural inductive biases, as the same mechanistic and behavioral signatures emerged in modified transformer architectures and even multilayer perceptrons. These results suggest that data diversity, rather than architectural inductive biases, is the primary driver of emergent symbolic computation in neural networks.