Learning through Internalization
Abstract
We study internalization, the process by which neural-network-based systems absorb an explicit computational procedure into their own weights, and how it facilitates learning. We investigate how transformers internalize the simulation of semiautomata by internalizing chain-of-thought (CoT) tokens, which classes of semiautomata are harder to internalize, and expose the flip side of internalization, that is, a progressive degradation of out-of-distribution performance. We then provide the first provable analysis of successful internalization: for the task of learning parities, we show that a simplified one-layer transformer provably first learns the target with explicit CoT supervision and then internalizes the autoregressive generation as CoT tokens are progressively removed, directly computing the parity. Learning such a representation directly from data without CoT supervision is computationally hard. Finally, we discuss how learning through internalization can be viewed as an instance of the \textit{Positive Distribution Shift} phenomenon recently introduced by~\citet{Med+26}.