Why Transformers Struggle with Distribution-Independent In-Context Learning
Abstract
Transformer models can adapt to new tasks from input--output examples at inference time, a capability known as in-context learning (ICL). A common interpretation is that transformers implement implicit learning algorithms in their forward pass, using the context to infer the task and predict the query. We test this interpretation beyond the training distribution in controlled regression settings where classical estimators can generalize exactly. Transformers achieve near-perfect in-distribution ICL, but fail sharply under scale and support shifts, exhibiting a cliff-shaped collapse that is absent from least-squares and kernel baselines. This failure is robust across polynomial degrees and a range of training interventions. We identify two architectural mechanisms that help explain this behavior. First, final normalization imposes a readout-side scale constraint, causing predictions to saturate at fixed boundary values. Removing this constraint eliminates hard saturation, but does not restore reliable ICL under coefficient or label scaling. Second, attention can become a context bottleneck: under distribution shift, softmax attention may concentrate on a small part of the prompt, limiting the context-dependent adaptation required for ICL. These results show that strong in-distribution ICL does not imply distribution-independent ICL, and suggest caution when using ICL instead of explicit retraining or adaptation under distribution shift.