A Unified Theoretical Framework for Task Recognition and Task Learning in In-Context Learning
Abstract
In-context learning (ICL) allows large language models (LLMs) to adapt to new tasks from a few examples without parameter updates. Previous work and our empirical studies suggest two modes in ICL: Task Recognition, which applies pre-trained knowledge to familiar tasks, and Task Learning, which generalizes to novel mappings during inference. However, a unified theoretical account of these behaviors remains lacking. We propose a new Bayesian prediction framework that models pre-training data as a Pitman–Yor mixture over latent tasks, which captures the growing and heavy-tailed structure of natural language. By formulating an information-theoretic optimization problem, we derive the structure of the optimal ICL predictor under explicit capacity constraints. Our analysis reveals a phase transition in optimal capacity allocation: a water-filling strategy that prioritizes high-frequency tasks while assigning zero task-specific information to rare ones. This gives rise to two regimes: for frequent tasks, the model performs Task Recognition, yielding exponentially fast error decay with prompt length. For rare or unseen tasks, the model performs Task Learning by executing an implicit inference algorithm, resulting in a slower power-law convergence rate. Our results provide a principled explanation for the dual nature of ICL and establish a direct connection between pre-training data distributions, model capacity, and in-context generalization behavior.