Optimal In-Context Learning of Autoregressive Processes under Heterogeneous Second-Order Moments of the Prompts
Abstract
In-context learning enables large language models to adapt to tasks directly from input sequences, without parameter updates. We investigate the mechanism underlying in-context learning of autoregressive processes in transformers under heterogeneous second-order moments of prompt distributions. Our study considers a two-layer architecture consisting of a linear attention head and a nonlinear MLP, with the outer layer trained via gradient descent. We prove that there exists a parametrization under which the model converges to the unique function that approximates the optimal second-order moment-based AR(1) estimator. We analyze bounds on approximation and convergence rates and confirm our findings experimentally. These results extend the first-moment-centric understanding of estimators realized by in-context learning to second-order moment-based estimators, offering further insight into ICL's robustness to new tasks.