Nonparametric In-Context Learning under Growing Geometric Complexity: Minimax Optimality and Local Geometry-Adaptivity of Transformers
Jaehee Seo ⋅ Jisu Kim
Abstract
Transformers have become a central architecture for \emph{in-context learning}, particularly through their strong empirical performance in large language models. This success suggests that transformers can extract task-relevant structure directly from prompts, even when the underlying data have complex geometric structure. However, existing theoretical analyses of transformer-based in-context learning are largely confined to simplified settings, such as Euclidean domains or single-manifold models. In this paper, we study in-context nonparametric regression under growing geometric complexity, modeled by a sample-size-dependent mixture of manifolds with heterogeneous local geometry. For this model class, we establish a minimax lower bound that captures the aggregate difficulty of its local components. We then construct an oracle tangent local-polynomial estimator and prove a matching upper bound by exploiting local geometric structure. The main technical step is to connect this estimator to a transformer architecture. We construct a two-stage linear-attention transformer consisting of a geometric preconditioner and all-chart reduced local-polynomial solvers, and show that its approximation error is negligible relative to the minimax regression rate. We also prove an in-context generalization bound for $\varepsilon$-near empirical risk minimizers within this transformer class. Together, the lower bound, oracle construction, transformer approximation, and generalization analysis identify the conditions under which the resulting in-context predictor adapts to local geometry and attains the aggregate minimax rate.
Chat is not available.
Successful Page Load