In-context classification: Calibration and noise in linear attention with softmax readout
Arman Rysmakhanov ⋅ Mary Letey ⋅ Yue Lu ⋅ Jacob Zavatone-Veth ⋅ Cengiz Pehlevan
Abstract
We study in-context binary classification in a reduced linear-attention model trained to minimize a logistic loss. Each prompt has a latent logistic teacher, and the learner must infer its direction from Bernoulli-labelled context examples. We derive a cavity-based asymptotic prediction for the test logit in a high-dimensional regime with context length $\ell \propto d$ and total training prompts $n \propto d^2$. Our formula separates calibration error, context noise, and a non-isotropic parameter-estimation noise that vanishes in the population limit. Our result further formulates the ridgeless separability regime, the effect of regularisation, and the effect of sampling temperature, all of which merit further study.
Chat is not available.
Successful Page Load