Logarithmic Depth Suffices for In-Context Gradient Descent
Abstract
Transformers can perform in-context learning at depths far smaller than the number of optimizer updates suggested by existing gradient-descent interpretations. Prior work shows that an L-layer Transformer can emulate L gradient-descent updates, but this leaves open whether depth is merely a step counter. We show that it is not: in linear self-attention, each layer compounds the algebraic capacity built by previous layers, yielding a tight Θ(log k) depth characterization for computing the final result of k gradient-descent steps on diagonal quadratic tasks. Controlled experiments validate this mechanism: trained linear-attention models exhibit the predicted logarithmic depth transition, and kernel and oracle comparisons identify polynomial degree as the governing bottleneck. Broader experiments show the same depth signal under standard Transformer components, broader ICL regression families, and early-exit probes on Qwen2.5.