Where You Backpropagate Matters: Token Hypothesis for Memory-Efficient Fine-Tuning
Abstract
Fine-tuning large transformers routes gradients through every token position, paying the cost in activation memory that scales linearly with sequence length and dominates the GPU budget. Existing efficient methods reduce the parameter side---low-rank adapters, prefix tuning, weight-decomposed updates---but leave the activation footprint untouched. We present the Token Hypothesis, a memory-efficient fine-tuning procedure motivated by a Neural Tangent Kernel analysis of how fine-tuning signal is distributed across token positions. We prove that every position is a low-dimensional capacity bottleneck whose width does not grow with the model, that positions with aligned hidden-state geometry are interchangeable in bidirectional models, and that causal positions form a monotone capacity chain in which later positions strictly dominate earlier ones. The analysis pins down how many positions are needed and which ones, both computable from forward passes alone: any subset suffices in bidirectional models, while the last few positions suffice in causal ones. We translate this into a procedure that backpropagates through only a small subset of positions while running the full forward pass over the rest. Empirically, the procedure matches or exceeds full-sequence fine-tuning across six task families---commonsense reasoning, mathematical reasoning, visual question answering, video question answering, GLUE, and image classification---and five backbones spanning bidirectional and causal architectures. It reduces activation memory substantially, translating to up to ∼ 20% lower total GPU memory (∼ 33GB) on long-context tasks, and composes cleanly with parameter-efficient methods such as LoRA and DoRA.