Rethinking Expressivity and Efficiency in Test-Time Training
Zeyun Zhong ⋅ Joya Chen ⋅ Manuel Martin ⋅ Frederik Diederichs ⋅ Jürgen Gall ⋅ Jürgen Beyerer
Abstract
Test-Time Training (TTT) enables long-context processing via continuous weight updates during inference, but current methods struggle to balance the expressivity of per-token update dynamics with the hardware efficiency of chunk-wise approximations. We propose E$^2$-TTT (Expressive and Efficient TTT) to bridge this gap. By deriving a closed-form state transition that exactly aggregates per-token momentum and decay coefficients within a chunk, E$^2$-TTT enables fully parallelized chunk-level training while preserving the temporal structure of the update rule that prior chunk-wise methods discard. We validate E$^2$-TTT by training models up to 1.3B parameters from scratch. Extensive experiments demonstrate that our method consistently outperforms previous TTT and hybrid attention baselines in language modeling and retrieval, while achieving significantly better extrapolation on the standard ``Needle in a Haystack'' test, maintaining $>90$ % accuracy on passkey retrieval at $8\times$ the training context length. Meanwhile, E$^2$-TTT can match the training throughput of efficient chunk-wise methods, demonstrating that it effectively reconciles expressivity with efficiency.
Chat is not available.
Successful Page Load