Towards continual in-context learning via training Cartridges at test time
Hongyi Huang ⋅ Kaitlyn Wang ⋅ Ajay Anubolu ⋅ Shayan Talaei ⋅ Arshia Soltani Moakhar ⋅ Azalia Mirhoseini ⋅ Amin Saberi
Abstract
Agents deployed over long-horizon tasks can learn from their own experiences with high sample efficiency from in-context learning alone, but ICL is naturally limited by context length and eventual context rot. Cartridges (\cite{eyuboglu2025cartridges}) and Attention Matching (\cite{zweiger2026attentionmatching}) show that a fixed context can be compressed $\geq 50\times$ into a small KV cache without significant loss in downstream accuracy. These methods only compact the context once, but an agent must iteratively compact newly acquired experiences. Therefore, we propose two new changes that enable iterative KV cache compaction at test-time. First, we augment cartridges training with distractor contexts that enable generalization to long-horizon knowledge acquisition. Second, we initialize cartridges training from a ridge-regularized Attention Matching cache, which speeds up training and improves downstream performance over the original Cartridges recipe. The resulting self-study-and-compaction loop matches or exceeds in-context learning on a synthetic knowledge-retrieval task and on two Continual Learning Bench tasks. On Blind Spectrum Monitoring, Qwen3-32B with iterative compaction beats its own ICL baselines and matches the ICL accuracy of Opus 4.7 and Gemini 3.1 Pro; on Database SQLite, Qwen3.8-27B matches ICL accuracy with fewer queries. We present this study and results across three settings as a research prototype, which has ubiquitous potential use cases across user chats or enterprise databases. Finally, we discuss limitations and applications of our approach in relation to broader challenges of continual learning.
Chat is not available.
Successful Page Load