Towards continual in-context learning via training Cartridges at test time
Hongyi Huang ⋅ Kaitlyn Wang ⋅ Ajay Anubolu ⋅ Shayan Talaei ⋅ Arshia Soltani Moakhar ⋅ Azalia Mirhoseini ⋅ Amin Saberi
Abstract
Agents deployed in a large project (e.g., over a large database or codebase) can learn from its own experiences with high sample efficiency from in-context learning alone, but ICL is naturally limited by context length and eventual context rot. Recently, Cartridges (\cite{eyuboglu2025cartridges}) and Attention Matching (\cite{zweiger2026attentionmatching}) show that a fixed context can be compressed $\geq 50\times$ into a small KV cache without significant loss in downstream accuracy. Motivated by these methods, we study whether repeated compression of an agent's experiences acting test time can be achieved through iterative Cartridge compaction. We propose algorithmic and parametric optimizations that enable our stable multiday Cartridge training algorithm. First, we augment Cartridges training with distractor contexts that enable generalization to long-horizon knowledge acquisition. Second, we initialize Cartridges training from a ridge-regularized Attention Matching cache, which speeds up training and improves downstream performance over the original Cartridges recipe. The resulting self-study-and-compaction loop matches or exceeds in-context learning on our synthetic knowledge-retrieval task and on two Continual Learning Bench tasks. On Blind Spectrum Monitoring, Qwen3-32B with iterative compaction beats its own ICL baselines and matches the ICL accuracy of Opus 4.7 and Gemini 3.1 Pro; on Database SQLite, Qwen3.8-27B matches ICL accuracy with fewer queries. This present this study and results across three settings as a research prototype, which has ubiquitous use cases across user chats and over enterprise databases.
Chat is not available.
Successful Page Load