Towards continual in-context learning via training cartridges at test-time
Abstract
Large language models can learn with high sample efficiency in-context, but in-context learning is naturally limited by context rot and the hard limit of the context window. Recent work such as Cartridges \cite{eyuboglu2025cartridges} and Attention Matching \cite{zweiger2026attentionmatching} show that a fixed context can be compressed greatly while preserving inference quality over it. Motivated by these methods, we study repeated context compaction with Cartridges at test time in the continual learning setting. First, we find that naive Cartridges fails to compose well with natural language instructions and generalize to longer appended contexts. We find distractor-augmented training, training with remaining window, and training additional residual MLPs greatly help solve this problem. We then find an ill-conditioned value fit in Attention Matching, stabilize it with ridge regularization, and use the resulting solution to initialize end-to-end SGD refinement which greatly improves accuracy and training stability. Building on these algorithmic and parametric optimizations, we develop a stable multiday self-study and compaction algorithm, which outperforms ICL on nine-day compaction of 1000 experiences in our synthetic cities task. On real-world tasks, we apply Cartridges to modern architectures (MoE, Gated DeltaNet hybrids, SWA) and scale self-study compute through synthetic data generation. On Continual Learning Bench (Blind Spectrum Monitoring task), we find multiday compaction with six cartridges using Qwen3-32B outperforms ICL, YaRN, recursive summarization, and match ICL of Opus 4.7 and Gemini 3.1 Pro. On the Database SQLite task, we find multiday compaction of Cartridges with Qwen 3.8 27B and reduces number of queries used while matching ICL accuracy. Our findings suggest that iterative context compaction can be a viable path to extend in-context learning.