Self-Compacting Language Model Agents
Abstract
Long-form reasoning and tool calling traces accumulate errors and stale content that anchor subsequent generations, a phenomenon known as \emph{context rot}. Existing scaffolds mitigate this with \textit{fixed-interval} compaction triggered at a token threshold, but such triggers are blind to trajectory structure and risk discarding partial results mid-derivation or mid-search. We propose SelfCompact, which pairs two inference-time elements: an inline \emph{compaction tool} the model invokes itself, and a lightweight \emph{rubric} specifying when to fire (a sub-task has resolved, or the trajectory is converging) and when to suppress (mid-derivation, or when stuck). Both are needed. The tool alone is unevenly used across open-weight models, often invoked at unhelpful moments or not at all; the rubric alone cannot act. Together, they elicit effective adaptive compaction without any fine-tuning. On competition math (IMO-Answerbench, HMMT Nov 25 / Feb 26) with four Qwen3 / Qwen3.5 models and agentic search (BrowseComp, BrowseComp-Plus) with three deployed agents, \method{} matches or exceeds fixed-interval summarization at a fraction of the token cost, improving over a no-summarization baseline by up to 16.7 points on math and 5--9 points on agentic search at 30--70\% lower per-question cost. The result reframes \emph{when to compact} as a meta-cognitive capability that scaffolds, not weights, can supply.