What Compression Preserves: Instruction Survival in Long-Horizon LLM Agents
Daniel Kirste ⋅ Alexander Kropiunig
Abstract
Production harnesses for long-horizon LLM agents contain a compaction operator: when the history outgrows the context window, the harness rewrites it and continues. The operator is treated as behaviour-preserving scaffolding, yet nothing bounds which standing requirements survive the rewrite. We measure this on 1,175 trajectories across $\tau$-bench retail and AppWorld, crossing three trigger thresholds, two compaction methods (truncation and LLM summarization), and three reminder frequencies against a no-compaction baseline, with every planted requirement scored by a mechanical checker. The failure mode is structurally predictable: requirements whose only cue lives in an early, compactable user message (an wrapper, a [DONE] marker) fall from 82%/87% without compaction to 22%/32% across compacted cells, while requirements the trajectory keeps re-cueing through tool traffic, such as rounding and snake_case keys, hold at 89–100%. The loss coexists with task success, so it is invisible to the signal harnesses monitor. Compaction also lowers task success by 24.7–33.6 percentage points, and its cost is protocol-dependent: uncached full-schema billing makes every compacted $\tau$-bench cell 13–42% more expensive in input tokens, whereas cached-prefix billing removes truncation's excess but not summarization's (19–35%). One mid-trajectory restatement raises $\tau$-bench adherence by 5.6 percentage points (95% CI [+0.9, +10.1]); further restatements add nothing. A target/cooldown sensitivity sweep and a Llama-3.3-70B replication reproduce the split. We release the harness, constraint catalogue, and all trajectories.
Chat is not available.
Successful Page Load