A Self-Evolving Framework for Efficient Terminal Agents via Observational Context Compression
Abstract
Terminal observations are not ordinary long-context text: they are heterogeneous, low-information-density execution traces in which sparse but exact evidence (e.g., error messages and file paths) is interleaved with large amounts of repetitive terminal output. For long-horizon CLI agents, retaining raw observations rapidly increases context cost and can dilute critical signals, while LLM-based summarization or fixed heuristics often fail to adapt across heterogeneous terminal tasks and may discard precise task-relevant evidence.We propose TACO, the first self-evolving Terminal Agent Compression framework, which treats compression rules as reusable, preservation-aware knowledge acquired from interaction trajectories. Rather than relying on manually designed rules, static pruning, or task-specific compressor training, TACO autonomously discovers, refines, and reuses structured compression rules from agent interaction trajectories. A Global Rule Pool accumulates effective rules across tasks, while task-time rule evolution adapts them online to the current workflow. This allows TACO to filter redundant terminal observations while conservatively preserving exact task-relevant evidence.Across six benchmarks, including TB 1.0, TB 2.0, SWE-Bench Lite, CompileBench, DevEval, and CRUST-Bench, TACO consistently maintains or improves task success across models and agent scaffolds. On TerminalBench, TACO yields 1%--4% absolute accuracy gains under standard evaluation and improves accuracy by 2%--3% under matched token budgets. On downstream benchmarks, it reduces total token consumption by 12%--27% while maintaining or improving task success. These results show that self-evolving observation compression can unlock latent capability in existing CLI agents by allocating context budget toward task-relevant evidence, without model fine-tuning or human-crafted compression rules.