Long-Horizon Agency Belongs in the Harness, Not the Context Window Only
Abstract
Long-horizon agency requires an AI system to preserve task state across hours or days of work, many tool calls, and repeated context resets. A common response is to scale a single resource: the model’s context window, together with inference-time reasoning. We argue that this response conflates long horizon with long context. Durable long-horizon state should not live only in the prompt; it should be externalized into the harness, the non-model infrastructure that manages tools, files, plans, checkpoints, sub-agents, permissions, and recovery. Publicly documented long-running agent systems already rely on such structures, including progress files, plan files, git history, rollback checkpoints, scratchpads, sub-agent decomposition, and browser-state checkpoints. We support this position with three lines of argument. First, memory-hierarchy history shows that pressure to scale one working resource repeatedly gave way to externally managed hierarchical state. Second, we identify five structural mismatches between long-context scaling and long-horizon agency: degraded effective context, lossy compaction, a single trust region, no transactional semantics, and no working-set decomposition. Third, we map long-horizon requirements to what context scaling provides and what harness-managed state must supply. We conclude by proposing eight harness state-management patterns and arguing that harness state should become a first-class research object for long-horizon agency.