PRO-LONG: Programmatic Memory Enables Long-Horizon Reasoning
Alexis Fox ⋅ Junlin Wang ⋅ Paul Rosu ⋅ Bhuwan Dhingra
Abstract
Long-horizon tasks require sustained perception, reasoning, and exploration, and are a persistent challenge for large language model (LLM) agents. This gap is reflected in their limited performance on continual learning benchmarks such as ARC-AGI-3. Various agent harnesses have been proposed to close the gap, and each commits to a strategy for handling long sequences of observations, i.e., what information to save from the environment and how to load it into model context, a choice we argue is particularly consequential. Existing methods for context management have typically been designed around a tradeoff, where preserving more information makes relevant details harder to retrieve. Recent progress in coding agents, however, has made new approaches to memory viable. We propose PRO-LONG, a minimal context management framework built around programmatic memory for LLM agents in long-horizon, exploratory settings. PRO-LONG addresses the tradeoff by keeping a complete, structured interaction log and using programmatic tools to search this history efficiently. On the full ARC-AGI-3 public game set, PRO-LONG improves over a base coding agent by an average of $\mathbf{18.0}$ percentage points across frontier models, and matches or exceeds state-of-the-art specialized harnesses (up to $\mathbf{76.1}$\% pass@1) while using $\mathbf{4.2}$--$\mathbf{5.8}\times$ fewer tokens. With Fable 5, PRO-LONG achieves $\mathbf{97.4}\%$ best@2 at a total cost of \$1,750.
Chat is not available.
Successful Page Load