Tracing the Thoughts of a Coding Agent Playing ARC-AGI-3: Lessons for Continual Learning
Abstract
We study how a coding agent learns across a sequence of abstract reasoning tasks. The agent runs on a frozen foundation model inside a fixed harness and acts by writing and running Python and shell scripts. The agent retains no state across turns other than its written artifacts, so every thought it forms, carries, corrects or abandons leaves a trace, where a thought is any belief, rule or plan committed to a file. We let the agent play ARC-AGI-3, a set of interactive reasoning games that provide no instructions. Each game is a sequence of levels, and a strategy that clears one level can fail on the next, so every new level is in effect a new task. The agent records what it learns about each game as Python scripts and text notes, while the harness keeps a complete log of every action and observation. Our contribution is a measurement protocol that traces each thought through these files, from the task where it forms to the task where it is corrected or abandoned, applied to seven evaluation runs with three backbones from two model families. We find that scripts written for one task are almost never called again in a later task (33 of 630 references cross a task boundary), because most scripts embed the state of the current level and become invalid when the level changes. Instead, the agent rewrites its knowledge into new scripts, retyping most of each version while keeping the general rules and dropping the level-specific details, and abandons 74\% of the scripts it wrote before a boundary. The notes, which only the model reads, are never revised. The agent appends to them without removing earlier claims, and the contradictions that accumulate are settled against the log. Because the log preserves everything, the agent can discard scripts freely and reconstruct their content when needed, so it forgets selectively, not catastrophically. The most costly error is a hard-coded value carried into a task where it no longer holds. The fix the agent found was to turn the constant into a parameter that must be supplied afresh for each new task. The above findings were obtained from the files the agent wrote, without access to the model, and constitute a white-box analysis of how a coding agent continually learns.