Proof-Gated Skill Compilation: Continual Learning for LLM Agents Without Trusting the Model
Abstract
LLM agents can learn continually by accumulating reusable skills, but every such system must decide what to persist, and most trust the model’s own report that a solution worked. We show this trust is misplaced: across 560 escalations to frontier “teacher” models from four vendors—closed and open weights—each prompted to verify its solution before returning, the teachers asserted verification 538 times and independent execution confirmed 52; more than 89% of every observed teacher’s affirmative verification claims were false under independent execution. We present proof-gated skill compilation, a continual-learning layer in which solutions persist only as executable programs a harness has independently executed against ground truth, and are served only when execution of the stored program reproduces the incoming instance’s own demonstrations. The gate rejected all 486 fabrications; none persisted. On a 90-task schema-recurrent benchmark derived from ARC-AGI-2 (∼4,900 ledgered trials), the library scores 98.9% on a disjoint-seed validation sequence against a 54.4% matched no-learning baseline, serving 78 of 90 trials by verified program execution at 100% accuracy, no model call, and 18 ms mean latency, with the escalation loop closing within a single mining pass. An exhaustive vocabulary-only control—the same twelve primitives and all 156 singleton and ordered-pair compositions under the same binding and demonstration checks, with no library, no model, and no memory—reaches 73.3%, and all 23 discordant tasks favor the learned library, ruling out search over the pre-existing vocabulary alone as the explanation. We argue that verification—at admission and again at serving—rather than storage or retrieval is the binding constraint on continual learning for agents, and we report the matched-baseline, harness-effect, memory-attribution, vocabulary-only, and fabrication-rate controls that any such claim should carry.