Attributing Agent Learning: Beyond Post-Experience Accuracy
Abstract
A post-experience gain does not reveal whether an agent reused context or acquired persistent state. We present a two-link attribution method requiring evidence that experience produced an identified state change and that the change produced later behavior. The protocol combines exposure controls, blocked persistence, context isolation, restart, and rollback. Across 288 runs of Claude Sonnet 4.6 agents implemented with the Strands Agents SDK, informative histories produced the target skill in 72 of 72 runs versus 0 of 72 matched-sham runs. Informative near-transfer accuracy exceeded sham by 0.538 (95% episode-bootstrap interval: [0.510, 0.569]); restart preserved the gain and rollback removed it. Sham histories nevertheless produced 45 non-target commits. The evidence supports persistent skill acquisition within the declared agent system, not model-parameter learning or open-ended skill discovery.