Agentic Memory Engineering: Learning Programmatic Memory for Web Agents
Abstract
Modern agentic systems typically store past experiences as unstructured textual traces, forcing the model to rely on implicit in-context reasoning rather than explicit control mechanisms. Under this paradigm, memory serves only as a passive reference that must be re-interpreted during every inference step. This lack of formal structure prevents the agent from directly applying learned logic to new decisions, often leading to redundant processing and inconsistent behavior in complex environments. To address these limitations, we introduce Agentic Memory Evolution (AME), a framework that transforms raw interaction history into executable symbolic logic. AME shifts the role of memory from a static text repository to an active component that directly intervenes in the action-selection process. By converting past successes and failures into modular rules, the system can filter out invalid candidates and prioritize actions with high predicted utility. These symbolic representations are not static; they are continuously refined through environmental feedback, allowing the agent's decision logic to evolve and improve as it gains more experience. We evaluate AME on the challenging WebArena benchmark, where it establishes a new state-of-the-art success rate of 68.7%. Beyond its performance on familiar domains, AME also shows promising cross-site generalization results on unseen websites. Experimental results show that by representing memory as executable logic rather than simple text logs, agents can accumulate reusable interaction patterns that significantly outperform traditional trajectory-based memorization.