Asynchronous Agentic Poisoning
Abstract
LLM agents are increasingly studied as systems that can self-evolve during deployment. A representative mechanism for realizing self-evolution is memory evolution, where agents maintain evolving memories derived from previous tasks and condition future inference on their memory. Although memory evolution is intended to improve agent performance, it also creates a feedback loop in which model-generated content can persist across tasks and influence future decisions. We study this loop as a delayed attack channel for memory-evolving agents. We introduce asynchronous agentic poisoning, a provider-side threat in which a malicious provider releases a model that appears safe under fresh-context evaluation but exhibits unsafe behavior after self-evolving through memory evolution. We realize this threat by fine-tuning the released model to silently embed a hidden marker into its otherwise normal responses, causing the marker to accumulate in the agent’s memory through memory evolution. When the accumulated markers are later retrieved into the input context, the model reacts to them as a backdoor trigger and shifts to attacker-specified behavior. This creates a self-triggered transition from safety-aligned behavior to attacker-specified behavior without requiring any post-deployment attacker interaction. Experimental results show that our method increases the harmful rate from below 5\% to around 90\% after memory evolution, while not substantially degrading the model’s self-evolution performance. These findings demonstrate that memory evolution creates a provider-side pathway for delayed compromise, allowing a model to appear safe at release time while becoming unsafe after deployment through the agent’s self-evolution process.