Agent-ToM: Learning to Monitor Autonomous LLM Agents via Theory-of-Mind Reasoning
Abstract
Monitoring autonomous large language model (LLM) agents for covert malicious behavior (e.g., covertly pursuing a hidden malicious objective) is challenging due to delayed, context-dependent, and long-horizon attack patterns. In adversarial settings such as sabotage, agents may pursue hidden objectives while maintaining superficially benign behavior, making detection difficult even with full trajectory access. Prior monitoring approaches primarily improve monitor scaffolding or ensemble aggregation, but treat each trajectory independently and do not improve from prior monitoring experience. Moreover, standard reasoning methods explain observed behavior but do not explicitly reason about agent beliefs, intentions, and goal alignment required to distinguish benign task execution from covert deviation. We propose Agent-ToM, a learning-to-monitor framework grounded in Theory-of-Mind (ToM) reasoning for security analysis of autonomous agents. Agent-ToM performs structured full-trajectory analysis by inferring beliefs, step-level intent hypotheses with calibrated confidence, expected actions, and deviations from task-consistent behavioral baselines. At inference time, it employs a Reason--Verify--Refine pipeline to construct and validate monitoring decisions. At training time, Agent-ToM learns from prior monitoring episodes by distilling critique signals into a persistent semantic guardrail memory that accumulates monitoring strategies, enabling reusable belief- and intent-conditioned constraints to be applied across episodes. We evaluate Agent-ToM on adversarial agent monitoring benchmarks (SHADE-Arena and CUA-SHADE-Arena). Agent-ToM achieves strong precision--recall balance and outperforms state-of-the-art monitoring baselines, including ensemble methods, while using a single coherent reasoning pipeline. These results demonstrate that \emph{learning at the monitoring layer}, combined with structured ToM reasoning and verification, provides an effective and deployable foundation for securing autonomous LLM agents.