Soteria: Formally Verified Planning with Runtime Enforcement for Safe LLM Agents
Deyuan (Mike) He ⋅ Ankush Desai ⋅ Sharad Malik ⋅ Aarti Gupta
Abstract
LLM agents that execute multi-step tool calls must satisfy two objectives simultaneously: completing the user's task (utility) and conforming to domain policies and correctness constraints (safety). These objectives are in tension -- blocking an unsafe action prevents a violation but can leave the agent stranded with no principled way to recover. Existing approaches sacrifice one objective for the other. We introduce Soteria, a framework that reconciles safety and utility through verified hierarchical planning and runtime enforcement of specifications. Before execution, the agent generates a structured plan that is formally verified against the specifications prior to any tool invocation. During execution, the verified plan serves dual roles: it ensures that the agent takes specification-conformant trajectories and provides guidance when unsafe actions are blocked. Across multiple benchmarks covering a diverse range of tool-use tasks, Soteria achieves perfect specification conformance while improving utility by up to $5\times$ over existing guardrail approaches, demonstrating that safety and utility need not be traded off.
Chat is not available.
Successful Page Load