Don't Prompt, Graph It: Reliable SOP Compliance via Workflow Graphs
Abstract
Standard Operating Procedures (SOPs) govern critical enterprise decisions, yet LLM-based approaches that directly prompt models with SOP documents suffer from rule confusion, hallucination, and inconsistent outputs, failure modes that are unacceptable in compliance-critical settings. We propose a different paradigm: don't prompt, graph it. Instead of including raw SOP text in every inference call, our method, AutoSOP, compiles the SOP offline into a structured workflow graph and executes it online through a rule engine harness: a deterministic rule engine that serves as the structural backbone for an LLM agent. The harness owns all procedural control, which step runs next and which branch is taken, while the agent contributes language understanding, interpreting the SOP offline and gathering step inputs online. This separation makes execution deterministic, fully traceable to the source document, and immune to prompt injection. On two public benchmarks, AutoSOP improves overall accuracy on RuleArena by up to 30 points (81.3% vs. 61.0% for the best baseline) and raises FlowBench session success by 28.7 and 20.0 points in single- and cross-scenario settings, while giving the reproducibility and auditability that enterprise applications require.