Sound Formula, Uncertain Facts: Solver-Directed Re-Verification for Policy Compliance
Abstract
Neuro-symbolic guardrails promise machine-checkable decisions by translating natural-language policies into solver constraints. Their guarantees are nevertheless conditional on the facts supplied at inference time. We study the setting in which a policy formula is established offline, while each case record is grounded online into variable bindings. An incorrect or incomplete case grounding can therefore change the verdict even when the solver reasons correctly. Existing redundancy checks compare independently generated groundings and flag disagreement. Yet disagreement is an imperfect proxy for verdict risk: the groundings may differ only on facts that cannot affect the verdict, whereas a shared omission can make them agree on the same wrong verdict. We introduce solver-directed re-verification, an agentic loop that uses solver-generated witnesses, candidate effects, and cores to seek the next decision-relevant uncertainty. Evidence sources attach checkable warrants to candidate values, inducing a set of evidence-admissible worlds. A verdict is certified only if the evidence rules out the opposite policy branch and confirms every decisive fact. Policy-blind subagents then gather warranted evidence from prose, logs, or registries. The solver updates the verdict and certificate after each query; unresolved cases return indet. This distinguishes decisions verified against evidence from those that have merely gone unchallenged. On an 80-case authored mortgage benchmark using Haiku 4.5, the Agent achieved 96.25\% accuracy, versus 77.5\% for a one-pass symbolic Pipeline and 72.5\% for a full-context LLM-as-a-Judge baseline. By directing evidence acquisition toward variables that could change the verdict or support its decisive proof core, the Agent opened 26.9 of 56 documents per case on average. These results provide a proof of concept that solver-directed re-verification can produce more accurate policy-compliance verdicts while reviewing fewer documents by targeting decision-relevant evidence.