Continual Evaluation for Enterprise Agent Governance: Architecture and Evidence from a Policy Layer
Abstract
Enterprise LLM agents in regulated sectors are increasingly deployed behind a governance layer: a service that inspects every prompt, tool call, and response and returns allow, transform, deny, or confirm. We describe such a layer built for financial-services and healthcare workloads, in which policy is per-organization data effective without redeployment, enforcement happens at the tool-call boundary including structured arguments, and failure postures default closed. On a 438-item span-annotated benchmark across five languages it catches 91.8% of ground-truth PII values and 89.4% of high-sensitivity identifiers while modifying none of a PII-free control set, at a median enforcement cost of 2.4 ms. We then argue that the hard problem in operating such a layer is not building it but keeping it correct: the enterprise stack around it is non-stationary (languages are added, tools registered, dependencies upgraded, policies rewritten), and its effective coverage drifts against its declared configuration for mundane engineering reasons. Our evidence is the apparatus that found it. A harness measuring the deployed enforcement path, rather than the detection component inside it, revealed that the pipeline was producing person and location detections at 0.85 confidence and discarding them before redaction, catching 10.9% of high-sensitivity entities while every functional test was green. Correcting five such defects raised that to 89.4% at unchanged latency and no loss of utility. We generalize the five into a drift taxonomy and argue that continual evaluation against the deployed path, cheap enough to gate every change, belongs in the enterprise agent harness alongside the enforcement layer itself.