Enterprise-Scale Agentic Classification with a Governed Execution Envelope
Abstract
Enterprise classification workflows often combine structured fields, ambiguous narratives, controlled classification schemas, and operational reliability requirements. We present an anonymized enterprise deployment case study of a governed execution envelope combining bounded agentic reasoning, deterministic policy controls, qualification, provenance-preserving orchestration, recovery, and targeted human review. A paired ON/OFF comparison of 860 subset-level observations found exact-or-partial agreement of 47.1% with governance disabled versus 69.9% enabled (+22.8 percentage points; paired-bootstrap 95% CI: +18.7 to +26.9). On a protected 25-observation cohort, credited outcomes increased from 6/25 to 22/25. Separately, a production-style no-ground-truth campaign exercised the architecture over approximately 65K records, with about 435K model/API calls, about 3.9B tokens, and deterministic governance participation recorded for a large majority of completed records. We argue that governed enterprise agents should be evaluated as execution envelopes, not isolated models.