ESCROW: Guarded and Dual-Objective Continual Maintenance for Agents in Policy-Governed Enterprise Workflows
Ruoqi Shu ⋅ Chen Dan ⋅ Xuhui Wang ⋅ Tianhua Xu ⋅ Mengxi luo ⋅ Yanming Mai ⋅ Bo Wan
Abstract
LLM agents increasingly run policy-bound enterprise workflows, where they must apply rules consistently and stay auditable. Deploying such an agent is the start of its long-term maintenance cycle: it must adapt to a stream of *operational signals*, yet reliably turning these sparse, unlabeled signals into reusable skill revisions is hard, and a careless update can trade one task category's accuracy for the overall gain, revive a resolved failure, or land at an undeployable cost. We present **ESCROW**, a post-deployment maintenance framework that updates an agent's external, reviewable skills under a *Strict Update Boundary*: the LLM proposes candidate revisions, but only an empirically evaluated version is deployed. It combines distributed diagnosis with consensus, a per-category non-regression guard, cross-cycle anti-regression, and accuracy-cost Pareto search, emitting a versioned, auditable diff per change. In real production on our internal financial document-auditing system, it attains the strongest evaluated accuracy-cost trade-off among baselines, with a transfer probe on public $\tau$-bench.
Chat is not available.
Successful Page Load