Continuous Process-Level Evaluation for Evolving Enterprise AI Agent Skills
Abstract
Enterprise AI agent skills are not static artifacts: they are continually revised as tool APIs change, LLM versions update, and skill specifications are refined in response to operational feedback. Yet the standard practice for validating each revision is to check only final-output accuracy—an approach that system- atically misses process-level behavioral drift introduced during evolution. We present a continuous evaluation framework for enterprise agentic skills that com- bines outcome-level and process-level quality checks, applied to two Business Value Determination (BVD) skill variants within an enterprise Value Aware Re- siliency (VAR) system. The framework independently computes ground truth per run, instantiates template test cases as persistent regression tests, and assesses tool selection, argument correctness, execution ordering, and database integrity using programmatic checks augmented by a narrowly scoped LLM judge when exact matching would be brittle. We report a fully automatic evaluation over 240 trials spanning two related but structurally distinct skills (Revenue and Pro- ductivity apportionment), two specification variants (SKILL.md and Skill.txt), two agent harnesses (Claude Code and Codex), and three models (GPT-5.6-Sol, Claude Opus 4.8, Claude Sonnet 4.6). Across the 240 automatic trials, 175 passed all applicable final numerical checks; among them, 162 (92.6%; Wilson 95% CI: 87.7–95.6%) still had at least one additional evaluator-detected deviation. Under a broader seven-check final-state definition, 151/164 passing runs (92.1%; 95% CI: 86.9–95.3%) still violated a trajectory check. Dependency attribution descrip- tively compressed a mean 6.34 failed checks per run to 2.65 roots. Specification- variant sensitivity also varies with model and harness: unadjusted bootstrap in- teraction intervals exclude zero for all three Revenue comparisons and for GPT on Productivity, but not for the two remaining Productivity comparisons. Tem- plate test cases with runtime-resolved placeholders provide reusable regression coverage across the evaluated specifications, models, and harnesses; longitudinal validation under actual API evolution remains future work.