Governing Meta-Agent–Driven Harness Improvement without an Executable Verifier: A Live Enterprise Study with Revisable Expert Feedback
Yu-Ting Huang
Abstract
Harness optimization revises prompts, tools, and orchestration from execution feedback, but usually assumes a repeatable correctness signal. We study meta-agent-driven harness improvement for anomaly discovery from employee screen activity in a live enterprise deployment. The task is open-ended because review-worthy behavior types cannot be exhaustively enumerated and potential alert intervals are not predefined. Correctness further depends on evolving policy and tacit expert judgment rather than an executable verifier. We therefore use successive versions of an expert-curated reference set as a \emph{soft verifier}. Each frozen version provides repeatable case-level judgments for the finite cases it contains but cannot determine correctness for arbitrary cases; expert-reconciled changes enter subsequent versions across development cycles. To narrow the harness revision search space while preserving human governance over the revision process, we externalize intermediate behavioral candidates through an explicit candidate interface. This interface also enables versioned checkpoints to be recombined to localize where performance differences emerged over harness evolution. Replacing the early alert-adjudication checkpoint with its late counterpart increased joint-system $F_1$ by 11.8--11.9 candidate-generation checkpoint with its late counterpart produced differences of only -2.2 and -2.4 points, with both confidence intervals spanning zero. This asymmetry is consistent with the intended concentration of evolving organization-specific judgment in alert adjudication. The comparatively small differences between candidate-generation checkpoints motivated an artifact-matched comparison of preserving this interface across separate reasoning contexts against consolidating the same substantive artifacts within a single monolithic context. Relative to this control, the interface-preserving configuration showed higher precision, recall, and $F_1$ point estimates, but the corresponding 95\% paired bootstrap confidence intervals crossed zero. Therefore, the incremental performance advantage from preserving the explicit interface remains unresolved in this study. We also design an offline--online evaluation protocol adapted to the practical constraints of this open-ended enterprise workflow and demonstrate its use in live deployment.
Chat is not available.
Successful Page Load