Beyond the Model Response: Evaluating Human–AI Engineering Teams as Controlled Episodes
Abstract
Agentic systems are increasingly evaluated as if the model response were the primary unit of performance, even when real deployment outcomes depend on source quality, task framing, expert challenge, correction, and authority to act. We present an episode-level evaluation framework developed through more than two years of large-language-model (LLM) use in a high-consequence engineering environment. The framework treats the complete human–AI episode—source state, task contract, candidate output, human challenge, correction and evidence recovery, and final disposition—as the unit of evaluation. It records context variables, pre-review control burden, issue classes, correction burden, post-review confidence, and disposition while preserving the authority of controlled sources and qualified humans. We use a bounded retrospective dataset to illustrate the framework: a mixed-provenance expansion contains 99 usable records, with a 41-task engineering-source subset across six task families. Historical HV, HECS, and issue aggregates are descriptive historical evidence only; the archive contains measured and backfilled fields and does not establish reliability or predictive validity. We therefore make the local scoring and backfill rules explicit and treat the previously reported aggregate association as an internal-consistency observation rather than evidence. The prospective protocol requires blinded inter-rater measurement, direct accounting of reviewer burden, frozen rules, holdout evaluation, and external replication. The central claim remains deliberately narrow: in high-stakes engineering, evaluating the answer alone can miss the process that determines whether an answer is allowed to become action.