Instrument before you evaluate: write-path invariants for deployed human-agent systems
Abstract
Whether a deployed human-agent system can be evaluated is decided before anyone designs an evaluation for it. It is decided by the write path: what the system records, with what attribution, and what it allows itself to overwrite. A transcript plus a rating is enough to evaluate an agent and not enough to evaluate a pair, because the human half of a human-agent team leaves almost no durable trace that is distinguishable from the agent's own output. We report from a small production system in which the human's contribution is the measured object, and name four write-path invariants that post-hoc evaluation of a deployment depends on: attribution stamped at write time, integrity enforced as a predicate rather than a policy, append-only state, and marking rather than pruning. They differ in one respect that matters. Attribution missed at write time cannot be recovered for sessions already served; the other three can be installed late, at the cost of a discontinuity that has to be reported. Three are given with the dated defect behind them, and the integrity guard has none to show, because it keeps no record of what it refuses. We add a fifth constraint that is not about storage and, like attribution, has no late repair: an archive can be analysed for the first time exactly once, which makes continual evaluation on live deployment data a validity problem rather than an engineering one. We close with a checklist a team can answer about its own system before it designs an evaluation. This is a field report and a checklist from one small deployment, not a study, and the counts are given in both directions in Section 2.