CLAIMRECEIPT: Auditable Evidence Contracts for Enterprise Agent Benchmarks
Abstract
Enterprise agent benchmarks increasingly inform deployment choices, yet a reported metric may be numerically reproducible from retained logs while still omitting claim-critical evidence or unfavorable runs. We separate two requirements: sufficiency, whether retained evidence uniquely determines a bounded claim, and coverage, whether the evidence spans the experiment set committed before outcomes were observed. We introduce CLAIMRECEIPT, a claim-relative receipt specification and selective verifier for multi-agent enterprise transactions. It binds typed public and auditor-only evidence to a signed experiment manifest and returns PASS, INVALID, or INCONCLUSIVE per claim. On 1,392 historical buyer–seller records, the verifier matches 5/5 manual audit verdicts, exactly replays 600 deterministic and 792 post-generation records, makes every one of 13 declared field groups non-redundant under tested ablations, and handles 11/11 frozen semantic faults with 0/8 false positives on benign variations. In a separate prospective epoch of 30 transactions, omitting one terminal receipt makes coverage inconclusive, while withholding private openings selectively blocks economic claims without erasing protocol verification. Instrumentation adds 0.021% of model-inference time and 9.9 KB per transaction. CLAIMRECEIPT turns enterprise-agent evidence from generic logs into an explicit contract over what a benchmark result is allowed to justify.