When Does a Software-Agent Benchmark Measure Engineering Judgment? A Construct and Task-Distribution Audit
Abstract
Coding-agent benchmarks often operationalize ability as resolving a specified issue under executable tests. Production engineering also requires judgment: discovering omitted constraints, rejecting a false premise, calibrating claims to evidence, and avoiding unnecessary intervention. We audit the construct and task-distribution validity of an evolving software-agent evaluation program grounded in licensed production repositories. The program's 530 authored tasks are a development corpus, not a frozen benchmark. The intended design represents behavioral and correctness signals separately, using deterministic checks plus repository-grounded guidance; the reported 159-task pilot is behavioral only, and four illustrative axis-divergence examples come from a separate later regrade workstream. The pilot reports means of 0.749 and 0.595 on the current common set, but the set is not random, configurations and paired uncertainty are incomplete, and a preliminary audit suggests hidden hazards are overrepresented relative to restraint cases. We argue that benchmark composition is part of the construct and propose a validity protocol covering provenance, archetype balance, repository clustering, judge reliability, contamination, and decision-relevant reporting.