Beyond the Pull Request: Evaluating Engineering Judgment in Coding-Agent Trajectories
Abstract
Repository-level coding benchmarks primarily ask whether an agent produced a patch that passes tests. Production engineering also depends on behavior that a test suite may not encode: tracing downstream consumers, challenging an unsafe premise, distinguishing evidence from assumption, preserving recovery paths, and avoiding unnecessary intervention. We present SDLC-Bench, an expert-authored, model-assisted benchmark for interpreting such behavior in coding-agent trajectories. Tasks are grounded in licensed private production codebases and resemble incomplete requests from colleagues, tickets, incidents, and design questions; the requested work product may be a patch, a diagnosis, a review, a consultation, or a plan. A judge agent reads the trajectory, diff, and executed repository checks against task-specific, repository-grounded guidance. We validate the judge against blinded software-engineer comparisons: when the judge and the human panel both express a preference they agree on 99.4% of pairs and reverse on 0.6%, while the judge declares a tie on 56% of pairs where humans have a clear preference; regrading the same attempts six times yields a within-attempt standard deviation of 0.02. On a 159-task common set under one coding-agent harness, Claude Opus 5 obtains a behavioral score of 0.749 and Claude Fable 5 obtains 0.595, with the gap larger on multi-turn tasks. A source-audited review of the largest per-task gaps attributes most of them to calibration and honesty rather than to code, and the few reversals to over-reach: invented risks, unsupported claims, and over-built solutions. This yields a two-sided account of engineering judgment—hazard tasks require surfacing real hidden constraints, restraint tasks require not inventing them—and a set of validity requirements: prospective restraint controls, frozen rubrics, cross-judge checks, task-clustered uncertainty, and evidence-linked labels. The benchmark makes coding-agent behavior under underspecified requests directly measurable.