From Benchmark to Adoption: Coding Agent Evaluation Needs User-Level Harness
Abstract
Coding agents are increasingly adopted in real-world software development, yet their evaluation remains largely benchmark-driven and model-centric. Existing approaches focus on correctness and task completion, offering useful signals of isolated capability but limited insight into whether agents can be reliably used, controlled, and trusted in practice. We argue that the problem is not only how coding agents are evaluated, but what is being evaluated: coding agents are evaluated as models, while they are adopted as harnessed systems. This gap becomes most visible at the point of adoption. In real workflows, whether a coding agent is actually used is not determined by base-model capability alone. Instead, users make agents usable through user-level harnesses: situated configurations of task scope, context, constraints, checkpoints, workflow adaptations, and verification practices that shape and validate model behavior. These harnesses often remain informal and ad hoc, yet can determine whether the same model succeeds in one setting and fails in another. To ground this position, we draw on three empirical studies with users. Study 1 identifies contextual conditions shaping when users adopt or avoid coding agents. Study 2 identifies execution and control requirements for agentic systems. Study 3 identifies verification and code-quality criteria beyond correctness. Together, these studies suggest that users evaluate not only agent outputs, but also where agents can be used, how their behavior is guided, and whether their outputs are acceptable in practice. We interpret these conditions as user-level harnesses that make coding agents usable, controllable, and acceptable in real workflows. Based on these findings, we argue that evaluation should move beyond model-centric and system-internal metrics toward harness-aware, user-grounded assessment. In particular, evaluations must make explicit the conditions under which coding agents are usable, controllable, and trustworthy in practice—across dimensions such as environment, users, capabilities, and code quality. We call on the AI research community to place user-level harness at the center of coding-agent evaluation.