Whose Behavior Is It? A Generalizability Study of Coding-Agent Behavior Across Models, Harnesses, and Tasks
Abstract
Behavior observed in a coding-agent trajectory can reflect the underlying language model, the harness that runs it (prompts, tools, control loop, limits), the task, or variation within a single run. Yet studies often describe such patterns as properties of a model or system, even though a claim attributed to the wrong source may not generalize when the surrounding conditions change. We ask which coding-agent behaviors travel with the model, which are imposed by the harness, and which are too task- or run-specific to support either attribution. We treat this as a measurement problem and apply generalizability theory to 24,986 publicly released SWE-bench Verified trajectories spanning 40 models, 9 harness families, and a common set of 500 tasks. After normalizing the trajectories into a common step-level representation, we measure 25 behaviors and decompose their variation across model, harness, task, and run. The answer is behavior-specific. Process behaviors such as reproducing the issue or executing code before editing are largely model-portable, even under similar harness instructions. Protocol behaviors such as explicit submission and editing existing tests are instead shaped by the harness. Other behaviors are highly situational, varying enough across tasks and runs that observations from a small number of trajectories do not reliably characterize the model. The agent's own final message is also a weak indicator of success: confident success claims frequently occur on unresolved tasks, and their correctness depends much more on the task than on the model. Together, these results show that behavioral observations have different scopes of validity: some support claims about a model, some about the harness, and others should not be generalized beyond the particular setting in which they were observed.