From Agent Traces to Behavioral Evidence: Benchmarking Observability Across Five Coding Agents
Abstract
Agentic coding tools are composite systems whose behavior spans model execution, agent loops, in- terfaces, tools, terminals, filesystems, repositories, and hosted infrastructure. Observing such behavior therefore requires evidence from system boundaries that are not uniformly represented in the teleme- try emitted by the tools themselves. We present an empirical benchmark of behavioral observability across five coding agents using six complementary observation mechanisms. We derive a 207-field tax- onomy of behavioral evidence across ten domains and use the Multi-Agent System Failure Taxonomy (MAST) to identify 35 fields for focused empirical verification. We evaluate native telemetry and ex- ternal observation under matched coding-agent executions and compare the evidence recovered across observation boundaries. Our taxonomy defines a broader space of observable field types, and collecting these fields across system boundaries provides a richer empirical basis for inferring agent behavior than native telemetry alone. The additional evidence includes developer-issued commands, filesystem changes and reversions, process liveness, delegation activity, and execution-context information. We addition- ally map SWE-bench, AgentBench, and Terminal-Bench to the ten behavioral domains to characterize which forms of behavioral evidence their evaluation protocols retain and evaluate. Our results show that behavioral analysis depends not simply on the existence of traces, but on the breadth of behavioral evi- dence represented and on whether observations from different boundaries can be related and interpreted consistently.