EvoCodeBench: Evaluating Coding Agents in Multi-Turn Iterative Interactions
Abstract
Coding agents are increasingly used as iterative development partners, but most benchmarks still evaluate one specification followed by one final assessment. This fails to capture whether agents can maintain executable artifacts as requirements evolve. We introduce EvoCode-Bench, a benchmark of 26 stateful coding tasks spanning 227 evaluated rounds. Each task preserves the agent's workspace across 5--15 rounds, expresses requirements through observable behavior, and uses cumulative executable tests to check both new requirements and regressions against still-active prior ones. We evaluate 13 coding agents with two complementary metrics: MT@4, a four-attempt fail-stop multi-round score, and SR, a single-round score from a reference-completed prior state. For most agents, SR exceeds MT@4 by 22--40 points, indicating that isolated instruction-following does not imply persistent reliability. The gap also changes model rankings: the highest-SR agent (78.9) ranks only third in persistent execution (44.0 MT@4). Even the strongest agents achieve only \~50\% success on multi-turn metrics, and aggregate pass rate drops below half of round-1 performance by round~5. Failure analysis reveals tier-dependent behavior: weaker agents fail early, while stronger agents survive long enough to expose specification-tracking and regression failures. These results show that single-round and persistent multi-turn coding evaluate different capability dimensions. We release the benchmark data and Harbor multi-turn infrastructure.