Coding Agents Succeed and Fail on the Same Issue
Abstract
Modern coding agents have demonstrated the capability to solve real GitHub issues, yet the same agent that solves an issue on one run can fail the same issue on the next run. We isolate this phenomenon in the public Nebius OpenHands SWE-rebench dataset, which keeps many repeats on the same issue, including many unresolved runs. We specifically investigate issues where the same agent succeeds on many runs, indicating it can already resolve the issue, yet still fails many runs, indicating a lack of consistency. Our testbed consists of 248 issues across 190 GitHub repositories (3,758 runs), where Qwen3-Coder-480B-A35B-Instruct and the OpenHands harness succeeded at least 5 times, and failed at least 5 times, per issue. Among the failed runs, we filter out unfinished or off-file failures, and runs with failing local tests, and what remains are what we call near-misses. Near-misses account for the majority (65.6%) of unresolved runs and appear on 85.9% of those issues. Near-misses look remarkably similar when compared directly to the agent's successful runs on 6 statistics, including message count, tool calls, edits, source files, test commands, and repeated tests. Despite their similarities on these counts, when the submitted patches of a near-miss and a success are compared with the outcome labels hidden, a judge almost always labeled them as different behavior. When two successes are compared in the same way, they are also often labeled as different behavior, albeit at a meaningfully lower rate than the former, indicating that two patches labeled as different behavior can both still be marked resolved. Practitioners aiming to build more reliable agents should not treat unresolved as one group. On these issues, diagnosing failures may require inspecting near-misses, which are hard to distinguish from successes on these counts but are usually still labeled as different behavior from a success.