Counterfactual Rollout Replay: Forkable Environments as Free Process Rewards for Software Engineering Agents
Abstract
Reinforcement learning has rapidly become the dominant paradigm for training software engineering (SWE) agents on long-horizon, multi-turn tasks such as resolving real-world GitHub issues. Yet existing pipelines suffer from a fundamental credit-assignment bottleneck: rewards arrive only at the end of trajectories that may span dozens of file edits, shell commands, and test invocations, leaving every intermediate decision indistinguishable from luck. We observe that containerised SWE tasks provide a particularly useful form of environment forkability: a commit hash plus a container image specifies a real executable software state that can be restored cheaply and exactly, with auditable provenance and without a simulator-to-reality gap. We exploit this property by introducing Counterfactual Rollout Replay (CRR), a training-time procedure that re-executes the environment from selected decision points with alternative actions and assigns each selected step an advantage equal to the sampled on-policy versus counterfactual return differential. CRR requires no learned process reward model, no human rubric annotation, and no oracle hindsight signal; it is free of auxiliary process labels and reward-model training, while still incurring extra environment-replay compute. Applied to a 14B base model on SWE-Gym and evaluated on SWE-bench Verified and SWE-rebench, CRR-trained agents reach higher resolution rates with substantially fewer trajectories than outcome-only GRPO and complement orthogonal advances such as PRM-based scoring and trajectory search.