From Patches to Trajectories: Privileged Process Supervision for Software-Engineering Agents
Murong Ma ⋅ Tianyu Chen ⋅ Yun Lin ⋅ Shuai Lu ⋅ Qinglin Zhu ⋅ Yeyun Gong ⋅ Zhiyong Huang ⋅ Peng CHENG ⋅ Yan Lu ⋅ Jin Song Dong
Abstract
Supervised fine-tuning (SFT) on long teacher trajectories is the dominant method for instilling investigation and reasoning capabilities into open software-engineering (SWE) agents. Under SFT, every retained response is an imitation target, so the student inherits not only the trajectory's outcome but also any flaw in its intermediate steps, including ungrounded leaps and redundant loops. High-quality training data must therefore be jointly effective (each step is grounded and narrows the agent's epistemic gap to the correct fix) and efficient (each step is information-bearing rather than redundant or looping). Existing recipes filter or relabel teacher rollouts using only a binary terminal verifier, which does not directly target these axes and provides no supervision on instances where the teacher fails. Every real issue ships with a developer-authored reference patch $p^\star$ that implicitly testifies to the file paths, runtime behaviors, and conventions a fix presupposes, but the standard pipeline discards it. We propose P2T (Patches-to-Trajectories), which uses $p^\star$ as privileged information during curation, and frames trajectory construction as a bi-objective program over per-step effectiveness and trajectory length. A reverse phase distills $p^\star$ into a latent process graph $G^\star$ of contextual facts and solution milestones, encoding dense intermediate anchors in constructive ordering. A forward phase curate trajectories from blinded teacher continuations, scoring per-step progress against $G^\star$ under a leakage-blocking groundedness check and committing the shortest segments that retain effectiveness. Using only 1.8k curated SWE-Gym instances, P2T improves both axes simultaneously over outcome-filtered SFT and its tool-error-masking variant: on SWE-bench Verified, it lifts Pass@1 by up to +10.8 points while cutting per-instance inference cost by 15%, with consistent gains on SWE-bench Lite and across two teachers. A size-matched ablation and qualitative analysis further isolate per-trajectory quality from data scale.
Chat is not available.
Successful Page Load