Confidence-Guided Policy Refinement for Iterative Code Debugging
Abstract
Group-relative policy optimization methods, such as GRPO, learn from the relative reward of trajectories sampled for the same prompt. When every trajectory in a group receives the same reward, the centered advantages vanish and the group contributes no policy-gradient signal, even though the environment was queried and feedback was obtained. We introduce Confidence-Guided Policy Refinement (CGPR), a training-time framework that uses this feedback to convert such uninformative groups into informative ones, rather than discarding them and resampling independently. CGPR builds a bounded refinement tree over the trajectories of an uninformative group, allocates the refinement budget with a UCB-style rule, and replaces non-anchor trajectories only when a candidate increases the reward dispersion of the group. Refinement occurs only during training and adds no inference-time computation. Across HumanEval+, MBPP+, APPS, and Codeforces, CGPR matches or improves upon GRPO on both pass@1 and pass@10. Relative to GRPO, CGPR achieves pass@1 gains of up to +4.2 percentage points and pass@10 gains of up to +3.4 points, with the largest improvements concentrated on the more challenging APPS and Codeforces benchmarks. CGPR also reduces the fraction of uninformative rollout groups relative to GRPO, and reaches informative training groups using fewer rollout tokens than DAPO-style filter-and-resample.