When Does Reflection Beat Reinforcement? Composing Prompt Evolution with Policy Gradients in Compound Agent Systems
Rohit Mohanty
Abstract
Reflective prompt evolution and group-relative policy optimisation are the two main ways of improving a multi-module language-model program from outcome feedback, and no published comparison asks when either wins once the two are combined. We report a controlled study of the two on a single harness, policy model (Qwen3-8B) and rollout budget grid, across a two-module HotpotQA program and a three-module competition-math program. We present four findings. First, composing GEPA with Dr. GRPO yields a task-dependent boundary with opposite signs, +0.10 mean reward on HotpotQA (10 of 12 cells) and $-0.16$ on math (0 of 12). Inspecting the returned prompts shows the cause is instruction wording, not reflection. GEPA returned its seed program unchanged in 12 of 12 math cells, and the seed's docstring instructions alone beat the trainer's defaults by +0.16 on HotpotQA and lose by 0.17 on math, whichever model GEPA optimised against. Where GEPA did edit the query instruction, the trainer's 64-token, raw-output query interface, not the edit, collapsed the reward. GEPA's own validation had scored that edit above the seed. Second, Reflective Credit Assignment, which turns an LM critique of a failed trajectory into a per-module advantage mask inside Dr. GRPO, produces no training benefit (-0.0006, 95\% CI [-0.005, +0.004] over 24 cells) although the same apparatus detects a learning-rate change. A method that only redistributes a group-relative advantage inherits that advantage's zeros, which occur on 25 to 66\% of steps here, and every reflection-in-RL method that works adds information or trajectories instead. Third, the scoring metric loses 53\% of achievable reward to answer formatting, 5.4 times GRPO's entire training gain over 8,192 rollouts. Fourth, on a fault-injection diagnostic with counterfactually verified ground truth no reflective localiser beats a constant predictor. A frontier LM judge asked to name the faulty module gains nothing from the execution trace on HotpotQA and is pushed below chance against origin-based ground truth on math, because the trace shifts its blame from the module that made an error to the module that failed to catch it. Blinded human labels reproduce the shift.
Chat is not available.
Successful Page Load