Reward Fusion is a Pairwise Contrast Problem: Why Ceiling Contact Can Harm Group Relative RL
Abstract
Reward fusion combines an exact outcome signal with a dense but noisy process signal before group relative policy optimization (GRPO). For a fixed rollout group, we show that the part of the local GRPO update due to reward is an exact sum of pairwise policy score differences weighted by normalized reward gaps. This identity separates two roles for process reward: it can order incorrect trajectories, but it can also erase or reverse the contrast between correct and incorrect trajectories. We distinguish \emph{ceiling collision}, where an incorrect reward ties the correct reward, from \emph{strict reversal}, where it exceeds it. Under stochastic monotonicity, both probabilities are nondecreasing in the quality of an incorrect solution; positive risk also requires process channel mass near the relevant threshold. A positive outcome margin gives a distribution free guarantee that every contrast across outcome classes has the correct sign. Empirically, process score tails track observed contact rates to about 1.5 percentage points. On hard math at 3B, fusion with a margin reaches percentage points over outcome reward alone; at 0.5B, CLIP degrades by percentage points while methods with a margin remain stable. A separate intervention shows that sufficiently frequent strict reversals can cause comparable degradation.