When Does Variance Reduction Help GRPO? A CRN Case Study in Mathematical Reasoning
Karthik Chandrasekar
Abstract
Can sharing the start of two solutions improve reinforcement learning for mathematical reasoning? We study a sampling method inspired by common random numbers (CRN): paired responses share a prefix and then continue separately. A frozen Qwen2.5-7B probe reports reward correlation $\rho=0.47$ and a 3.9-fold variance gap between paired reward differences and independent, group-centered rewards. These are different quantities, so the gap does not measure the benefit of coupling alone. Under an idealized model with unchanged reward distributions, correlation reduces the variance of the unnormalized group-relative advantage by $\rho/(G-1)$, where $G$ is the group size. Using $\rho=0.47$ and $G=8$ gives 6.7\%; this is an illustration, not a measured reduction during training. In one training run per method, CRN-GRPO and Dr.GRPO have similar training rewards over 134 matched steps. At step 100, MATH-500 greedy accuracy is 73.8\% versus 70.8\%, but the comparison does not account for variation across training seeds. The study shows why a large reduction in a reward statistic is insufficient evidence of improved policy optimization.
Chat is not available.
Successful Page Load