What pass@k Cannot Measure: A Construct-Validity Stress Test for RL Post-Training Evaluation
Subham Rath ⋅ Rajat Dandekar ⋅ Sreedath Panat ⋅ Raj Dandekar
Abstract
pass@$k$, the fraction of problems a model solves within $k$ sampled attempts, is the field's default protocol for deciding whether reinforcement-learning (RL) post-training on verifiable rewards improved a model. At the population level, pass@$k$ depends only on a problem's probability of a correct sample and has no term for how that probability is distributed across distinct outputs. We show this structural gap is not academic. Training Qwen2.5-1.5B-Instruct on grade-school math with Group Relative Policy Optimization (GRPO) and with rejection-sampling fine-tuning (RFT, training on the model's own shortest verifier-passed rollout) moves three complementary diversity measures (token-level entropy, answer-level entropy, and unique answers per prompt) in opposite directions, with zero overlap across three seeds per arm. The gap survives restricting to verifier-correct completions (lexical diversity among correct solutions is 15% lower for GRPO, holding after controlling for completion length) and to a count-controlled incorrect-answers-only check that removes correct answers from the comparison entirely (a 2.2–2.9% gap once the incorrect-sample pool size itself is controlled for, versus the order-of-magnitude-larger raw gap that pools correct and incorrect answers together). Yet pass@8 and pass@32 show no consistent winner on GSM8K: a paired bootstrap puts the minimum detectable pass@1 effect at 80% power near $0.02$, well below the observed gap. A deliberately hard MATH-500 subset with substantial headroom (pass@16 below $0.67$) shows a related pattern: the arms are separable at pass@1–2, with no difference detected from pass@4 on. Compared against the starting checkpoint, no trained arm significantly improves hard-problem coverage: RFT is significantly worse, while GRPO is statistically indistinguishable from the starting checkpoint, and GRPO's pass@1 edge over RFT there reflects a smaller loss rather than a capability gain, a lesson about missing controls separate from a failure of pass@$k$ as a metric. On GSM8K, only pass@1, with no role in detecting diversity by construction, separates the arms cleanly, and it rewards the arm whose correct solutions are the least diverse. We argue this is a concrete instance of the standard evaluation protocol lacking construct validity for a property practitioners routinely read off it, at this model scale and training budget, and outline what would be required to test its generality.
Chat is not available.
Successful Page Load