Are Easier or Harder Examples Better? Rethinking Data Selection for Reward Models and Preference Optimization
Kevin Christian Wibisono ⋅ Aya Ismail ⋅ Pedro O. Pinheiro ⋅ Yixin Wang ⋅ Kyunghyun Cho ⋅ Nataša Tagasovska ⋅ Rajesh Ranganath
Abstract
Despite being crucial for effective LLM alignment, data selection remains understudied. Prior work on reward model (RM) training and policy optimization (e.g., DPO, GRPO) identifies *example difficulty*, the reward gap between chosen and rejected responses, as a key factor, but findings conflict: some favor easier examples with larger gaps, others harder ones. To isolate difficulty from confounders, we *assume access to a reference RM* and systematically study data selection across RM, DPO, and GRPO training. When difficulty is measured via the reference RM, *easier pairs consistently outperform harder ones*, especially for smaller base models: using only the top 20% easiest examples often matches or exceeds full-dataset performance while cutting post-training costs $5\times$. However, *this advantage hinges on reward estimation quality*. As the difficulty signal is corrupted by noise or estimated with a weak proxy RM, the easy-example advantage shrinks and can reverse. A signal-to-noise analysis explains why: larger-gap examples yield more reliable gradient directions, an advantage that weakens with noisier reward estimates. These results suggest that *conflicting findings in prior work partly stem from differences in reward reliability and signal mismatch*.
Chat is not available.
Successful Page Load