Not Every Rubric Teaches Equally: Policy-Aware Rubric Rewards for RLVR
Utkarsh Tyagi ⋅ Xingang Guo ⋅ MohammadHossein Rezaei ⋅ Daniel George ⋅ Anas Mahmoud ⋅ Jackson Lee ⋅ Bing Liu ⋅ Yunzhong He
Abstract
Reinforcement learning with verifiable rewards has made post-training effective when correctness can be verified automatically, but many useful model behaviors require satisfying several qualitative criteria rather than a single verifier. Rubric-based rewards extend RLVR to these open-ended domains by grading prompt-specific criteria and aggregating them into a scalar reward. However, the common static aggregations conflate a criterion's desired end-state importance with its current usefulness for learning. We show that this conflation is central in rubric RL: many criteria that matter to the final answer are already saturated or currently unreachable for the policy, while criteria with informative rollout variation are not reliably the ones assigned the largest human weights. We introduce a Policy-Aware Rubric Reward framework for RLVR, \textbf{POW3R}, that keeps human weights and reward category balance as the target objective while reallocating within-category training pressure toward criteria that distinguish the current rollouts. Across three base policies on each of two datasets spanning multimodal and text-only settings, POW3R takes first place on $27$ of $33$ (base policy, metric) cells we evaluate, leading both mean rubric reward and the harder strict ``every-rubric-passed'' pass-rate over vanilla GRPO with rubric-based rewards, and reaches the same plateau in $3$--$4\times$ fewer training steps. These findings suggest that rubric rewards should separate what should matter in the final answer from what can teach the current policy.
Chat is not available.
Successful Page Load