Correlated Criteria Drive Reward Over-Optimization: Redundancy-Aware Scheduling for Checklist Rewards
Abstract
Post-training on open-ended tasks usually replaces a checkable answer with a checklist of criteria that a language model grades. We asked a narrow question about that setup: when a policy over-optimizes the checklist, which criteria is it exploiting? In a testbed where the true quality of every answer is known by construction, the answer is redundancy. Checklists contain groups of criteria that one presentational move satisfies together, and since the reward is a sum, that move is paid once per member, so its value scales with the size of the group rather than with the quality it adds. The same fact explains a negative result we did not expect: holding criteria out at random does nothing at all, because under a group-shared mask it rescales the group-relative advantage and leaves the optimum where it was. What does work is to schedule by redundancy, estimating online which criteria co-fire from verdicts the grader has already returned and scoring one representative per cluster. On Finance, Consulting and STEM checklists this recovers up to 8.4 points of audited quality and cuts over-crediting by up to 39.1 points, where uniform hold-out, round-robin hold-out and discriminative reweighting recover none. It also has a clean boundary: the schedule repairs redundancy in the checklist and does nothing about credulity in the grader.