IFDECORATOR: Wrapping Instruction Following Reinforcement Learning with Verifiable Rewards
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as a promising approach to enhance Instruction Following (IF) capabilities of large language models (LLMs). However, RLVR for instruction following remains sample-inefficient and prone to reward hacking, where LLMs exploit verification shortcuts rather than fulfilling the core intent. To address these challenges, we frame RLVR for instruction following as an integrated environment that unifies dynamic task generation and robust training. We introduce Instruction Following Decorator (IFDecorator}, a framework coupling difficulty adaptation, intent alignment, and hack diagnostics. It features: (1) a cooperative-adversarial data flywheel that yields challenging yet solvable tasks for sample-efficient training; (2) IntentCheck, a gating module to mitigate reward hacking; and (3) TripWires, a proactive diagnostic tool for eliciting hacking behaviors and quantifying their prevalence. Extensive experiments show our Qwen2.5-32B-Instruct trained with IFDecorator using only 3,625 training examples achieves 87.43% on IFEval (outperforming GPT-4o) and improves FollowBench by 4.2%, while preserving general capabilities. Diagnostics confirm that our method effectively reduces reward hacking. The approach generalizes across model architectures and scales. We will release code and data for future research.