CM2: Reinforcement Learning with Checklist Rewards for Multi-Turn and Multi-Step Agentic Tool Use
Zhen Zhang ⋅ Kaiqiang Song ⋅ Sean Wang ⋅ Yebowen Hu ⋅ Weixiang Yan ⋅ Chenyang Zhao ⋅ Henry P Zou ⋅ Haoyun Deng ⋅ Sathish Reddy Indurthi ⋅ Shujian Liu ⋅ Simin Ma ⋅ Xiaoyang Wang ⋅ Xin Wang ⋅ Song Wang
Abstract
AI agents increasingly solve real-world tasks by reasoning over multi-turn user interactions and invoking external tools. Yet reinforcement learning (RL) remains difficult in this setting: realistic objectives are often open-ended and lack verifiable rewards, RL for multi-turn, multi-step agentic tool use is underexplored, and executable tool environments are costly to build and maintain. We propose CM2, an RL framework that replaces verifiable outcome rewards with checklist rewards. CM2 decomposes each turn's intended behavior into fine-grained binary criteria with explicit evidence grounding and structured metadata, turning open-ended judging into more stable classification-style decisions. To balance stability and informativeness, CM2 uses sparse reward assignment with dense evaluation criteria, and trains in a scalable LLM-simulated tool environment to avoid engineering large tool sets. Experiments show that CM2 consistently improves over supervised fine-tuning. Starting from an 8B Base model and training on an 8k-example RL dataset, CM2 improves over the SFT counterpart by 8 points on $\tau^2$-Bench, 10 points on BFCL-V4, and 12 points on ToolSandbox. The results match or even outperform similarly sized open-source baselines, including the judging model. CM2 thus offers a scalable recipe for optimizing multi-turn, multi-step tool-using agents without verifiable rewards.
Chat is not available.
Successful Page Load