Aligning LLMs Toward Multi-Turn Conversational Outcomes Using Iterative RLHF
Abstract
Training large language models (LLMs) as multi-turn conversational agents remains a significant challenge, particularly in goal-oriented settings. The difficulty stems from sparse, long-horizon objectives and the discrepancy between response-level planning and token-level generation. In this paper, we present a formal reduction of the multi-turn RL problem into a \emph{sequence of single-turn RLHF-style problems}. This is achieved by setting a learned multi-turn Q-function as the reward model for the single-turn problem. We demonstrate and prove a key insight: solving this single-turn RLHF problem with standard token-level GRPO is equivalent to an approximate policy improvement step within the multi-turn problem. This insight naturally leads to \emph{Iterative GRPO}, a batch online approximate policy iteration algorithm that alternates between collecting a batch of data from the current policy, fitting Q-functions from these logged conversation trajectories, and improving the policy via single-turn RLHF. A major practical advantage is that Iterative GRPO directly leverages stable, off-the-shelf single-turn RLHF tools, making it straightforward to implement. In addition, the method occupies a middle ground between fully online and fully offline approaches, retaining the adaptability of online updates while gaining the stability benefits of offline training. Empirically, we demonstrate the effectiveness of Iterative GRPO on six multi-turn conversational environments.