Multi-Stage Planning from Single-Stage Data: Reinforcement Learning Helps Composition but Requires Anchoring
Abstract
Long-horizon planning often requires composing familiar local transitions while remaining grounded in a specified goal. We study this setting from single-stage data: on a controlled no-skip graph benchmark, training provides only adjacent-stage demonstrations while evaluation requires composing them on unseen long-horizon start–goal pairs; an event-chain diagnostic decomposes success into legality, stage-bridging, and goal-conditioned termination. Across a minimal Transformer, Qwen2.5-3B, and an external validation on MuSiQue multi-hop QA, supervised fine-tuning (SFT) learns local transitions but degrades with horizon, with failures localized to stage-bridging. Reinforcement learning (RL) narrows this composition gap, especially at longer horizons, but unanchored RL drifts—producing valid stage progress while ignoring the instructed target. Stable gains require two complementary forms of anchoring: instruction-style prompting strengthens prompt-side goal-conditioning, especially at shorter horizons, whereas KL regularization to the SFT reference policy prevents RL-induced drift across horizons; reward shaping mainly provides denser credit. Finally, we identify a Decomposition Paradox: models trained only on long compositions can reach intermediate targets yet fail the corresponding prefix subtasks because stopping at instructed intermediates is never supervised.