TGPO: Trace-Guided Policy Optimization for Robot Task Planning via Verifiable Subgoal Generation
Abstract
Robot task planning in real-world environments requires mapping abstract natural language instructions to executable action sequences under long horizons and complex constraints. While large language models (LLMs) provide strong commonsense reasoning, they often fail to generate reliable and feasible plans. In contrast, symbolic planners ensure feasibility and optimality but require well-specified goals and cannot directly interpret high-level human intent. We formulate robot task planning as learning to generate verifiable subgoals in the Planning Domain Definition Language (PDDL), bridging language understanding and symbolic planning. To address the challenges of sparse and noisy supervision, we propose Trace-Guided Policy Optimization (TGPO), a reinforcement learning framework that improves structured subgoal generation through (i) verifier-grounded rewards, (ii) external correction of intermediate reasoning traces, and (iii) constrained policy updates that incorporate corrected traces into training. We evaluate TGPO on large-scale household planning tasks with long horizons, abstract instructions, and complex constraints. TGPO significantly outperforms prompting-based and reinforcement learning baselines, with the largest gains on abstract tasks. Furthermore, TGPO integrates naturally with symbolic planners and language-conditioned executors, enabling robust long-horizon planning and combinatorial generalization in realistic environments.