Foundation Pareto Flow Policy for Multi-Objective Reinforcement Learning
Abstract
The overarching goal of reinforcement learning (RL) is to efficiently optimize the agent's policy for various environments, tasks, dynamics, and even user preferences, which often consist of multiple and potentially conflicting rewards/objectives. This necessitates the learning of Pareto-optimal policies that can navigate different trade-offs among objectives. However, existing multi-objective RL (MORL) methods, whether trained online or offline, are typically limited to task- or objective-specific policy learning and do not address generalization of learned skills to unseen tasks/objectives. In this work, we propose the Foundation Pareto Flow Policy (FP2) for MORL, inspired by the pre-mid-post-training paradigm underlying recent foundation models. Specifically, we instantiate FP2 as a multi-modal diffusion transformer equipped with a likelihood-free flow matching loss, enabling unified, scalable, and instructable learning of diverse Pareto-optimal trajectories. We further develop a three-stage training pipeline to gradually optimize the policy: (1) pre-training on a large-scale MORL dataset, (2) mid-training on a carefully curated high-quality dataset, and (3) post-training on few collected trajectory data from new tasks and trade-offs. Empirical results demonstrate that FP2 consistently outperforms previous state-of-the-arts on the D4MORL benchmark, achieving superior zero-shot interpolation and extrapolation on complex Pareto manifolds as well as efficient few-shot adaptation to new task/dynamics scenarios.