Beyond Outcome: Trajectory-Driven Prompt Optimization via Multi-Dimensional Rewards
Abstract
Automatic Prompt Optimization (APO) aims to improve the capabilities of Large Language Models (LLMs), yet existing methods remain limited by their result-oriented nature. Training-free approaches often converge to suboptimal prompts, while training-based methods suffer from reward sparsity and optimization instability. More importantly, both paradigms focus on obtaining the final prompt rather than learning the prompt optimization process itself. To address this dilemma, we propose \textbf{ProPO}, a trajectory-driven \textbf{Pro}cess learning framework for \textbf{P}rompt \textbf{O}ptimization guided by multi-dimensional composite rewards. Specifically, ProPO introduces a historical trajectory reflection mechanism for high-quality exploration. Concurrently, we adopt adaptive rubric rewards to provide fine-grained supervision and alleviate reward sparsity. More importantly, we design a novel Process Logic Reward that supervises the consistency of the reflection process. Extensive experiments demonstrate that ProPO, integrated with GRPO, establishes a new SOTA performance. Notably, ProPO achieves a remarkable \textbf{70.00\%} accuracy on the highly complex AIME-2026 task and an absolute gain of \textbf{+25.25\%} over the zero-shot baseline on Subjectivity Classification (Subj). Furthermore, empirical analyses reveal that our multi-dimensional reward mechanism provides finer-grained and denser signals, allowing the model to learn the prompt optimization process, and ultimately enabling it to generate significantly higher-quality prompts.