PACE: Progress Actively-internalized Conditioning Execution for Long-horizon Manipulation
Abstract
Vision-language-action (VLA) models have made rapid progress in robotic manipulation, yet the policy acts with no sense of the pace at which its own execution unfolds—whether it is advancing, stalling, or has quietly gone wrong—especially for long-horizon tasks. Recent methods supply progress through external reward models or as an auxiliary prediction inside the policy, but the signal runs alongside action generation rather than actively shaping it. Moreover, progress is inherently tied to the terminal state, making estimation fragile when the policy lacks an internal awareness of where the task should end. We propose PACE, which actively internalizes task progress as a predictive interface bridging commitment and execution, equipping the policy with a calibrated sense of the pace at which it advances toward task completion. Specifically, we first equip the VLA with a foresight anchor, a predicted terminal representation from observation and instruction, from which progress emerges as a geometric projection. Building on this anchor, the model perceives a compact past–present–future progress state capturing how much the last step advanced, how far execution has come, and how far the next step should go. These predicted progress further actively conditions the action decoder to guide the prediction, while the gap between committed and realized increments provides a built-in consistency check for lightweight failure detection. Experiments across LIBERO, SimplerEnv, and real-world single-arm and bimanual platforms show that PACE achieves 98.7% on LIBERO and 74% on real-world manipulation, with consistent gains especially on long-horizon tasks. Robustness and stable generalization across diverse scenarios further confirm the advantage of internalizing progress within the policy, validating a calibrated sense of execution pace as an effective paradigm for long-horizon robot control.