Chain-of-Correction: Progress-Aware Policy Steering via Anchor-Grounded Predictive Reasoning
Abstract
Imitation learning policies for robotic manipulation inherently suffer from compounding errors. To mitigate this, inference-time policy steering introduces predictive reasoning to evaluate and refine action proposals before execution. However, existing methods typically anticipate consequences in unstructured pixel or latent spaces, leading to physically inconsistent predictions and unreliable error detection. Furthermore, their blindness to global progress often yields locally plausible corrections that fail to complete the overall task. To address these limitations, we propose Chain-of-Correction (CoC), an inference-time steering framework driven by a single VLM. Shifting away from unstructured state predictions, CoC explicitly leverages anchor-grounded scene graphs to track exact 3D physical relations. Built upon this representation, a progress-aware dual-check mechanism verifies historical execution and predicts future task advancement to systematically intercept myopic action proposals. Through a failure-augmented supervised fine-tuning pipeline, CoC seamlessly unifies scene graph extraction, progress evaluation, and precise kinematic correction, eliminating the need for task-specific world models. Extensive experiments across the COLOSSEUM benchmark, a newly introduced offline protocol for predictive reasoning, and real-world tasks demonstrate that CoC significantly outperforms state-of-the-art baselines in fine-grained error detection and robust error recovery.