Action Chunking Proximal Policy Optimization with Feedback Correction
Abstract
Action chunking can improve credit assignment and exploration in reinforcement learning by shortening the effective decision horizon, but existing action chunking reinforcement learning methods often rely on chunked state-action value functions which become challenging to learn as action dimensionality grows. Moreover, executing action chunks open loop sacrifices within-chunk reactivity, which is especially problematic in contact-rich robotic control. We present Action Chunking PPO (ACPPO), an extension of PPO that introduces temporal abstraction through a chunked actor while retaining a standard state-value critic, thereby avoiding chunked Q-functions. We further propose ACPPO-Corr, which augments the chunk planner with a stepwise feedback corrector that adjusts planned actions online and provides reactivity within a chunk. Across 25 simulated robotics tasks from IsaacGym and Bi-DexHands, spanning locomotion, arm manipulation, and dexterous hand-object interaction, ACPPO-Corr achieves the strongest aggregated benchmark performance and remains best on both decision-frequency-sensitive and decision-frequency-neutral task subsets. Ablations show that moderate chunk lengths work best and that regularizing the corrector balances long-horizon planning and local feedback. These results suggest that action chunking can be made effective in fully online PPO when open-loop temporal structure is paired with closed-loop correction.