Chunking the Policy: Extending the Policy-Update Horizon with a Compact Transformer
Dong Tian ⋅ Markus Scholz ⋅ Olaf Landsiedel
Abstract
Action chunking adds temporal structure to reinforcement learning, yet the policy-update horizon---the number of consecutive policy-generated actions jointly evaluated during actor optimization---remains understudied. We present HOPS (Horizon-Extended Optimization over Policy-Generated Sequences), a fully online actor--critic that autoregressively composes fixed-size action chunks with a causally masked Transformer and evaluates them using twin prefix-conditioned critics. HOPS proposes four actions during interaction while optimizing 4--16-action sequences. Across all 50 Meta-World ML1 tasks, a compact two-layer policy reaches a terminal-success IQM of $0.881$ after $2\times10^{6}$ interactions, compared with $0.736$ for the main system-level baseline (T-SAC) under the same online-from-scratch protocol. Our evaluations on FANCY GYM Box Pushing provide complementary evidence of sample efficiency. Further, a benchmark-wide ablation across all 50 Meta-World ML1 tasks holds the interaction-time and critic-training horizons fixed while comparing $H_{\max}\in{4,8,16}$; at $2\times10^{6}$ interactions, the $H_{\max}=16$ setting attains the highest aggregate terminal-success IQM among the evaluated settings. Together, these results support the system-level effectiveness of HOPS and highlight the policy-update horizon as an important design dimension for interaction-efficient learning.
Chat is not available.
Successful Page Load