H-GenPO: Hierarchical Generative Policy Optimization via the Option-Critic Framework
Abstract
Diffusion-based policies have demonstrated remarkable performance in online reinforcement learning (RL) by virtue of their expressive, multimodal action distributions. However, existing methods operate on a flat architecture that must generate a new action at every timestep, with no mechanism for temporal commitment to a behavioral mode—making them ill-suited for context-dependent tasks that require sustaining coherent behavioral modes, regardless of how expressive their action distributions are. We propose Hierarchical Generative Policy Optimization (H-GenPO), the first framework to integrate on-policy diffusion-based RL with the temporal abstraction of the Option-Critic architecture. H-GenPO adopts a two-level hierarchy in which a high-level policy over options selects among a discrete set of options, while a unified flow matching model—conditioned on a learned option embedding—serves as the expressive low-level intra-option primitive. The shared intra-option policy design encapsulates diverse behavioral repertoires within a single set of parameters, enabling scalable hierarchical control without the parameter growth associated with maintaining separate policy networks per option. We evaluate H-GenPO on 8 standard continuous control tasks in Isaac Lab and 3 custom context-dependent tasks designed to require coordinated behavioral mode switching. Our empirical evaluation demonstrates that H-GenPO achieves the best mean rank among all baselines on both standard and context-dependent benchmarks, while also exhibiting interpretable option specialization that emerges without any option-level supervision.