InfCoiL: Coordinated Planner-Controller Learning for Closed-Loop Physics-Based Human-Object Interaction
Abstract
Text- and goal-conditioned human-object interaction (HOI) provides an intuitive interface for specifying high-level intents in controllable character animation and embodied AI. Mapping such intents to physically plausible, contact-rich behaviors naturally requires hierarchical planning and control, yet existing planner-controller systems remain largely open-loop and lack adaptation to physical execution feedback. To address this, we present InfCoiL, a closed-loop framework for physics-based HOI from coordinated planner-controller learning. InfCoiL integrates a morphology-aware flow-matching planner with a unified multi-morphology controller in an autoregressive plan-and-execute loop, enabling diverse, controllable, and executable interactions across diverse objects and morphologies of humanoid embodiments. To further mitigate planner-controller distribution shift, we alternatively fine-tune both components under closed-loop reinforcement learning. Particularly, we introduce Semantic Flow-GRPO, which optimizes the flow-matching HOI planner with semantic covariance over heterogeneous HOI features and group-relative rewards from physical rollouts, while training the controller for reliable execution under contact-rich dynamics. Extensive experiments demonstrate that InfCoiL achieves state-of-the-art performance in both motion quality and task success. The code will be released upon acceptance.