Long-Term Composition of Human-Object Interactions
Abstract
Real-world human activities unfold as sequences of temporally dependent human-object interactions (HOIs), yet existing methods address either atomic HOIs in isolation or text-driven motion composition without object or scene context. To bridge this gap, we introduce the task of long-term action composition of HOIs, where a human navigates between and interacts with multiple objects within a 3D scene. The primary challenges lie in producing smooth transitions, scene-aware locomotion, and plausible hand-object interaction. To this end, we propose LT-HOI, a diffusion-based model that jointly conditions on past and future contexts—the past providing kinematic continuity across transitions, the future enabling anticipatory navigation toward upcoming interacting objects. We further integrate Diffusion Noise Optimization for collision avoidance and introduce a hand-object displacement loss to improve contact quality. On the ParaHome dataset, we establish the first benchmark for this task and demonstrate that LT-HOI effectively composes long, temporally coherent HOIs.