HAI: Hierarchical Anchored Interaction for Multi-View Bimanual World Models
Abstract
Multi-view bimanual robot world models must predict future observations while preserving scene identity, following two synchronized but non-exchangeable arms, and maintaining consistency between global and wrist-local views. Existing conditioning schemes often collapse left- and right-arm actions into a view-agnostic control signal, making long-horizon rollout prone to scene drift, wrong-arm responses, and cross-view inconsistency. We propose \textbf{HAI}, a \textbf{Hierarchical Anchored Interaction} architecture for controllable multi-view bimanual world modeling. HAI organizes prediction as a structured information flow. First, hierarchical action-view conditioning routes the bimanual action chunk to camera streams, injects coarse chunk-level intent, performs structured multi-view modeling on the coarse-modulated features, and then injects fine per-frame action control. Second, anchored dynamic generation fuses the fine action-conditioned descriptors with persistent per-view scene anchors before decoding future observations. Experiments on AgiBot and DROID show that HAI improves long-horizon rollout quality over the baselines. Ablations confirm gains in scene stability, wrist-view controllability, and cross-view consistency. Further experiments on policy-improvement diagnostics indicate more useful synthetic rollout signals for downstream VLA policies, achieving an averaging 67.2\% relative gain in success rate. The project page is available at \url{https://hai-anon.github.io/}.