ForceLDM: A Force-aware Latent Dynamics Model for Contact-Rich Manipulation
Abstract
Contact-rich manipulation requires precise control and complex physical interaction with the environment, posing significant challenges for robot policy learning. Vision-only policies struggle with such tasks, as contact states and force feedback are difficult to infer from vision alone. Existing methods mitigate this issue by incorporating force or tactile sensing. However, they often treat these signals as auxiliary observations rather than using them to model future interaction dynamics, leading to inefficient modality utilization and limited generalization. To address this, we propose ForceLDM, a force-aware latent dynamics model that enables the policy to anticipate future contact and motion dynamics in latent space and use them to guide action generation. Rather than predicting future observations in pixel space, which often requires large amounts of data and risk overfitting irrelevant details, ForceLDM learns a compact task-relevant representation through knowledge distillation. Specifically, a teacher network with privileged access to future force/torque and optical flow signals extracts future dynamic features, which are then distilled into a student network trained only on current observations. This enables the student to reason about upcoming contacts and physical dynamics at deployment without relying on future sensory inputs. To improve robustness and generalization, we further introduce a curriculum-based progressive noise injection strategy to mitigate over-reliance on future features. Experiments on five real-world contact-rich manipulation tasks demonstrate that ForceLDM significantly outperforms state-of-the-art baselines, generalizes effectively to novel objects, and remains robust to environmental perturbations.