MapPolicy: Structure-Aware Imitation Learning for Robot Manipulation via Physically Constrained Scene Map
Abstract
Robot manipulation under partial observability requires policies to act from incomplete visual evidence while still reasoning about object structure, spatial layout, and physical relations. Existing imitation learning methods mainly encode currently visible observations, which makes them brittle when task-critical parts, contact regions, or support relations are occluded. We introduce Scene Map, a structured scene representation that models object shape, part structure, spatial layout, and physical constraints as a graph of manipulation-relevant primitives and relations. Scene Map turns occluded but task-relevant structure into explicit prediction targets and provides structural priors for inferring unseen components from partial observations. Building on Scene Map, we propose MapPolicy, an imitation learning framework that fuses structured map features with visual and robot state features for action prediction. MapPolicy uses structure-modulated graph attention to propagate global structural and spatial information through primitive features, and a physical constraint loss to regularize learned features with geometric and kinematic consistency. Across 49 manipulation tasks in three simulation benchmarks and five real world tasks, MapPolicy consistently improves over strong imitation learning baselines, with especially clear gains under severe occlusion. Our code will be released publicly to facilitate reproducibility and further research.