ENCORE: Few-Shot Agentic Discovery of Manipulation Abstractions
Abstract
Language instructions often omit manipulation details such as grasp pose, contact sequence, and recovery behavior. These details appear in demonstrations, but standard imitation learning is unreliable with only a few trajectories, while coding-agent approaches are not designed to inspect long multimodal episodes. We present ENCORE, a framework that converts demonstrations into a structured evidence pack containing multi-view keyframes, gripper events, local frame strips, end-effector trajectories, and raw actions. A coding agent analyzes this pack, implements a policy against a fixed perception and action API, and improves it during a limited set of development rollouts. The resulting program is then frozen and evaluated on held-out initial states without access to the benchmark success signal. On LIBERO-PRO, ENCORE achieves 95.8% success across position and task perturbations, compared with 87.4% without demonstrations and 89.3% for the same released baseline code run on our inner model. We also deploy policies generated from five demonstrations on bimanual cube handover and cup inversion. These results show that structured demonstration evidence helps coding agents recover manipulation strategies that are difficult to infer from language or trial and error alone.