From Sample to Subset Construction: Coverage-Aware Curation of Robot Demonstrations
Abstract
Scaling robot demonstrations does not guarantee better imitation: redundant trajectories dilute learning, while rare, imperfect segments may encode critical transitions. Prior curation methods overlook such dependencies by scoring samples independently at the trajectory or segment level. We reframe robot demonstration curation as a coverage-aware sequential selection problem and propose PROSE, which selects data by maximizing marginal utility under a fixed budget. PROSE unifies three signals: (i) influence on closed-loop return, (ii) kinematic reliability, and (iii) coverage gain measuring novelty with respect to the retained set, allowing full trajectories and fine-grained segments jointly participate in scoring and filtering. Across three Robomimic simulation tasks and three real Franka tasks, PROSE achieves the best success rate on every task we evaluate against three state-of-the-art curators (CUPID, SCIZOR, Demo-SCORE), with average closed-loop gains of +15\% in simulation and +35\% in real-world. Ablations attribute the gain primarily to the coverage step.