RoboCoach: World Models as Active Coaches for Compositional Robot Skills
Abstract
Long-horizon manipulation spans a combinatorial space of reusable subtask com-positions, yet compounding errors and scarce end-to-end demonstrations make failures difficult to localize and costly to repair. We observe that compositionality not only compounds errors, it also provides natural boundaries for localizing and repairing them. Thus, we present RoboCoach, a world-model-guided specialization framework organized around a Route-Imagine-Diagnose-Improve (RIDI) loop. A router dispatches the first unfinished subtask to an addressable VLA expert, and CoachWorld, our action-conditioned world model, predicts closed-loop execution of the expert. A progress judge advances the route upon subtask completionor attributes the first timeout to the active subtask-expert pair. RoboCoach aggre-gates these diagnoses to pointedly allocate the demonstration budget and update the corresponding experts. Across two simulation suites and two real-robot platforms, weakness rankings inferred from CoachWorld track deployed performance, and our coaching method outperforms matched baselines under matched data and update schedules. With 150 additional subtask demonstrations, success rises from 13.3% to 75.0% on Franka and from 40.0% to 83.8% on AgileX. The specialized experts also transfer to four held-out compositions, whereas the vanilla VLA achieves 0% success. Together, these results show that world models can serve as active coaches efficiently, turning imagined failures into targeted supervision for modular policy improvement. Project Page: https://RoboCoach-AI.github.io/