Episode-Level Feedback for Data Selection in Continually Updated World Models
Abstract
A held-out prediction loss measures how well a world model fits observed transitions, but it need not predict whether a downstream planner completes a task. Native task success is closer to the intended behavior, yet its average does not show which fixed test episodes improved or regressed. Two updates can have the same success rate while solving different initial conditions. This paper proposes a study of data selection using retained, gained, and lost successes on a shared episode bank, compared against a matched uniform sampler. The design uses at least three tasks, multiple fixed seeds, paired rollouts, and native task endpoints. Before comparison, expert replay and action timing must be checked, and the complete planning system must demonstrate both task success and a response to changes in model scores. Nonzero success alone is insufficient: a strong behavior prior can leave little room for model updates to affect execution. RoboCasa and RoboMimic audits motivate these safeguards but establish neither the effectiveness of outcome-based selection nor equivalence between selection rules.