Count Generalization Is Easy, Actually: Warp a Frozen Skill Instead of Retraining It
Paul Mattes ⋅ Wojciech Samek ⋅ Marc Toussaint
Abstract
Reinforcement learning policies for robot manipulation degrade when the number of objects grows beyond the training range. This is a covariate shift that changes the dimension of the state, and a frozen pretrained skill does not cross it zero shot. We propose WAVE (Warping Actions and Views of an Expert), which extends such a skill to many objects without updating it. WAVE proceeds one attempt at a time. It selects a single object and presents the frozen skill with a scene in which the other objects are absent. Around the skill it learns a reactive warping policy that reads a local view whose width does not depend on $N$ and emits a multiplier and an offset for every input and output channel of the skill as the scene changes under the skill's own placements. WAVE is task general and skill agnostic. It requires only that the skill was trained on the same embodiment and a goal pose for every object. A completion guarantee conditioned on how many objects an attempt places on average predicts the count at which generalization ends. We train WAVE on five cube arrangement tasks with at most nine cubes and evaluate it zero shot on up to thirty. Depending on the task we observe a two and a half to five fold extension of the object count at which a learned policy still completes the whole task, measured against the best results published by methods that learn the contact behaviour themselves. In comparison, fine tuning the skill, learning from the same local view without the skill, and an entity transformer trained from scratch all collapse within their training counts.
Chat is not available.
Successful Page Load