Implicit Goal Conditioning via Value Disaggregation
Abstract
Long-horizon tasks are an important challenge for modern AI and robotics, yet pose a challenge for reinforcement learning. A historically popular approach has been to decompose a task into subgoals, and learn a subgoal-conditioned policy to execute each subtask in sequence. We introduce an alternative called Value Disaggregation (VaDar), which trains a single subgoal-independent policy to maximize a linear combination of per-subgoal critic values. Because subgoal information is needed only at training-time, VaDar removes the need for possibly costly or cumbersome subgoal-selection during policy deployment. Moreover, through careful experiments, we show that VaDar (1) is always competitive with, and frequently outperforms, goal-conditioning and other natural baselines and (2) succeeds on tasks with non-sequential goal structure and multiple optimization objectives. In particular, VaDar is the first algorithm to solve challenging, long-horizon tasks in the Robocasa suite, even when starting from a policy pretrained from behavior cloning with near-zero initial task success. Taken together, VaDar's success suggests that subgoal information benefits reinforcement learning by enhancing the efficacy of critic learning, whereas policy goal-conditioning is often unnecessary.