Compositional Policy Optimization with Language Models
Abstract
Large language models (LLMs) possess remarkable ability to understand natural language descriptions of complex robotics environments. Earlier studies have shown that LLM agents can use a predefined set of skills for robot planning in long-horizon tasks. However, the requirement of prior knowledge of the skill set required for a given task constrains its applicability and flexibility. We present a novel approach L2S (short for Language2Subtasks) to leverage the generalization capabilities of LLMs to decompose the description of a complex task in natural language into definitions of reusable subtasks. Each subtask is defined by an LLM-generated dense reward function and a termination condition, which in turn lead to effective subtask training and chaining. However, LLMs lack detailed insight into the specific low-level control intricacies of the environment, such as threshold parameters within the generated reward and termination functions. To address this uncertainty, L2S (1) enables LLM reflection feedback loop to improve task decomposition and code generation, and (2) trains parameter-conditioned subtask policies that perform well in a broad spectrum of parameter values. As the impact of these parameters for one subtask on the overall task becomes apparent only when its following subtasks are trained, L2S selects the most suitable parameter value during the training of the subsequent subtasks to effectively mitigate the risk associated with incorrect parameter choices. During training, L2S autonomously accumulates a subtask library from continuously presented tasks and their descriptions, using guidance from the LLM agent to effectively apply this subtask library in tackling novel tasks. Our experimental results show that L2S is capable of generating reusable subtasks to solve a wide range of robot manipulation tasks.