Few-shot Task Learning via Compositional Concept Inference
Hanming Ye ⋅ Yiding Song ⋅ Yilun Du
Abstract
Humans can learn new tasks from just a few demonstrations. Prior methods like behavior cloning struggle to replicate this ability because they imitate demonstrations without learning the underlying concepts, such as goals, positional relations, and constraints. To address this challenge, we propose COIN, an approach that represents behavior as a composition of reusable concepts. A new task is learned by inferring its concepts from demonstrations then using a policy to implement them. We show that this approach is effective at few-shot learning across diverse domains. On object rearrangement, goal-oriented navigation, and human motion generation, COIN recombines previously seen concepts to solve novel tasks at inference time. On the LIBERO robotics benchmark, COIN achieves a $96.9$% average success rate on LIBERO-90, surpassing pretrained policies such as $\pi_0$ with roughly $10\times$ fewer parameters. This performance gap widens further when fine-tuning on unseen tasks with limited demonstrations, highlighting the advantage of concept inference over end-to-end policy adaptation in the few-shot regime.
Chat is not available.
Successful Page Load