Deployment-Time Online Imitation Learning from Corrective Demonstrations
Abstract
Behavioral cloning (BC) is a powerful paradigm for learning robotic policies from demonstrations. However, most existing approaches operate in an offline setting and do not address a critical requirement for real-world deployment: the ability to continually improve policies after deployment. While interactive imitation learning methods such as DAgger enable improvement by collecting additional data, they rely on repeated retraining over an ever-growing dataset, leading to increasing computational cost and limited scalability. In practice, robotic systems must adapt continuously during deployment while remaining operational, making retraining on all accumulated data impractical. In this work, we introduce an online formulation of interactive BC, where the policy is updated in real time with expert corrections to address failure cases, with no access to past data. This setting presents a fundamental challenge due to catastrophic forgetting, which degrades previously acquired behaviors when learning from new data. To address this, we propose a novel method based on online ridge regression and random projections, enabling efficient policy updates without gradient-based retraining. We evaluate our approach on a range of manipulation tasks and show that it enables continuous and stable improvement after deployment, outperforming state-of-the-art BC baselines.