LIBERO-PeRM: Benchmarking Personalized Robotic Manipulation
Abstract
Personalization is essential for general-purpose robots to transition into household environments and achieve true practicality. While general robotic manipulation has been extensively studied, personalized robotic manipulation (PeRM) lacks a unified formulation and remains significantly underexplored. In this paper, we identify user preference as the cornerstone of PeRM and introduce LIBERO-PeRM, a novel benchmark designed for personalized preference learning in embodied agents. Specifically, LIBERO-PeRM highlights five key research topics in personalization: 1) how policies follow preferences in zero-shot settings; 2) how to learn single preferences from demonstrations; 3) how efficiently policies acquire preferences as demonstrations increase; 4) how learned preferences generalize across layouts and preference settings; and 5) how policies compose multiple preferences and resolve conflicts. To this end, we propose a unified formulation of personalized manipulation with three preference levels that jointly cover common daily personalization scenarios. Building on this formulation and the LIBERO framework, we develop an extensible procedural generation pipeline that generates paired demonstrations with and without preferences. For benchmarking purposes, we create five task suites (312 tasks in total) with corresponding high-quality demonstrations, rich metadata, and preference satisfaction metrics to probe the aforementioned research topics. Experimental results show that current VLA policies struggle with zero-shot personalization, while SFT makes preferences learnable; preference learning benefits from mixed training, broader layout diversity, and same-layout no-preference data, but generalization, decomposition, and conflict resolution remain difficult.