HyperSkill: Multi-Modal Skill Learning on the Unit Hypersphere
Erdemt Bao ⋅ JunChen ⋅ Weijun Qin ⋅ Shaopeng Li ⋅ Ming Li ⋅ Mengchen Zhao ⋅ Ziqian Zeng ⋅ Cen Chen ⋅ HUIPING ZHUANG
Abstract
Language-conditioned manipulation requires skill representations that are both multi-modal and predictable: the same instruction may admit diverse executions, while the high-level planner must select a stable skill for low-level control. However, vector-quantized skill codes compress continuous variations into isolated symbols, introduce biased straight-through gradients, and lack geometry for relating semantically similar skills. We propose $\textbf{HyperSkill}$, a hierarchical skill-learning framework that models skills as continuous embeddings on the unit hypersphere. We use a learnable von Mises--Fisher mixture prior to capture skill-level multi-modality, with geodesic repulsion encouraging diverse mode coverage on the hyperspherical manifold. To connect training-time skill inference with inference-time planning, we introduce Geodesic Skill Alignment (GSA), which aligns posterior skills inferred from demonstration segments with prior skills predicted by a high-level planner using squared geodesic distance. Together, these components provide a geometrically structured skill interface that preserves execution diversity while enabling causal skill prediction at inference time. For low-level control, a Conditional Flow Matching generator produces smooth, multi-modal action trajectories conditioned on skill embeddings. Experiments on LOReL and Kitchen demonstrate that HyperSkill achieves state-of-the-art success rates on the evaluated benchmarks. Ablation studies further validate the importance of each proposed component.
Chat is not available.
Successful Page Load