Learning Actions, Not Noise: Direct Visuomotor Policy Learning on Action Manifolds
Abstract
We investigate action manifold policy, a generative imitation learning framework that directly learns to predict actions on the lower dimensional manifold in contrast to traditional flow or diffusion-based models that learn velocity fields or noise which reside in unstructured higher dimensions. We believe that training a network to predict actions seems more intuitive than predicting velocity or noise. We use a diffusion transformer-based architecture combined with Heun sampling to generate executable action sequences running in a receding horizon closed-loop control. We run this policy on the Push-T benchmark first, and then demonstrate its robustness on the more diverse Robomimic environment spanning a range of tasks across difficulties. We find promising evidence to support the effectiveness of the policy. Under state-based conditions and on relatively easy tasks, the policy outperforms flow matching policies and provides similar performance to diffusion policies. Under vision based and contact rich tasks, the action manifold policy outperforms both the baselines. The results pave the path for further investigation and provide preliminary evidence for action manifold policy as a viable candidate against its traditional generative imitation learning counterparts.