SwiftFlow: An Efficient One-Step Policy Learning via Improved Mean Flow for Robotic Manipulation
Abstract
Generative modeling has become a mainstream paradigm for learning based robot manipulation research. Among these, diffusion-based policies achieve strong performance but typically require multiple iterative denoising steps to generate final action, leading to inefficient inference. Although flow-based methods eliminate the need for iterative denoising process, existing approaches often face challenges in maintaining stable optimization and high-quality action generation. In this paper, we propose SwiftFlow, an efficient one-step policy learning framework built upon improved mean flow for robotic manipulation. Specifically, SwiftFlow formulates policy learning as a conditional generation problem by jointly leveraging 3D point cloud observations and proprioceptive robot states. It stably learns mean velocity field through a principled regression objective and optimization process, enabling more reliable action generation while preserving efficient inference. To further regularize the learned implicit generation dynamics, we construct a mean velocity smoothness regularization term, which constrains rapid local variations in the learned mean velocity field. Since this regularization is imposed only during training, it hardly introduces additional inference overhead. Extensive experiments on Adroit and MetaWorld benchmarks demonstrate that our proposed method achieves stronger task performance together with highly efficient one-step inference, validating its effectiveness for robotic manipulation tasks.