Efficient Diffusion Policy Fine Tuning with Latent Noise Representation Bridging
Yuanchu Liang ⋅ Yue Yang ⋅ An-Chi He ⋅ Hanna Kurniawati
Abstract
Pre-trained diffusion and flow policies have emerged as powerful backbones for robotics, yet adapting them to novel tasks via reinforcement learning (RL) remains computationally unstable, sample inefficient and lacks theoretical guarantees. Prior approaches that fine-tune the latent noise space bypass costly backpropagation through time but struggle with "noise aliasing," where determining the optimal noise input to the pre-trained generative model requires repeatedly refitting separate Q-value networks. This paper resolves these challenges by introducing a novel theoretical framework demonstrating that imposing a linear structure on the environment induces a Latent Noise Linear Markov Decision Process (LNL-MDP). By treating the pre-trained model as a virtual dataset, our LNL-MDP fine tune framework falls under the FineTuneRL theory with near-optimal guarantee. To translate this idealized theory into a practical methodology, we derive the Bridge Equation which eliminates the noise aliasing problem. Leveraging this insight, we propose Latent Representation Bridging for Diffusion (LARBRID), an algorithm that efficiently learns the representation space from the LNL-MDP and optimal noises to steer the pre-trained model to act optimally in given tasks. We demonstrate that LARBRID is exceptionally sample efficient, yielding up to $2\text{-}3\times$ performance increases on challenging control problems across a diverse suite of robotic locomotion and manipulation tasks.
Chat is not available.
Successful Page Load