WTF?! Simulation-Free Reinforcement Learning with Wasserstein-Tilted Flow Maps
Abstract
Fine-tuning aims to update a pre-trained flow-based generative model to improve the downstream reward of its generated samples. Existing methods typically frame this problem as sampling from a reward-tilted distribution, which arises as the solution to a KL-regularized reward-maximization problem. Here we take an alternative approach and introduce an optimal transport regularizer built directly from the pre-trained drift; we show that the resulting fine-tuning problem is equivalent to a deterministic optimal-control problem on the flow. Assuming access to a pre-trained flow map, we exploit this equivalence to devise a simulation-free reinforcement-learning algorithm for fine-tuning generative flows. We call the resulting framework Wasserstein-Tilted Flow Maps (WTF), the first end-to-end fine-tuning recipe native to flow maps. The output of our approach is itself a fine-tuned flow map, retaining few-step reward-aligned inference at deployment. Numerical experiments at text-to-image scale highlight both the efficiency and the efficacy of our approach. More broadly, our work argues that accelerated samplers such as flow maps are essential infrastructure for efficient post-training, and that the dominant KL-regularized formulation of fine-tuning is only one of many choices worth revisiting.