Manifold Drift in Flow Preference Optimization: A Root Cause of Reward Hacking
Abstract
Preference optimization is a standard alignment method for generative models, yet extending it to continuous-time dynamics remains non-trivial. In flow matching, reward-driven updates modify the transport trajectories but lack inherent constraints to pretrained data manifold, often pushing generated (terminal) samples off the pretrained support. We formalize this failure mode as manifold drift, a phenomenon where preference-aligned trajectories diverge from the pretrained support. Theoretically, we prove that while optimal flow matching exactly preserves the terminal manifold, existing methods like FlowDPO fail to do so whenever reward-driven updates contain components normal to the manifold surface. As a remedy, we propose ThermoDPO, a temperature-controlled objective that anchors pairwise preference optimization on preferred samples. Across temperature regimes, this objective connects rejection sampling fine-tuning and FlowDPO while upper-bounding a reconstruction-based surrogate for manifold drift. To counteract diminished signals at low temperatures, we further introduce a weighted variant, ThermoDPO-weighted, which we evaluate across all experiments for its superior trade-off between preference alignment and manifold preservation. Empirically, ThermoDPO-weighted demonstrates a better trade-off on synthetic manifolds than FlowDPO variants. On real-world image benchmarks, it improves OCR-oriented alignment over the pretrained model and FlowDPO variants while remaining competitive with strong baselines in held-out model-based and human evaluations without visible sample degradation.