GenRM-Flow: Generators are Process-aware Reward Models in Flow Matching
Abstract
Aligning video generation models with human preference depends critically on the quality of the reward model that scores generated videos. Current video reward models, however, source their reward signal from a model's representational rather than generative capacity: they either repurpose vision-language models as pixel-space scorers, or train auxiliary scoring heads on top of frozen generator features. The generator's actual ability to predict velocities, denoise latents, and produce videos plays no role in scoring its own generations. We argue that this overlooks a structural property of offline preference learning. Preference optimization objective approximating an offline preference dataset mathematically targets a strictly defined KL-constrained optimal policy. By isolating the reward term from this analytical solution and mapping the exact generation log-likelihood to its flow-matching surrogate, the empirical reward emerges deterministically as the relative difference in continuous-time velocity-prediction errors. We formalize this equivalence into GenRM-Flow, a framework where a preference-tuned video generator inherently operates as a process-aware reward model directly in the noisy latent space. We demonstrate that this evaluation mechanism is structurally universal: distinct fine-tuning objectives, whether contrastive like DPO and IPO or non-contrastive like RWR, are merely alternative optimization pathways that share a unified empirical reward readout. Furthermore, the extracted process reward seamlessly substitutes external VLM scorers in Flow-GRPO. Operating purely on native latent outputs without auxiliary RL machinery, this unified generative reward drives measurable improvements in downstream video generation.