Steered Not Stirred: Effective Tilting in Generative Models using Surrogate Reward Mixing
Alec Sargood ⋅ Olivia Gallup ⋅ Georg Maierhofer ⋅ Noam Ghenassia ⋅ Paul Duckworth ⋅ Krisztina Sinkovics ⋅ Joey Bose
Abstract
Continuous-time generative models, from diffusion models, flow matching, and now flow maps, owe their practical value to high-fidelity generation and steerability. In many applications, a ground-truth latent oracle reward $R^* $ is unavailable or expensive to evaluate, e.g., human judgments or experimental measurements. Although the oracle reward is not directly observable, it can often be approximated through multiple cheaper surrogates that, individually, provide incomplete characterizations but collectively reveal substantial information about the downstream task. In this paper, we present a method for composing such surrogates to steer a pretrained model toward the distribution tilted by an unavailable $R^* $. Given a small feedback set containing either $K$-way oracle rankings or numeric oracle evaluations, we introduce Mixing And Reward Tilting for INference-time Improvement (MARTINI), which learns a surrogate composition for plug-and-play use with standard inference-time steering methods. Theoretically, we demonstrate that the learning objective for MARTINI controls the KL divergence between steering under $R^* $ and the learned composition of surrogates. Furthermore, we characterize the extent to which MARTINI can recover $R^* $, in both a distributional and reward sense, over the base model. Across FFHQ and ImageNet text-to-image alignment, MARTINI consistently shifts generation toward high-$R^* $ regions without ever querying $R^* $ at inference.
Chat is not available.
Successful Page Load