Free Reward Tilting via Embedding Diffusion
Abstract
Steering a trained diffusion model toward a desired property typically requires reward gradients, additional network evaluations, or fine-tuning. We show that by jointly diffusing data with a semantic embedding, the per-reward cost of fine-tuning can be amortized into a single training run, after which each reward is available for free at inference time: the tilted denoiser is simply the original evaluated at shifted inputs, with no reward-specific training, gradients, or additional denoiser evaluations. A sufficiently rich embedding makes a wide range of semantic properties available for exact tilting. We present two complementary realizations: a jointly trained image-embedding model with decoupled noise levels, and direct application to existing factorized architectures such as Kandinsky, without any retraining. On ImageNet, the joint model achieves FID comparable to REG while enabling prompt-based image editing and adjustable aesthetic-score control, at less than 1\% overhead in parameters, training cost, and inference time.