Learning Preference Representations for Preference-Conditioned Image Generation
Abstract
Modern image generative models produce high-quality, prompt-faithful images, yet the same prompt can correspond to different desired outputs for different users. We study \emph{preference-conditioned image generation}: generating images that follow a target prompt while reflecting a user's latent visual preferences inferred from liked and disliked histories. This problem is challenging because real preference histories are sparse and costly to collect, user tastes mix semantic and stylistic factors that are difficult to verbalize, and diffusion generators are not naturally designed to condition on multi-image preference histories. We propose \textsc{PrefGen}, a framework for learning structured preference representations and injecting them into diffusion-based generators. To address the supervision bottleneck, we construct a large-scale synthetic-agent preference dataset with dense, coherent liked/disliked histories, using it as scalable supervision for preference representation learning. We train a multimodal large language model with preference-oriented visual question answering and analyze its hidden states to separate two complementary signals: an intra-user embedding for liked-versus-disliked distinctions and a cross-history consistency embedding for stable user-level tendencies. To bridge MLLM representations and diffusion conditioning, we introduce a distributional alignment objective based on maximum mean discrepancy and inject the aligned preference signal through a lightweight cross-attention branch. We evaluate \textsc{PrefGen} with a tiered protocol: \textsc{PrefBench} as a synthetic in-distribution diagnostic, leakage-free Pick-a-Pic transfer as a controlled real-user benchmark, and user-in-the-loop studies with participant-curated histories as human-facing evidence. Across these settings, \textsc{PrefGen} improves preference alignment over strong personalization baselines while preserving prompt fidelity and competitive image quality.