Disentangling Dual Image References in Frequency Aware Diffusion Models for Personalized Generation
Abstract
Personalized image generation aims to synthesize text-driven images conditioned on reference images, while mainly casting the generation as image customization for foreground and style transfer for background. Previous arts of diffusion models suffers from the text misalignment with background for image customization and foreground for style transfer during the denoising process. Such facts, as we observed, rooted from the entanglement among hybrid frequency bands during the denoising process. In particular, the low-frequency band primarily captures color information, mid-frequency band encodes global structure, while high-frequency band corresponds to local texture. Besides, the early (high-noise) denoising stage mainly focuses on the low-and-mid frequency bands, while the late (low-noise) denoising stage progressively exhibits the high-frequency band, resulting in the dominance of low-and-mid over high frequency bands, especially when entangled within a single image reference, tends to reinforce each other at the early stage, hence the pitfall of overwhelming suppression of text prompt guidance. To address such salient limitation, in this paper, we study personalized generation based on dual references - customization and color and style reference - and propose a paradigm to disentangle these Dual image references within Frequency-aware Diffusion Models, dubbed Dual-FDM, to simultaneously tackle two crucial personalized image generation tasks: customization style transfer and color style transfer, by disentangling different frequency bands via mask strategy within frequency domain. For customization style transfer, we replace the mid-frequency band of the background in the style reference with that from the foreground of the customized reference. For color style transfer, we substitute the low-frequency band of the background in the style reference with that from both the foreground and background of the color reference. Both the substituted frequency bands are used as the key and value to reconstruct the query foreground and background of the denoised personalized image. We further compute the ratio of information entropy of the substituted frequency bands to adaptively modulate the denoising timesteps between the early and late stages. Extensive experiments validate the superiority of Dual-FDM over the state-of-the-art diffusion models for personalized image generation. Our code can be accessed from the supplementary material package.