Identifying and Mitigating Diversity Collapse in Zero-Shot Personalization with I2I Editing Models
Abstract
Reference-conditioned image-to-image (I2I) editing models provide a promising zero-shot approach to personalized image generation by jointly conditioning on a reference image and a text prompt. We show that recent I2I editing models, including FLUX.1 Kontext and FLUX.2, achieve strong subject identity preservation and prompt alignment, but exhibit seed-wise diversity collapse: under the same reference image and prompt, different random seeds often produce highly similar outputs. We further find that this collapse is associated with over-alignment of latent-token trajectories during denoising. Based on this observation, we introduce Selective Hidden Perturbation (SHiP), a training-free inference-time intervention that perturbs only the attention-output representations of latent tokens during early denoising while keeping reference-image and text tokens unchanged. Across diverse subjects and prompts, SHiP improves seed-wise diversity on both FLUX.1 Kontext and FLUX.2 while largely preserving identity and prompt fidelity. On FLUX.1 Kontext, SHiP increases Vendi Score from 1.546 to 2.066 and LPIPS from 0.475 to 0.621, with only minor decreases in GPT-ID and GPT-Text. Code is available at https://anonymous.4open.science/r/i2i-personalization-diversity-0260/.