Beyond the Dirac Delta: Mitigating Diversity Collapse in Reinforcement Fine-Tuning for Image Generation
Jinmei Liu ⋅ Haoru Li ⋅ Zhenhong Sun ⋅ Chaofeng Chen ⋅ Yatao Bian ⋅ Hongdong Li ⋅ Bo Wang ⋅ Daoyi Dong ⋅ Zhi Wang
Abstract
Reinforcement learning (RL) has emerged as a paradigm for fine-tuning large-scale generative models, such as diffusion and flow models, to align with complex human preferences and user-specified tasks. A fundamental limitation remains *the curse of diversity collapse*, where the objective formulation and optimization landscape inherently collapse the policy to a Dirac delta distribution. To address this challenge, we propose **DRIFT** (**D**ive**R**sity-**I**ncentivized Reinforcement **F**ine-**T**uning for Versatile Image Generation), an innovative framework that systematically incentivizes output diversity throughout the on-policy fine-tuning process, reconciling strong task alignment with high generation diversity to enhance versatility essential for applications that demand diverse candidate generations. We approach the problem across three representative perspectives: i) **sampling** a reward-concentrated subset that filters out reward outliers to prevent premature collapse; ii) **prompting** with stochastic variations to expand the conditioning space, and iii) **optimization** of the intra-group diversity with a potential-based reward shaping mechanism. Experimental results show that DRIFT exhibits clear Pareto dominance in task alignment and generation diversity, achieving 7.19\%$\sim$93.40\% higher diversity at matched alignment and 13.23\%$\sim$60.13\% higher alignment at matched diversity.
Chat is not available.
Successful Page Load