TokenRepel: Generating Diverse Image Sets Without Sacrificing Quality
Abstract
We present an inference-time approach for enhancing the visual diversity of image sets generated by text-to-image (T2I) models, while preserving, or even improving, the average image quality. Our method leverages the intrinsic geometry of hidden latent representations in pretrained models to diversify a batch of images generated under the same prompt. First, we introduce a token repulsion mechanism to encourage diversity by increasing the angle between the tokens of hidden features across different samples. This is achieved via a negative Riemannian gradient step defined over a distribution on the unit hypersphere. Second, to preserve the quality and prompt adherence of the output images, we encourage the updates to remain within the high-density region learned by the generative model. Third, we extend this constraint to the output of each intermediate layer, enabling more fine-grained iterative guidance. We evaluate our method on three representative large-scale pretrained T2I models: Flux1.dev (flow matching), SDXL (diffusion), and Flux2.Klein (few-step distilled flow matching). Experimental results demonstrate substantial improvements in visual diversity under identical prompts, while maintaining averaged individual sample fidelity. Our approach is simple, effective, and broadly applicable. It requires no external models to guide the generation process and introduces only modest additional computation overhead during the early stages of sampling.