Breaking the Static: Dynamic Text Conditioning for Diverse Image Generation
Abstract
High-performance text-to-image~(T2I) diffusion models suffer from diversity degradation. In this study, we attribute this challenge to a fundamental temporal mismatch: while visual features evolve dynamically in a coarse-to-fine manner in the denoising process, the text condition remains entirely static. Through empirical analysis, we demonstrate that decoupling the text prompt to prioritize low-frequency components in the early denoising stage better aligns with the visual denoising process, while simultaneously enhancing generation diversity. Driven by this insight, we introduce Dynamic Semantic Interpolation (DSI), a training-free strategy that utilizes low-frequency semantics during the initial stages to foster the sampling space and progressively recovers the full semantics, ensuring text-to-image alignment. We also provide a rigorous theoretical justification for DSI from the perspective of conditional entropy, explaining its ability to maintain a broader early sampling space. Extensive experiments across diverse state-of-the-art models demonstrate that DSI significantly boosts generation diversity while preserving text alignment and aesthetic quality.