Directional Noise Conditioning for Diffusion Models
Abstract
Conditioning diffusion models on visual features is essential for controlled image generation and counterfactual reasoning across diverse applications, from creative design to medical imaging. Existing conditioning methods, such as timestep embedding and cross-attention mechanisms, suffer from limited control over feature accuracy, require substantial architectural modifications, and struggle to maintain conditioning precision throughout the denoising process, as the model's focus shifts toward fine-grained details in later steps rather than the desired conditioned features. We present a novel approach that addresses these limitations by learning to shift the initial noise distribution in the latent space. Our method assigns each visual feature a distinct direction in the latent space and applies proportional shifts during both training and inference, enabling precise control over continuous and discrete visual attributes. We introduce a combined loss function that explicitly enforces shift accuracy alongside denoising, ensuring consistent conditioning across all denoising steps. Furthermore, we extend this framework to datasets with hidden or inaccessible visual features by employing a Variational Autoencoder to extract latent representations in an unsupervised manner. This enables high-fidelity counterfactual generation on complex, real-world datasets where explicit feature annotations are unavailable or prohibitively expensive to obtain.