Learning When to Denoise: Optimizing Asynchronous Schedules for Latent Diffusion
Bingshuo Qian ⋅ Xiang Cheng
Abstract
Recent latent diffusion methods jointly model multiple representations of an image, denoising them with asynchronous noise schedules in which one representation temporally leads another. However, existing asynchronous schedules are often hand-designed, relying on grid search over simple function classes. We propose to learn the asynchronous schedule directly. We parameterize the schedule as a convex monotone bijection $f : [0,1] \to [0,1]$, where convexity together with $f(0)=0$ and $f(1)=1$ guarantees the desired leading-representation property by construction while still admitting a wide family of schedule shapes. A short probe stage jointly optimizes the schedule and diffusion model under a mild speed regularizer; the resulting smooth schedule is averaged over the stable probe window, frozen, and reused for full training, introducing no additional overhead during the main training stage. On ImageNet 256$\times$256, our learned schedule reaches unguided FID 2.37 at 1M iterations and FID 1.05 with AutoGuidance at 200 epochs, outperforming the best hand-tuned baseline trained for four times longer. Code is available at \url{https://anonymous.4open.science/r/LAS_public-01E4/}.
Chat is not available.
Successful Page Load