Generative Modeling With Learnable Noise Memory
Gabriel Nobis ⋅ Tim Korjakow ⋅ Luca Schmidt ⋅ Maximilian Springenberg ⋅ Thomas Demeester ⋅ Wojciech Samek ⋅ Manfred Opper ⋅ Tolga Birdal ⋅ Rembert Daems
Abstract
Mandelbrot--Van Ness fractional Brownian motion (Type I fBM) interpolates between the driving noise of continuous-time diffusion models and random straight-line paths via the Hurst index $H\in(0,1)$, which controls the correlation of its increments. Due to its non-Markovian nature, tractable fractional diffusion models are driven by a Markovian approximation of fBM (MA-fBM). However, the approximation used in prior work exhibits increasing error as $H\to1$ and therefore fails to preserve the random straight-line limit, making it unsuitable for designing a generative model that interpolates between Brownian motion and random straight-line paths. We propose $\varepsilon$-MA-fBM, an improved approximation that preserves this limiting behaviour and reduces the approximation error. Building on this approximation, we define a generative model driven by $\varepsilon$-MA-fBM that spans four regimes: subdiffusive dynamics ($H<1/2$), standard diffusion ($H=1/2$), superdiffusive dynamics ($H>1/2$) and random straight lines ($H\to1$). To reduce the training cost of identifying the most suitable regime for a given dataset and noise schedule, we jointly learn the Hurst index and neural network parameters by minimizing a derived KL divergence between the reverse process and its parameterization. On MNIST and CIFAR-10, Type I $\varepsilon$-MA-fBM outperforms Type II MA-fBM and all considered Brownian-driven baselines in FID and FD-DINOv2. The overall best performance is achieved with Type I and a learnable Hurst index saturating for MNIST at $H=0.9$ and for CIFAR10 at $H=0.82$.
Chat is not available.
Successful Page Load