Diffusion on Hyperbolic Space with a Curvature Determined Confinement
Prithika Narayanan ⋅ Daria Godorozha ⋅ Mateusz Mikolajczak ⋅ Aditya K Shethwala
Abstract
Score based generative models need a forward stochastic process whose law converges to a reference distribution that can be sampled. Hyperbolic space $\mathbb{H}^n$ is a natural state space for hierarchical data, but Brownian motion on it is transient and has no invariant probability law. We analyse the confined diffusion $dX_t=-\tfrac{c}{2}\nabla r^2\thinspace dt+\sqrt{2}\thinspace dB_t$, whose invariant law $\mu_c\propto e^{-cr^2/2}\thinspace d\mathrm{vol}$ exists for every $c>0$, and trace its spectral structure into the design, training and validation of a generative model. The spectral gap $\lambda_1(c)$ sets the rate at which the forward process forgets the data, and the curvature threshold $c=n-1$ separates two regimes. Above the threshold the gap grows linearly, with $c-(n-1)\le\lambda_1(c)\le c-\tfrac23(n-1)+o(1)$. Below it the gap collapses faster than any power of $c$ as $c\downarrow0$, while an explicit lower bound keeps it positive at every $c$. On $\mathbb{H}^2$ the upper and lower bounds match: $\lambda_1(c)\asymp\sqrt{c}\thinspace e^{-1/(2c)}$. A commutation identity shows that for every $c$ the gap is attained in the first spherical harmonic sector, with the slowest radial mode relaxing exactly $c$ faster than the slowest angular mode. The first angular harmonic of the data is therefore the last component the forward process forgets, and the sampler must reconstruct it from the least informative endpoint. The sampler's score is the stationary drift plus a learned residual read in an orthonormal frame, which avoids the $\Theta(e^{r})$ amplification that makes untrained extrinsic network heads diverge in a one seed ablation. Simulated eigenfunction observables relax at the computed gaps to within four percent in dimensions two to four. On hierarchical data in $\mathbb{H}^2$ and $\mathbb{H}^3$, synthetic and built on the WordNet mammal taxonomy, the generation error at a fixed horizon is smallest at or within a factor two of the threshold and rises on either side of it, for different reasons. Below the threshold the forward process does not reach its prior, which a statistic of the forward process alone predicts before training. Above it the rise is consistent with a harder score estimation and sampling problem. Across repeated training seeds the intrinsic model is comparable to a wrapped normal baseline trained with exact score targets.
Chat is not available.
Successful Page Load