Derivative-Weighted Capacity Bounds for Kernel Score Estimation
Thanh Nguyen
Abstract
In score-based generative modeling, nonparametric score estimation is the statistical problem that governs the fidelity of reverse-time sampling. Kernel score estimators offer a principled operator-theoretic framework, but they rely on feature derivatives whose sampling variability escapes the control of the conventional design effective dimension. We establish capacity-dependent convergence bounds by decoupling the fluctuations of the empirical operator from those of the score-matching right-hand side and retaining a derivative-weighted covariance trace for the latter. Under bounded features, polynomial capacity conditions, and a source condition of order $0 \le r \le \frac{1}{2}$, spectral regularizers of sufficient qualification achieve the $L^2(\rho)$ rate $\widetilde{O}_P(M^{-(r+1/2)/(2r+1+\alpha)})$, where $M$ is the sample size and $\alpha \in (0,1)$ is the derivative-weighted capacity exponent, strictly improving on capacity-independent baselines. The same rate holds in expectation and for a Nystr\"om variant whose rank need only resolve the design complexity. For diagonal and curl-free kernels on the flat torus with eigenvalue decay $|\omega|^{-A}$, the capacity exponents are $\nu = d/A$ and $\alpha = (d+2)/A$: differentiation costs exactly two effective dimensions, a derivative penalty $\alpha - \nu = 2/A$. For every such kernel and source order, a Fano lower bound over the matched Sobolev log-density class shows that the rate is minimax optimal, in probability and in expectation, up to logarithmic factors. Because the exponents are invariant under two-sided density bounds, they persist along torus heat flow, which yields score-estimation guarantees at each diffusion time and, within the matched smoothness window, expected-risk bounds that are uniform in time.
Chat is not available.
Successful Page Load