Active Learning of Conditional Generative Models via the Transport Neural Tangent Kernel
Abstract
Many high-throughput scientific experiments produce distributional outputs, requiring conditional generative models to interpret the data, and to predict outcomes at new inputs. Acquiring data at new conditions is often expensive, motivating the need for automated experimental design. However, existing active learning frameworks assume scalar or vector-valued outputs. We first show that the Wasserstein test risk of a conditional generative task decomposes into a mean term and a shape term, demonstrating why mean-based active learning is blind to distributional shape. To capture both, we propose the \emph{transport Neural Tangent Kernel} (tNTK), a tractable kernel that measures how sensitive a generative model's predicted distribution is to its parameters. The tNTK recovers the standard NTK in the deterministic limit and upper-bounds the posterior Fréchet variance of the prediction under small parameter perturbations. It admits a closed form for affine Gaussian transport and tractable Monte Carlo estimators otherwise. Treating the tNTK as a Gaussian process covariance function, we greedily select inputs that minimize the total GP posterior variance across the pool, yielding a distribution-aware active learning strategy. Across four classes of conditional generative models, including conditional Monge gap and conditional flow matching on synthetic distributions, weather datasets and single-cell perturbation responses, the tNTK outperforms the baselines that use the conditional mean, demonstrating the value of distribution-aware active learning for generative-modeling tasks.