Distilling Conditional Image Generators into Spatial Effect Maps
Eun Ryung Lee ⋅ Seyoung Park
Abstract
Conditional image generators learn $P(Y\mid X)$ for images given clinical, demographic, or experimental covariates, but their covariate effects are encoded implicitly in samples, denoisers, gradients, or score fields. We propose \emph{generative effect distillation} (GED), a framework that treats a conditional image generator as a teacher and projects its conditional law onto an interpretable image-response regression student $m_{\phi}(x,s)=\beta_0(s)+x^\top\beta(s)$ whose parameters are spatial effect maps. The teacher-to-map target is defined as a statistical functional of the teacher conditional law under a target design and discrepancy, rather than a set of synthetic images or a compressed generator. We develop mean, effect-gradient, and structured score distillation objectives; establish projection identities linking the distilled map to covariance-weighted teacher means and, under Stein conditions, to average teacher gradients; and derive finite-query risk bounds that separate teacher bias, spatial approximation, query design, Monte Carlo, and student optimization error. For the associated finite-basis coefficient-matrix oracle, a Gaussian lower bound matches the query-design and Monte Carlo stochastic dependencies. Across eleven blocks spanning synthetic, semi-real, pretrained SDXL-Turbo/FFHQ-256, CelebA conditional-diffusion, and face-image analyses, GED attains the lowest standardized integrated squared error (SISE) among the direct-regression and augmentation baselines in every block, with paired Student $t$-tests rejecting equality with the strongest non-GED baseline at $p<0.05$; on a CelebA conditional residual diffusion teacher, score-access GED via a single-step Tweedie denoising procedure attains SISE $0.009$, a factor of $25$ below the strongest non-GED baseline.
Chat is not available.
Successful Page Load