Diffusable Latents from Structure-Agnostic Distillation
Abstract
Distilling pretrained foundation models into an autoencoder bottleneck improves latent diffusability, enabling diffusion models to converge faster and reach higher sample quality. Standard distillation aligns the latent at each position to a co-located teacher feature, tying the latent layout to the teacher's. We show this constraint is unnecessary: aligning a single pooled image-level descriptor to the teacher's matches, and even slightly improves on, dense position-wise distillation. We compare first-order and relational pooled objectives across latent shapes and teacher modalities. Within a modality, first-order matching suffices and extends naturally to 1D token-sequence latents. Across modalities, first-order matching breaks down, but a relational objective based only on between-image similarities still improves diffusability when distilling a text encoder into an image autoencoder.