Factored Generative Models through Mechanism Diversity
Abstract
A generative model is factored when each latent dimension independently controls one factor of variation: changing a single latent predictably changes one semantic attribute while leaving the rest unchanged, enabling controllable generation, compositional generalization, and reproducible representations. Existing approaches either constrain the latent distribution, e.g., requiring it to shift with an auxiliary variable, or regularize the model, e.g., sparsity, quantization, or Hessian penalties, neither of which matches how modern conditional generative models actually work: a conditioning signal u reshapes the generator g(·, u) while the latent prior stays fixed. We prove an identifiability theorem showing that generating mechanism diversity, the natural variation that arises when u sufficiently reshapes g, is sufficient for the model to be provably factored, with no parametric assumption on the latent distribution. To actively enforce this condition, we propose Mechanistic Contrastive Learning (MCL), a model-agnostic contrastive objective over generator Jacobians. Empirically, MCL achieves state-of-the-art latent concept disentanglement on three benchmarks equipped with a latent diffusion model, and improves prediction quality and zero-shot cross-task transfer in latent action world models with a 1.4B video generative model as the backbone.