Do Diffusion Models Learn to Generalize Basic Visual Skills?
Abstract
Diffusion-based visual generative models are powerful, but understanding their generalization behavior remains challenging. Current evaluations rely on human preference scores for text-to-image models and distribution-level metrics such as Frechet Inception Distance (FID) and Inception Score (IS) for class-conditional models, but these do not test whether models learn basic visual skills or largely match training patterns. To overcome this gap, we introduce the Visual Generative Lab (VGL), a controlled experimental framework for understanding how diffusion models learn and generalize basic visual skills, including size, position, rotation, count, color, shape, and their composition. By generating synthetic training data for each skill from explicit rules, VGL enables precise measurement of generalization using task-specific, rule-based metrics. Training models from scratch on these data reveals three consistent patterns. (1) Models extrapolate rotation better than size, skill, or count on far-out-of-range queries; healthy rotation seeds match most extrapolation angles within a few degrees, although a non-trivial fraction of seeds collapse near the wraparound, while skill and position saturate at small distances outside training regardless of seed. (2) For compositional generalization, coverage of basic visual skill combinations matters more than dataset size. (3) Basic visual skills are learned jointly, so an out-of-distribution request in one skill can degrade the others even when they remain in range.