Characterizing the Aesthetic Defaults of Generative Image Models
Abstract
Contemporary text-to-image (T2I) models are widely observed to exhibit a recognizable, model-specific look even when prompted without any stylistic instruction (e.g., the so-called ``Midjourney look''). What such a default aesthetic actually is, however, has not been characterized: dominant T2I evaluation paradigms score realism, prompt alignment, or pairwise preference, but cannot say which concepts drive separation, how they shift across releases, or how they relate to real-image distributions. We address this with a concept-level decomposition of generator outputs along an interpretable, style-trained sparse vocabulary. We introduce LouvreSAE, a sparse autoencoder over CLIP image embeddings whose dictionary is shaped to fall on stylistic rather than object features, and LatentAesthetic, an evaluation methodology that uses it to estimate, ground, and compare generators' default aesthetics under content-controlled, style-neutral prompts. Applied to 26 generators spanning four years, our methodology surfaces stable per-generator profiles, a uniform pull toward lifestyle and editorial photography over art-historical imagery, and a steady cross-family contraction of the aesthetic envelope, supporting hypotheses of an algorithmic monoculture in image generation. Code, weights, and data are available at https://hf.co/datasets/neurips-subm/anonymous for the evaluation of future models.