Beyond a Single Score: An Audit of Aesthetic Evaluation in Text-to-Image Pipelines Across Subcultural Visual Languages
Abstract
Text-to-image (T2I) models are increasingly used by online creative communities to generate imagery in the shared visual languages those communities have built, from Cottagecore's pastoral warmth to Cyberpunk's neon dystopias. The aesthetic predictors embedded in T2I pipelines filter training data and score generated outputs, yet whether different predictors agree on which visual languages are "high-quality" is largely untested. We present SubCulture-2.6K, a dataset of 2,610 Stable-Diffusion-generated image–prompt pairs spanning six subcultures (Cottagecore, Cyberpunk, Dark Academia, Goblincore, Synthwave, Y2K), released with dominant-color palettes, CLIP ViT-L/14 embeddings, and per-image scores from four deployed predictors (LAION-Aesthetics V2, PickScore, ImageReward, HPS-v2). Using this dataset we report four findings. First, subcultures are visually separable in the representations these predictors read from, so any cross-subcultural score difference reflects a value judgment, not a perceptual failure. Second, LAION-Aesthetics V2 imposes a large ordered preference across subcultures, with Y2K rendering retained at the standard quality cutoff at a small fraction of the rate of the most-favored subculture. Third, peer predictors share LAION-Aesthetics V2's dispreference for Y2K but disagree sharply on which subculture they reward most, and they cluster much more closely with each other than with LAION-Aesthetics V2. Fourth, an interpretable-feature regression accounts for roughly half of the LAION-Aesthetics V2 effect, leaving a sizeable residual subcultural taste. Because LAION-Aesthetics V2 sits both upstream and downstream of recent Stable Diffusion variants, this audit shows that aesthetic predictors are not interchangeable: the choice of which one to use is itself a choice about whose visual language is rewarded.