GWScore: A structural diversity metric for consistent text-to-image generation
Abstract
Recent advances in visual storytelling have leveraged powerful text-to-image diffusion models to generate sequences of images featuring consistent subjects and styles. While such methods achieve impressive identity preservation across the sequence, they often produce scenes with limited structural diversity---images that are overly similar in subject pose and spatial layout. Existing evaluation protocols primarily measure identity consistency using metrics such as DreamSim, which can inadvertently conflate identity preservation with subject pose. The inability to independently evaluate structural diversity inhibits progress on identity preserving image generation methods. We address this limitation by introducing a new metric based on the Gromov Wasserstein distance to quantify structural diversity independently of identity. Our metric provides an interpretable and generalizable measure of spatial variation across generated images, enabling a more balanced assessment of identity consistency and pose diversity. Using this framework, we systematically evaluate several recent visual storytelling and subject-consistent generation methods. Furthermore, we evaluate the effect of latent initialization methods on structural diversity in several popular text-to-image models.