Optimal Transport on Graphs for Fine-Grained Text-to-Image Evaluation
Abstract
Recent text-to-image (T2I) models have made substantial progress in generating complex scenes with detailed objects and relations specified by text prompts. Yet, existing popular evaluation metrics, such as CLIPScore, remain largely focused on global semantic alignment, rendering them insufficient for precisely assessing whether individual objects and their relations are faithfully represented in the generated images. To address this gap, we introduce GMTI (Graph Matching for T2I) Score, a relation- and object-aware metric for fine-grained T2I evaluation. Specifically, our approach represents entities and relations as graphs corresponding to both the input prompt and generated image, using a CLIP model to align them, and then measures their consistency via optimal transport based graph matching. Across CLEVR-bind, NCD-Composite, NegBench, and our newly introduced synthetic benchmark IconScene, GMTI Score demonstrates inherent entity-binding capabilities and provides reliable and accurate evaluation for multiple T2I generation scenarios, such as relational and negation-based tasks.