TerraVis: Towards Evaluation of World-Grounded Visual Consistency in Text-to-Image Generation via MLLM Workflows
Abstract
Recent text-to-image models have made substantial progress in realism, aesthetics, and prompt-image alignment. Yet visually appealing images can still violate real-world plausibility, exhibiting malformed object structures, impossible anatomy, physically implausible interactions, or impossible spatial relationships. These failures are not well captured by existing quality, aesthetics, preference, or alignment metrics. To address this gap, we introduce TerraVis, a framework for evaluating world-grounded visual consistency in generated images. TerraVis defines a structured taxonomy of world-consistency violations covering object-, interaction-, and scene-level failures, and evaluates images through a multi-stage workflow. Specifically, TerraVis uses an MLLM to first perform eligibility checking, determining whether an image is suitable for world-consistency evaluation. It then conducts taxonomy-guided violation detection over fine-grained violation types and classifies detected violations as minor or severe to derive an overall world-consistency score. Across diverse open-source and proprietary text-to-image models on two widely used benchmarks, TerraVis shows the strongest correlation with human judgments of world consistency among the evaluated metrics. Our results further reveal that models with high quality, aesthetics, preference or alignment scores can still exhibit frequent world-consistency failures. TerraVis therefore provides a complementary evaluation perspective, enabling fine-grained diagnosis of generated images and model comparison beyond existing evaluation dimensions.