KnowVis: A Dual-View Benchmark for Diagnosing World-Knowledge Grounding in Text-to-Image Models
Abstract
Text-to-image (T2I) models have made substantial progress in visual realism, aesthetic quality, and instruction following. However, real-world prompts often go beyond explicit visual descriptions and require implicit facts, structured knowledge, and domain-specific commonsense. Existing evaluations mainly focus on explicit prompt-to-image semantic alignment or specific knowledge subdomains, lacking a unified and diagnostic framework for evaluating world-knowledge grounding in T2I models. We introduce KnowVis, a dual-view diagnostic benchmark for world-knowledge grounding in T2I generation. KnowVis is built on a hierarchical knowledge taxonomy with case-level annotations, and contains two complementary views: a Skill-Tree Benchmark for fine-grained atomic knowledge evaluation and a Generalized Benchmark for open-ended knowledge composition and expression. We further design a structured MLLM-based VQA judge protocol and validate its reliability on the Skill-Tree validation split, enabling scalable evaluation and diagnostic error attribution. The protocol uses Required VQAs to assess atomic knowledge correctness in the Skill-Tree Benchmark, and combines Required, Optional, and dynamically discovered knowledge expressions with human-judge collaborative evaluation in the Generalized Benchmark. Experiments show that KnowVis reveals notable gaps in world-knowledge grounding across existing T2I models and provides structured diagnostic signals for knowledge-intensive image generation.