VisualWorldBench: A Fine-Grained Multi-Task Benchmark for Evaluating Visual World Knowledge in MLLMs
Abstract
Visual world knowledge is a core capability of multimodal large language models (MLLMs), underpinning knowledge-intensive visual question answering and fine-grained visual reasoning. However, existing benchmarks fall short along two complementary dimensions. First, category coverage is limited: most benchmarks cover only a few thousand fine-grained entity categories, often within narrow domains, far from the open-ended scale of real-world visual entities. Second, task diversity is limited: most benchmarks rely on single task formats such as closed-set classification or attribute QA, and therefore cannot characterize the multi-dimensional structure of visual world knowledge. We introduce VisualWorldBench, a taxonomy-grounded multi-task diagnostic benchmark built on a hierarchical taxonomy of 21K+ leaf-node classes across 8 visual domains, comprising 32K+ instances across four complementary tasks: closed-set recognition, contrastive selection, precise localization, and open-world naming. To our knowledge, this is the largest category coverage among visual world knowledge benchmarks to date. VisualWorldBench reveals substantial capability gaps hidden under conventional benchmarks: the strongest closed-source model achieves only 80.4% average accuracy, and single-model spread across tasks can exceed 40 percentage points, indicating that visual world knowledge is multi-dimensional rather than monolithic. We further release VisualWorldBench-50K, a training resource covering 50K+ categories. Fine-tuning on VisualWorldBench-50K yields over 21 percentage points improvement on VisualWorldBench, with consistent transfer to public fine-grained recognition and visual knowledge benchmarks.