MMGO-Bench: A Multimodal Graph Optimization Benchmark
Abstract
We present MMGO-Bench, a controlled benchmark for evaluating whether vision-language models (VLMs) can solve graph optimization problems directly from graph images, instantiated with the weighted shortest-path problem. MMGO-Bench contains 90 graph instances across three difficulty tiers and systematically varies graph layout, node-label representation, textual support, prompting strategy, and answer format. We evaluate four open-weight VLMs (Gemma 3 4B, Qwen2.5-VL 7B, LLaVA, and MiniCPM-V) and Claude Haiku using separate model calls for path generation and numeric-weight prediction. Across 43,200 multimodal and 1,080 text-only evaluations, performance declines sharply with graph difficulty; with alphabetic labels and short descriptions, path accuracy falls from 38.3% on easy graphs to 12.8% on hard graphs. Differences across layout conditions reach 16.9 percentage points in path accuracy on identical underlying graphs, while complete textual edge descriptions approximately double aggregate path and numeric accuracy. Node-label representation is also associated with a large drop in numeric prediction: under long descriptions, numeric labels reduce Claude Haiku's numeric accuracy from 94.7% to 53.2%, although we cannot rule out answer-extraction effects. Finally, prompting effects vary substantially across models, with chain-of-thought substantially improving Gemma 3 4B while sometimes reducing Claude Haiku's numeric accuracy. These results indicate that multimodal graph reasoning depends strongly on visualization and prompting choices. We release the benchmark, images, ground truth, and evaluation code at https://github.com/GraceJulius/MMGO-Bench.