AesGI-Bench: Benchmarking and Evaluating the Aesthetic Quality of AI-Generated Images via Large Multimodal Models
Abstract
The rapid advancement of generative models has led to a new era of AI-generated images (AIGIs). Early generation models often suffered from basic defects, such as geometric distortions and poor text-image alignment. Modern generation models have largely overcome these issues, narrowing the perceptual gap between AI-generated and real-world images, thereby shifting evaluation from basic correctness to high-level aesthetic excellence. In this paper, we introduce AesGI-Bench, a comprehensive and multidimensional benchmark for fine-grained aesthetic evaluation of AIGIs. Our benchmark covers 18 representative generation models, 19,485 images, and 117K pairwise comparisons from perspectives of visual aesthetic, technical quality, and style alignment. Based on AesGI-Bench, we propose AesGI-Assessor, a lightweight and human preference-aligned aesthetic evaluation model for AIGIs. Built upon Qwen3.5, AesGI-Assessor achieves a pairwise-to-score learning paradigm, using tie-aware pairwise loss for training and three lightweight score heads for multi-dimensional aesthetic score prediction. Although trained from pairwise human preferences, it learns to predict precise continuous scores for individual images, enabling both fine-grained image-level evaluation and model-level ranking. Furthermore, we explore the potential of large multimodal models (LMMs) as automatic aesthetic evaluators. Experiments show that LMM-based evaluators better align with human aesthetic preferences than conventional metrics. AesGI-Assessor achieves the highest agreement with human judgments, validating its effectiveness as a lightweight, preference-aligned evaluator for fine-grained AIGI aesthetic assessment. The database and codes are publicly available.