CompJudge: Fine-Grained Comparative Evaluation using Multimodal LLM for Subject-Driven Generation
Abstract
Subject-driven generation (SDG) evaluation using multi-modal LLMs (MLLMs) provides better interpretability and stronger alignment with human judgment than traditional embedding-based methods. However, MLLM-as-a-Judge scores are limited in differentiating between error and error-free SDG images. In this paper, we propose CompJudge, which overcomes these limitations by introducing comparative capabilities to MLLM-as-a-Judge and recalibrating scores based on comparison results. To rigorously evaluate the correctness of SDG evaluators, we introduce CompILIAS, a diagnostic benchmark of triplets (reference, identity-preserving image, identity-degraded image) with verified binary ground-truth labels indicating whether the SDG image preserves the reference identity. We demonstrate improvements of SDG evaluation over existing approaches: 8.9pp on CompILIAS and 10.62--15.98pp on the DreamBench++ dataset. These improvements consistently generalize across diverse MLLM backbones and across subject categories.