MMCompass: Diagnosing Position Bias in Generative Multimodal Reward Models
Abstract
Generative multimodal reward models are increasingly used to rank responses, guide alignment, and serve as automatic judges, making their reliability a central concern. However, existing multimodal reward benchmarks remain limited in annotation rigor, domain coverage, response model diversity, and robustness-oriented evaluation, making it difficult to assess reward models under realistic and diagnostic settings. To address these gaps, we introduce MMCompass, a new benchmark for evaluating generative multimodal reward models. MMCompass contains carefully verified preference pairs spanning 10 multimodal domains, with responses collected from 16 vision-language models. The benchmark is constructed through a multi-stage pipeline including AI-assisted difficulty filtering, blind multi-annotator verification, and expert review. To enable more diagnostic evaluation, we adopt three complementary metrics: overall accuracy, symmetric accuracy, and verdict guess rate, which jointly measure judgment quality, order robustness, and position bias. Evaluations on MMCompass show that position bias is widespread in our evaluated setting, even among strong multimodal judges and that standard accuracy alone often overestimates reliability. Beyond this diagnosis, we further provide CompassRM as a mitigation baseline for improving cross-order consistency. Built with dual-order supervised fine-tuning and the proposed GXPO (Group Cross Policy Optimization), CompassRM improves robustness over its instruct-model backbones on MMCompass, VLRewardBench, and MMRewardBench, while also improving overall judgment accuracy.