RUBRIC-MME: Real-User Behavior-grounded Rubric for Multimodal Interaction Capability Evaluation
Abstract
Multimodal large language models (MLLMs) increasingly serve as always-on assistants in continuous, multi-turn, goal-driven user interactions, yet existing benchmarks evaluate them through pre-authored capability probes, single-turn correctness, or aggregate leaderboard scores. As frontier models saturate such benchmarks, a widening gap emerges between leaderboard performance and real interaction quality, and current evaluation infrastructure can neither characterize nor diagnose this drift. We introduce RUBUIC-MME, a multimodal multi-turn interaction benchmark designed to close this gap along three axes: an authenticity axis, a competence axis, and a diagnosability axis. To close the authenticity axis, we propose behavior-grounded benchmark construction, a privacy safe pipeline that extracts distributional regularities from real deployment logs, including scenarios, persona-goal pairings, intent typologies, multi-turn structures, and uses them as priors to guide public video filtering and a scenario–persona–intent–dialogue generation chain, yielding a fully synthesizable benchmark for two interaction modes, streaming video and multi-image sequences. To close the competence axis, we define a multi-level structured rubric that evaluates models simultaneously at the turn level, including grounding, relevance, and helpfulness, and the session level, including intent-shift recovery, cross-turn consistency, and goal completion, sliced by capability, scenario, and modality. To close the diagnosability axis, RUBRIC-MME further embeds an automated analysis pipeline that clusters failure modes into per-model diagnostic reports with actionable improvement directions, forming an evaluation–analysis–guidance loop. Evaluating representative frontier and open-source MLLMs, we find that session-level competence diverges sharply from turn-level accuracy, and that there are also significant differences in performance across different scenarios and capabilities. We position RUBRIC-MME as a step toward benchmarks that function as evaluation infrastructure rather than ranking artifacts.