MM-SCALE: Evaluating Evidence-Grounded Moral Judgment in Vision-Language Models
Abstract
Vision-Language Models increasingly make socially consequential judgments from image-text inputs, yet current evaluations often stop at verdict-level accuracy: whether the final label is human-aligned, not whether it follows from the correct evidence. We introduce MM-SCALE (Multimodal Moral SCALE), a benchmark for evidence-grounded moral judgment built around a design choice: each image is paired with multiple action scenarios, so models must compare how different actions interact with the same visual context rather than score images in isolation. MM-SCALE contains 8,444 image contexts and 21,977 action scenarios, each annotated with a 5-point moral acceptability rating and a modality-grounding label indicating whether the human judgment relied on text, image, or both. Three tasks evaluate whether models (i) assign calibrated scalar moral scores, (ii) preserve human orderings within a shared visual context, and (iii) ground rationales in the evidence source humans found critical to judgment. Across models, aggregate metrics overstate alignment: NDCG@5 remains high across models, while within-image pairwise accuracy remains only 0.53--0.62 even on scenario pairs whose human mean ratings differ by at least one point. CoT has inconsistent effects on calibration and does not reliably improve within-image ordering. Evidence-grounding evaluations further show that models often describe visual context without making it critical for the verdict. These results reveal an evidence-grounding failure in current VLMs and position MM-SCALE as an evaluation testbed for moral judgment under shared image contexts. This dataset includes potentially harmful and sensitive visual content. Images are intended solely for evaluation.