Benchmarking Multimodal LLMs for Medicine Across African Languages and Context
Abstract
Current evaluations of LLMs for health rely predominantly on English-centric, text-only benchmarks that fail to capture the multilingual, multimodal, and culturally localized demands of real-world clinical practice, particularly across African health systems. To address this evaluation gap, we introduce an expert-generated, multimodal, multilingual benchmark for clinical visual question answering (VQA) and machine translation (MT) grounded in African healthcare contexts. The benchmark comprises 300 multimodal African-sourced clinical cases spanning five specialties, translated from English to 13 formal and local languages yielding 4,200 QA pairs. Evaluating frontier and open-weight models reveals substantial performance degradation: average multiple-choice question (MCQ) accuracy shows cross-lingual disparity between high resource and local African languages ranging from 5.4\% in closed weight models to 38.6\% in open weight models. These disparities persist for short answer question (SAQ) responses and machine translation (MT). Our findings demonstrate that clinical question-answering in English does not transfer reliably to African language and contexts, underscoring the need for culturally grounded, multilingual, multimodal benchmarks for health.