$\text{MP}^3\text{Bench}$: Do Large Language Models Excel as Medical Agents when Facing Situated Patients?
Gangao Liu ⋅ Daxing Zhao ⋅ Jianghua Huang ⋅ Ziding Liu ⋅ Yanjun Shen ⋅ 吴凯 ⋅ Qingsong Hua ⋅ Wei Lin ⋅ Peng Li
Abstract
While large language models (LLMs) excel at medical knowledge question-answering, conversational medical agents must navigate stochastic patient behaviors and clinical ambiguity. Existing benchmarks evaluate against uniform clinical standards, neglecting the idiosyncratic constraints of individual patients. We introduce $\text{MP}^3\text{Bench}$, a medical dialogue benchmark designed to decouple generic clinical fluency from situated effectiveness. Driven by 282 patient profiles (P$^3$: Profile, Pathology, Psychology) synthesized from real-world free-form medical dialogues and medical databases through dual-level clustering, $\text{MP}^3\text{Bench}$ clearly quantifies the gap between Generic Competence and Situated Competence. Evaluating nine LLMs revealed a noteworthy phenomenon: despite their achieving expert-level scores ($83\sim94$) on generic medical rubrics, their performance collapses by over 30 percentage points on case-specific situated rubrics. Our analysis reveals that current LLMs act as compliant medical encyclopedias, struggling to translate clinical fluency into personalized care. These findings highlight the critical need for advanced agents capable of hypothesis-driven reasoning and dynamic patient interaction.
Chat is not available.
Successful Page Load