ClinMAS: A Knowledge-Grounded Multi-Agent Simulation Framework for Evaluating Clinical Reasoning in LLMs
Abstract
Clinical diagnosis begins with doctor-patient interaction, during which physicians iteratively gather information, order examinations, and refine diagnosis through patients’ responses. This dynamic reasoning process is poorly represented by existing LLM benchmarks that focus on static question-answering. To mitigate these gaps, recent methods explore dynamic medical frameworks involving interactive clinical dialogues. Although effective, they often rely on limited, contamination-prone datasets and lack granular evaluation. In this work, we propose ClinMAS, a knowledge-grounded multi-agent simulation framework for evaluating clinical reasoning in LLMs. Grounded in a disease knowledge graph, our method dynamically generates patient cases and facilitates multi-turn interactions between an LLM-based doctor, an automated patient agent, and an examiner agent. Our evaluation goes beyond diagnostic accuracy by incorporating fine-grained efficiency analysis and rubric-based assessment of diagnostic quality. Experiments show that ClinMAS effectively exposes critical clinical reasoning gaps in state-of-the-art LLMs, offering a more nuanced and clinically meaningful evaluation paradigm.