A Stratified Multi-Rater Evaluation of LLM-Based Virtual Standardized Patients with a Deployed Data-Generation Platform
Abstract
The Objective Structured Clinical Examination (OSCE) is widely used to assess clinical communication skills in medical education, requiring trainees to take histories and reason clinically with standardized patients (SPs). Large language model (LLM)-based virtual standardized patients (VSPs) offer a practical alternative to human SPs. These are typically built using persona-injected prompting, which casts a model into a clinical role via a single system prompt. However, existing evaluations of such VSPs often rely on a single backbone LLM, small human rater panels, or LLM-as-a-judge scoring without an utterance-aligned human-expert reference. Consequently, they are limited in their ability to identify the specific shortcomings of individual models. In this study, we evaluated four LLMs (GPT-4o, Claude 3.5 Sonnet, Llama-3.1 70B, and HyperCLOVA X)—spanning commercial, open-source, and Korean-developed models—against a human-expert reference across three Korean OSCE cases (hematuria, fever in pregnancy, and chronic cough). In a double-blind, fully crossed design, 17 senior medical-student raters scored every model utterance on three sentence-level and six encounter-level Likert metrics, yielding 7,480 data points. We analyzed the scores using a stratified linear mixed-effects model based on an 11-category history-taking and 3-level linguistic-form taxonomy. At the encounter level, three of the four LLMs matched the human-expert reference on Consistency and Comprehensibility. HyperCLOVA X additionally matched the reference on Engagement, whereas Llama-3.1 70B underperformed the human reference across all six metrics. Furthermore, we introduce a web-based OSCE platform that supports virtual practice across diverse clinical cases. This platform continuously collects student–VSP dialogues under research and privacy consent; the resulting dataset is available at https://huggingface.co/datasets/neurips-2026-cpx/neurips-2026-cpx.