REAL-MED: Benchmarking LLM Agents on Real-World Medical Tasks
Abstract
Real-world medical work extends far beyond answering isolated clinical questions: it often requires clinicians to process complex materials, interact with information systems, and produce structured deliverables. As AI systems evolve from medical chatbots into medical agents, evaluation should likewise move beyond \emph{answer correctness} toward the \emph{quality of end-to-end workflow deliverables}. We introduce \textsc{Real-Med}, a benchmark for evaluating AI agents on real-world medical tasks. \textsc{Real-Med} contains 405 expert-designed tasks across four categories: Clinical Support, Patient Management, Pharmacy Management, and Medical Education & Research. Each task is grounded in a realistic scenario with concrete materials, tool interfaces, and explicit deliverable requirements, such as clinical documents, medication plans, spreadsheets and review reports. To support rigorous evaluation, each task is annotated with fine-grained expert-written rubrics, averaging 15.4 independent rubric items per task, and every task is cross-validated by at least three medical experts. We evaluate 11 strong mainstream models under a unified harness with the same resources, tools, and execution protocol. Our results show that current agents can often satisfy surface-level task requirements, but still struggle with detail-sensitive medical reasoning, dosage-calculation traps, shallow research, formatting constraints, and reliable final delivery. These findings reveal a substantial gap between medical QA ability and robust execution in realistic medical workflows. Code is available at \url{https://anonymous.4open.science/r/REAL-MED}.