SafeDrug: A Benchmark Dataset for Safety-Critical Pharmacological Reasoning in LLMs
Abstract
Large language models (LLMs) have shown strong performance on biomedical tasks, yet evaluating their reasoning in safety-critical pharmacological contexts requires large-scale, structured datasets. We present SafeDrug, a benchmark dataset for systematic evaluation of pharmacological reasoning across multiple task dimensions. SafeDrug integrates heterogeneous sources, including adverse drug event records, drug–drug interaction knowledge, literature-derived evidence, and population-specific information such as child, adult, and older adult groups, into a unified reasoning-oriented QA format. The benchmark comprises two complementary subsets. \textbf{SafeDrug-large} provides broad coverage at scale, while \textbf{SafeDrug-small} is a high-quality human-annotated subset for precise evaluation. It supports tasks such as prediction, risk assessment, mechanism explanation, evidence-grounded QA, and drug replacement across both single-drug and multi-drug settings. Our evaluation shows that LLMs capture outcome-level associations but struggle with mechanistic reasoning, evidence alignment, and population-aware predictions. SafeDrug provides a foundation for reproducible benchmarking and advances research on evidence-grounded and population-aware pharmacological reasoning in LLMs.