ChemoLit: A Traceable Literature-Derived Dataset for Chemotherapy Outcome Research
Abstract
Published chemotherapy studies contain substantial evidence on treatment response, survival, toxicity, and regimen characteristics, but these results are distributed across articles and are rarely available in a consistent form for computational analysis. We developed ChemoLit, a literature-derived dataset that organizes published chemotherapy evidence while preserving study-level provenance and explicitly separating reported clinical evidence from simulated patient-level records. ChemoLit-Study contains 33 studies across six cancer groups and represents 31,117 participants, with structured fields for study design, treatment regimen, response, survival, toxicity, and source information. Each study is linked to its DOI, and the current release contains no duplicate study identifiers, duplicate DOIs, or missing DOI fields. Because aggregate publications cannot provide the individual patient records needed for many machine-learning experiments, we additionally constructed ChemoLit-Sim as a clearly labeled benchmark rather than presenting simulated observations as clinical data. The benchmark contains 6,000 simulated patient records derived from 24 treatment arms across 12 randomized studies. Across 30 simulation seeds, the mean absolute error between simulated and source objective response rates was 2.36 pp, while the corresponding error for reported time-to-event medians was 0.47 months. Under the fixed release seed, these errors were 2.21 pp and 0.44 months, respectively. ChemoLit therefore provides two complementary resources: a traceable study-level evidence table for analysis of published chemotherapy outcomes and a separate simulated benchmark for computational experiments that require patient-level structure. This separation is intended to support data readiness without converting literature-derived aggregate evidence into purported real patient data.