Harmonizing Public Lung Adenocarcinoma Data for Reproducible Machine Learning
Abstract
Lung adenocarcinoma (LUAD) studies can draw on public clinical, molecular, treatment, and outcome data, but those data are distributed across repositories, releases, samples, and event tables. This fragmentation creates a practical problem for machine learning (ML): a model-ready table can silently lose provenance, mix patient and sample units, or discard the longitudinal information needed for later patient-specific modeling. In this study, we develop a real-patient data collection and harmonization protocol for LUAD/NSCLC and evaluate its data readiness with ML. We register eight authentic public resources, audit a TCGA-LUAD phenotype release containing 151 raw fields, and map high-value variables to a 44-field canonical schema across nine linked tables. We do not generate simulated patients or fill missing patient attributes with values from other people or study-level averages. To test whether the resulting clinical fields are sufficient for prediction, we use a deterministic 100-patient real TCGA-LUAD audit slice and construct censoring-aware 1-, 2-, and 3-year mortality tasks. Eligibility falls from 85 patients at 1 year to 46 at 3 years as follow-up requirements become stricter. A random forest reaches an AUROC of 0.768 at 1 year but 0.469 at 3 years, while adding sparsely observed smoking and history variables does not improve the 3-year benchmark. In a Cox model, stage is associated with mortality (hazard ratio 2.06, 95% CI 1.38–3.07, p<0.001). These results show that data readiness depends on identity, missingness, timing, and endpoint definition, not simply on the number of available variables. The released protocol, source manifest, field dictionary, analysis scripts, and benchmark outputs provide a reproducible base for later treatment-response and digital-twin studies using real patients only.