Disentangling Distribution Shifts in Polymer Property Prediction with Metadata, Physical Descriptors, and Group-Aware Evaluation
Hiroto Yokoyama ⋅ Takahiro Umemoto ⋅ Akiko Kumada ⋅ Masahiro Sato
Abstract
Predicting experimental polymer properties involves two distinct distribution shifts: generalization to unseen author-linked groups, and the systematic gap between simulated and experimental values; naive random splits conflate the former with predictive skill through group leakage. We combine hierarchical physics-informed descriptors (QM/FF/MD) with process and measurement metadata (Meta) from PoLyInfo, evaluated by group-aware cross-validation over author-linked record groups (Group\_ID). For gas diffusion coefficients, Meta reduces Normalized RMSE under Group\_ID splits from $0.92$ to $0.75$ (conditional group-cluster 95\% interval for the reduction $[0.107,0.271]$), reproduced across nonlinear models and descriptor classes; a permutation-based diagnostic is dominated by measurement-definition variables (gas type, temperature), with each process-variable marginal change $\leq0.002$. For density, MD-calculated values systematically underestimate experiments (NRMSE $0.88$); supervised calibration removes most of this discrepancy (best NRMSE $0.50$--$0.54$ per descriptor class; $0.56$ for a leave-one-group-out pooled offset), and the conditional 95\% interval for the Meta difference is $[-0.022,0.040]$. Measurement context thus emerges as the key metadata; within this cohort, much of the simulation--experiment gap is captured by a correctable empirical offset.
Chat is not available.
Successful Page Load