Sequence- and Ensemble-Aware Polymer Simulation: A Fidelity Choice for Automated Discovery
Gabriel Vogel ⋅ Jana M. Weber
Abstract
Polymer informatics and AI-based discovery workflows often represent copolymers by repeat-unit chemistry or average composition, although synthesized copolymers are ensembles of chains with distributions over length and sequence. Because simulation-derived labels are increasingly used in AI workflows, we ask how sequence architecture changes a label and how variation across sampled ensembles quantifies uncertainty in label differences. We introduce a workflow that samples copolymer ensembles from G-BigSMILES and preserve explicit monomer identities through all-atom melt simulation and analysis. We study styrene/$n$-butyl acrylate, holding chemistry, monomer composition, and mean chain length fixed while varying sequence architecture from random to statistically blocky to ideal diblock. Five independently constructed boxes per architecture and additional reruns quantify finite-ensemble and setup-plus-MD variation. Increasing blockiness increases local composition heterogeneity beyond this variation, but does not produce distinguishable changes in density, mean per-chain radius of gyration, or short-time chain-relative repeat-center mobility. The ideal diblock also produces a consistent change in local contacts between different chains. These results indicate that sequence resolution affects selected local labels, such as composition heterogeneity and interchain packing, while finite-ensemble and setup variation should be measured to avoid overinterpreting small simulation-label differences.
Chat is not available.
Successful Page Load