Physics-Guided Curriculum Learning for Nonlinear Mixture Modeling and Discovery
Abstract
Many functional products are formulations designed to achieve properties that emerge only upon combining components. Discovering such behavior, however, requires learning the intermolecular interactions that drive non-ideal mixing. Yet mixture datasets are dominated by near-ideal systems; for example, an Arrhenius mixing rule reproduces the composition-dependent property ranking for 82.3\% of entries in a large fluid-viscosity dataset. Consequently, aggregate evaluation metrics in standard machine-learning benchmarks can conceal model failures in rarer, highly non-ideal regimes that are often the most valuable targets for discovery. Here we introduce REMIX (Representation of Excess in MIXtures), a physics-guided framework that anchors predictions to expected mixing behavior and focuses learning capacity on the excess contribution arising from molecular interactions. A composition-guided curriculum progressively shifts learning from pure-component behavior towards increasingly interaction-dominated mixtures, while hard-sample mining concentrates training on persistent errors. For mixtures where classical mixing rules fail, REMIX reduces prediction error approximately threefold relative to the rule and by 25–40\% compared with the strongest machine-learning baseline. In active discovery campaigns, REMIX-guided Bayesian optimization identifies all top-50 mixtures within approximately 300 acquisitions, compared with 800 for the leading baseline. These results demonstrate that learning what established physical models miss can focus both modeling and experimentation on formulations with the greatest discovery value.