Flux: Online, Fine-Grained Data Scheduling for Training Machine Learning Interatomic Potentials
Yuanchang Zhou ⋅ Chen Wang ⋅ Hongtao Xu ⋅ Mingzhen Li ⋅ Guangming Tan ⋅ Weile Jia
Abstract
Large-scale training of machine learning interatomic potentials (MLIPs) increasingly relies on data parallelism over heterogeneous atomistic data. In this setting, atom count captures an important part of the per-rank workload, but structures with similar atom counts can still induce different graph workloads through variations in graph count, cutoff edges, higher-order geometric features, and graph-collation overhead. This makes conventional atom-count batching and offline load tables both incomplete and inflexible, especially when the model architecture, cutoff radius, or data mixture changes. We present Flux, an online workload-aware scheduler for distributed MLIP training. Flux estimates structure-dependent workload signals from lightweight physical and geometric priors during data loading, calibrates them against runtime and memory costs, and forms balanced, memory-feasible local mini-batches across data-parallel ranks. Without requiring precomputed graph metadata, Flux can be used as a drop-in replacement for the standard distributed sampler, making it flexible and portable across training setups. Across large-scale atomistic datasets and three representative MLIP architectures, eSEN, MatRIS, and AllScAIP, Flux improves training throughput by up to 2.23--4.76$\times$ without degrading convergence.
Chat is not available.
Successful Page Load