Designing Information-Dense Synthesis to Accelerate Molecular Discovery
Kasper K Jakobsen ⋅ Eli N. Weinstein
Abstract
Machine learning holds dramatic potential for accelerating molecular discovery, as models sequentially generate molecules, plan experiments, and learn from the results. However, many of the most outstanding challenges in chemistry and biology require discovering molecules with properties that are very rare. In this sparse setting, existing active machine learning methods often offer little improvement over random guessing. We propose a method to efficiently search vast regions of molecular space using algorithmically controlled stochastic synthesis. Rather than design, make and test individual molecules, we design and make complex mixtures, test them as a pool, then deconvolute the molecule-activity map from the measurements. We maximize information yield using Bayesian experimental design, modeling the experiment as an encoding of the unknown molecule-activity map, and jointly optimizing a decoder to reconstruct the optimal molecule(s). Our approach can dramatically accelerate learning on sparse molecular discovery problems, by extracting more information per experiment. Theoretically, it can reduce the number of experiments required to find the optimal molecule among $d$ candidates from $\mathcal{O}(d)$ to $\mathcal{O}(\log d)$ or $\mathcal{O}(1)$. Empirically, on simulated protein fitness landscapes, it requires $10\times-100\times$ fewer experiments to find an optimal molecule compared to existing Bayesian optimization methods.
Chat is not available.
Successful Page Load