Limitations of EC-Labeled Contrastive Learning for Reaction-Enzyme Retrieval
Abstract
Contrastive learning has emerged as a leading computational method for linking enzyme sequences to chemical reactions. However, it has not been assessed whether such a linkage is genuine or how well contrastive models can learn to represent the data. In this work, we compare multimodal models trained on EC-labeled reaction and sequence data to unimodal models trained only on EC-labeled reaction data across a diverse set of reaction encoders. We find that unimodal models convincingly outperform multimodal models, that bespoke reaction encoders have marginal to no benefit over general encoder architectures, and that contrastive models do not greatly improve upon untrained baselines, facing reproducible performance caps on out-of-distribution queries. This leads us to believe that EC-labeled contrastive models do not learn a meaningful linkage between sequence and reaction space. Finally, we hypothesize that EC-labeled contrastive models rely heavily on pattern matching and are fundamentally limited by the curation of the EC system and the model's inability to learn the underlying chemistry.