Large Language Models Enable Training-Free Retention Time Prediction Across Chromatographic Systems
Woojae Kim ⋅ Seokho Kang
Abstract
Retention time (RT) provides complementary information for compound annotation in liquid chromatography–mass spectrometry (LC-MS). The impracticality of experimentally measuring RT for every compound under each chromatographic system has motivated the development of RT prediction methods. Since the dependence of RT on chromatographic conditions limits the transferability of prediction models across different chromatographic systems, each system typically requires a separate prediction model trained on a sufficiently large RT-labeled dataset. Although transfer learning and multitask learning alleviate this difficulty, RT prediction remains challenging for chromatographic systems with scarce RT data. Here, we propose $\textbf{RT-ICL}$, a large language model (LLM)-based method that leverages in-context learning (ICL) to enable training-free RT prediction for data-scarce chromatographic systems. To predict the RT of a query compound under a target chromatographic system, we prompt the LLM with task instructions and LC conditions for the target system, molecular information about the query compound, and structurally similar reference compounds with RTs measured from the same system. This method can be applied to different chromatographic systems without requiring task-specific fine-tuning. We demonstrate that the proposed method achieves performance superior or comparable to that of transfer learning and multitask learning baselines on various data-scarce chromatographic systems.
Chat is not available.
Successful Page Load