Revisiting the Evaluation of Molecular Property Understanding in LLMs from a Relative Perspective
Abstract
Large language models (LLMs) have been evaluated for their ability to predict molecular properties, with prior studies reporting limited performance under value-level metrics such as accuracy and MAE. We revisit this evaluation from a relative perspective. Rather than assessing only whether models predict exact property values, we examine whether their predictions preserve the relative ordering between molecules. Crucially, each molecule is predicted independently, with the model receiving one molecule per query and no access to any other molecule. The pairwise comparison is applied only during evaluation. Across five LLMs and eight descriptors on 3,657 molecules, the mean value accuracy is only 0.18 across models, yet the mean ordering accuracy reaches 0.75. This gap cannot be explained by SMILES string length alone. It is further supported by functional group modification experiments in which LLM-predicted property changes consistently follow the ground-truth direction. These results suggest that value-level metrics alone provide an incomplete picture of LLMs' molecular understanding, and that relative ordering offers a complementary evaluation.