Know Thy Neighbour: Benchmarking Protein Language Model Embedding-Based Annotation Reveals Anisotropy Substantially Impacts Retrieval
Abstract
Functional annotation using protein language models (`pLMs') usually proceeds by transferring the function of a query's nearest neighbour in a reference database. A major drawback is that embedding-based annotation transfer has no null model i.e. no accepted way to set a cosine-similarity threshold to control false positives. This contrasts with traditional bioinformatics methods, including sequence- and structure-based alignment tools, which rank hits with E-values derived from explicit null models. Here, we benchmark a variety of embedding methods against MMseqs2 and Foldseek on different curated datasets based on Swiss-Prot and SCOPe. Setting per-model cosine similarity thresholds controlling for a 1\% false-positive rate, every pLM method retrieves fewer remote homologs than alignment-based tools (at most 37\% versus 85\% on Swiss-Prot). Strikingly, for the ESM-2 and ESM-C families, shuffled decoys have \emph{higher} cosine similarities to their nearest database hits than real remote homologs do. We trace this failure to extreme anisotropy, showing ESM-2 and ESM-C embeddings of SCOPe and Swiss-Prot database proteins occupy fewer than 15 effective dimensions out of 640--2560 (under 1\%), far lower than other tested pLMs. However, applying ZCA whitening fit on the reference database alone substantially increases sensitivity for every pLM and lifts threshold-free Top-1 superfamily retrieval to levels comparable to Foldseek. As ZCA whitening requires only the reference database and is compute cheap, it should be considered and tested when considering any pLM embeddings-based annotation task.