Exact Per-Gene Attribution Rankings for Bacterial Phenotype Prediction
Abstract
Biological foundation models are quickly becoming useful across applications, from designing therapeutics to annotating genes. Whole-genome embeddings extracted from them offer a continuous, numerical representation of an organism from which downstream properties such as phenotype may be decoded. Ideally, the features a model uses to predict those properties align with the causal biology. We present a simple method for interpreting linear probes built on mean-pooled embeddings and evaluate it on phenotype prediction from whole-genome bacterial embeddings. We used a protein foundation model to obtain gene-level embeddings, mean-pool them into a genome embedding, and train linear probes to predict phenotype from. Exploiting the linearity of the probe and the monotonicity of the sigmoid, we obtain an exact ranking of the contribution that each gene makes to a given prediction. By annotating the ranked genes, we can then examine whether the model recapitulates the known biology of the predicted phenotypes. For aerobic metabolism and sporulation, it largely does, highlighting oxygen-handling enzymes and spore coat and germination proteins. Our method also enables the discovery of potential shortcuts that can arise from correlated factors, such as phylogeny, which we observed when predicting Gram-negative bacteria; the model used motility genes more strongly than cell envelope genes. We further find that the biological signals a model uses depend on which layer the embeddings are taken from. Finally, because the method links genes to phenotypes, it may also be used to form hypotheses about the role of genes that are associated with a phenotype but cannot be annotated by traditional bioinformatic methods. Code is available at \url{https://anonymous.4open.science/r/phenoprobeanon-E517}.