Sparse crosscoders identify antibody features gained during fine-tuning of protein language models
Abstract
Fine-tuning adapts a machine learning model pre-trained on a large corpus to a narrower context. In biological sequence modeling, fine-tuning has been applied to specialize general protein language models to the antibody domain, and further to paired-chain antibodies, for which little data is available. While fine-tuning improves masked language modeling performance on antibody sequences, it is unclear what antibody-specific features are learned during fine-tuning and which protein features may be lost. Better understanding this process would help us develop fine-tuning techniques to learn more relevant antibody features on smaller datasets. Here, we apply sparse crosscoders, a model diffing technique from mechanistic interpretability, to the protein language model ProtBert and its antibody fine-tuned variants IgBert-unpaired and IgBert. The crosscoders learn shared latents whose contribution to each model we quantify. We find that most latents are shared across all three models and latents that are specific to the fine-tuned models fire predominantly on antibody sequences. These include interpretable detectors of antibody regions and V genes, which may explain the observed germline bias of antibody language models. In contrast, we find no latents that distinguish IgBert from IgBert-unpaired. This could mean that fine-tuning on paired antibody sequences alters the model's representations less than expected, or that more sensitive interpretability techniques are needed to detect the difference.