I Can’t Believe It’s the Encoding: Better than Deep Learning via an Interpretable Linear Model in Proteomic Prediction
Abstract
Interpreting the output of deep learning models remains a major challenge, particularly in biology, where understanding the underlying mechanisms is essential for clinical applications. Here, we focus on predicting plasma protein abundance from genetic data, which plays a crucial role in biomarker discovery and is an important first step toward downstream disease-risk prediction. A recent paper has proposed a deep learning approach that significantly outperforms penalized linear regression baselines by modeling non-linear interactions. While this seems—at first glance—the usual success story of deep learning, our work paints a different, more nuanced picture. In fact, we show that the performance gap between deep learning and a linear model baseline is completely closed by (i) using an interpretable Bayesian method, (ii) performing a two-stage data selection and, most importantly, (iii) one-hot encoding the selected features. More precisely, as for (i), we build on a genomic Vector Approximate Message Passing (gVAMP) framework; as for (ii), we first select protein-associated genomic regions in a smaller genetic dataset, and then re-map them onto a denser feature set; and the one-hot encoding in (iii) closes alone 70% of the performance gap with deep learning. Our approach comes with optimality guarantees typical of approximate message passing, it scales to a dense set with over 8 million features, and it produces interpretable genomic regions which strongly replicate in independent Finnish and Icelandic cohorts. Finally, having access to a refined set of associated genomic regions enables the use of off-the-shelf non-linear methods, such as XGBoost, which then improves prediction performance by 13.6% upon the deep learning baseline across the full set of 2,923 proteins in the UK Biobank Pharma Proteomics Project.