Using Interface Similarity to Better Understand and Improve Antibody-Antigen Complex Prediction
Abstract
Computational prediction of antibody-antigen complexes is an important problem in therapeutic design. Co-folding methods have shown great promise at tackling it, with strong performance across multiple benchmarks. However, the success and failure modes of these models are not well understood. To understand their performance, we aim to quantify the inherent difficulty of benchmarks based on the interface similarity (IS score) between their test and train data. To further verify that this information overlap can be useful, we use it alongside a simple alignment docking method (IS-Adock). We find that for the high homology regime, this results in strong dockQ performance, and that therefore this information can and should be used by co-folding models. We next assess Boltz-2 and ESMFold2 on the same benchmarks, and we find that their performance has a relatively weak correlation to IS-score. Specifically, they perform poorly in multiple cases where a similar interface exists, and strongly when one does not. This is in contrast to protein-ligand co-folding, where a predictable performance drop-off is observed as similarity decreases. While a considerable effort has been made to evaluate the generalisability of co-folding methods in low-homology cases, less attention has been paid to the high-homology regime, where it is still important that the models work well. In order to make better use of this information, we train an SE(3)-equivariant encoder to pick out similar interfaces given just an epitope (EquiTope). We show that EquiTope is able to rapidly pick out highly similar pairs, which, combined with align docking, can be used to reliably create successful docks.