DegradeQuery: Counterfactual Tuple Pretraining for Context-Aware PROTAC Degradation Prediction
Abstract
Proteolysis-targeting chimeras (PROTACs) induce protein degradation through coordinated interactions among a degrader molecule, a target protein, and an E3 ubiquitin ligase. Yet computational prediction is often framed as if degradation were an intrinsic property of the degrader alone, overlooking the tuple-level context that determines activity. This mismatch is especially consequential in public PROTAC datasets, where degradation labels are sparse but molecule–target–E3 records are abundant. We introduce DegradeQuery, a tuple-conditioned framework that uses unlabeled records as structured relational evidence rather than incomplete labeled examples. DegradeQuery first performs counterfactual tuple pretraining, learning to distinguish observed molecule–target–E3 tuples from alternatives generated by replacing the target, the E3 ligase, or both. This objective induces a conditional compatibility representation before degradation labels are used, enabling the model to capture how degrader activity depends on biological context. The pretrained representation is then fine-tuned for high/low degradation prediction. On the official PROTAC-8K benchmark, DegradeQuery achieves 0.9065 AUROC and 0.8500 accuracy, outperforming reported semi-supervised results without pseudo-labeling, teacher models, distillation, or ensembles. Ablations, scaffold holdout, target–E3 holdout, and target-wise few-shot adaptation further support PROTAC degradation prediction as a tuple-conditioned rather than molecule-only problem.