The Quagmire of Immunogenicity-Guided Antibody Design: Labels, Not Model Capacity, Are the Limit
Amelia Villegas-Morcillo ⋅ Chirag Raman ⋅ Jana M. Weber ⋅ Marcel Reinders
Abstract
Generative sequence-based antibody design is increasingly steered by learned property predictors, so a design objective is only as well specified as the labels behind its oracle. Immunogenicity is among the most clinically consequential properties to condition on, and among the least well specified: the supervisory signal is a single anti-drug-antibody percentage (ADA%) per molecule, aggregated from published trials. To analyze that signal, we first extended and corrected the widely used 217-entry dataset to 231 therapeutics with paired sequences, spanning formats from mouse to fully human. Then, we manually reconstructed the clinical evidence for 110 of them from primary sources. We find that an antibody format lookup outperforms every sequence-based learned predictor on the 231-antibody dataset. Apparent accuracy tracks the format composition of the evaluation split rather than sequence signal. Furthermore, the labels understate their own evidence: molecules labeled at $\leq$6 ADA% span up to 86% across trials. We recomputed labels from the trial-level measurements and found that the advantage of the format-based approach disappears, while sequence-based predictors remain largely unaffected. We conclude that progress is limited by the label itself, not by model capacity, and recommend reporting ADA% at the trial level, treating the target as a distribution rather than a point value, and evaluating against a format-based baseline.
Chat is not available.
Successful Page Load