Inference-Time Antibody Refinement with Agentic Strategy Search
Abstract
Antibody humanization replaces non-human framework residues with human-like ones while preserving the murine CDR loops that determine binding. Discrete-diffusion models such as HuDiff cast this as conditional sequence generation, but their one-pass sampling commits residues before the full sequence context is available. We propose AIR, an inference-time refinement framework that lets a pre-trained discrete-diffusion model audit and revise its own predictions. Each cycle remasks framework residues flagged by the model's confidence, an external biological scorer, or both, and resamples them within a more complete sequence context. On the Humab25 benchmark, AIR traces a controllable trade-off between humanness and fidelity. Pairing external nativeness scoring with a low-rate self-consistency stage approaches the humanness and total preservation of experimentally validated humanizations. As an alternative to grid search for choosing an AIR variant, we also explore agentic strategy search: an LLM proposes programs composed from AIR's refinement primitives and uses evaluation feedback to guide further proposals. This search discovers AIR-Cascade, which combines changing criteria, proximity protection, and edit counts to reach laboratory-level humanness at comparable total preservation on the benchmark. AIR requires no retraining and operates as a drop-in wrapper for existing discrete-diffusion humanization models, with either hand-designed or search-discovered refinement strategies.