Do Molecular Probes Encode Ligand Preferences?
Abstract
AlphaFold3-like models produce rich internal representations of protein--ligand interactions, yet the structure and information encoded in these representations remain poorly understood. As these representations are increasingly used in downstream drug discovery, understanding what they encode is important for their reliable interpretation and reuse. We use Protenix and a fixed panel of 128 molecular probes to analyze ligand-conditioned PairFormer states across kinase targets. After removing a shared response manifold, same-target probe subspaces capture a larger fraction of direct-ligand residual-state energy than probe banks from other targets (0.529 vs.\ 0.422). More strikingly, ligand-specific coordinates over probe identity learned on eight calibration targets transfer to 24 unseen targets, recovering target--ligand interaction variation with cosine 0.537, while simple chemical-similarity routing fails. These results suggest that molecular probes expose a structured and partially transferable organization of ligand preferences in cofolding representations.