Hyaline: Kinase-State Prediction Can Learn Taxonomy Instead of Conformation
Abstract
Random structure splits can overstate progress in kinase conformational-state pre- diction because repeated structures of the same kinase occur in both training and test data. We audit this failure mode using public KLIFS-derived DFG labels and find that 77.8% of kinase identities occur in only one DFG class. A negative- control classifier receiving no sequence and no coordinates—only a one-hot ki- nase identifier—reaches mean AUROC 0.907 across 100 class-balanced random five-fold repetitions (SD 0.009; empirical stability range 0.890–0.923). Kinase family alone reaches 0.888 AUROC (SD 0.009; range 0.870–0.901), showing that the shortcut is not confined to exact identifiers. A fixed DFG-to-αC-helix distance provides a pooled structural reference at 0.844 AUROC, while a learned distance-plus-angle classifier reaches 0.834 under grouped leave-one-kinase-out (LOKO) evaluation. Hyaline therefore argues that categorical negative controls and kinase-grouped evaluation are necessary before claiming transferable conformation recognition.