PIGRAM: An Interpretable Patch--Motif Interaction Grammar for Protein--Nucleic-Acid Recognition
Abstract
Protein--nucleic-acid recognition underlies aptamer discovery and nucleic-acid therapeutics, yet computational models often trade mechanistic interpretability for predictive flexibility. Hybrid interpretable--latent architectures risk fallback takeover, where the latent branch silently becomes the true predictor while the explicit pathway remains decorative. We introduce PIGRAM, a patch--motif grammar model that decomposes proteins into local residue patches and nucleic-acid partners into secondary-structure motifs, learning a sign-separated grammar over explicit physicochemical and geometric attributes to preserve rare favorable interactions. Heterogeneous rule families are selectively refined into a coarse-to-fine atlas, and a frozen-grammar residual Transformer provides sparse, bounded corrections without displacing the grammar as the primary predictor. On a strict sequence-pair-disjoint benchmark, PIGRAM achieves a Pearson correlation of 0.494, approaching the unconstrained Transformer (0.511) while exposing signed rules, fine subtypes, and route-gated corrections for each prediction. The learned rules recover bidirectional patch--motif mechanisms, and residual corrections concentrate selectively on grammar-hard regions. In fluorescence ELISA experiments, suppressive-rule relief improved GFP binding, whereas favorable-rule disruption weakened NELF binding, providing direct experimental support for grammar-guided design. Together, these results show that interpretable grammar can serve as both a predictive scaffold and an experimentally actionable design principle for aptamer optimization.