Benchmarking the Protocol, Not Just the Model: Failure Modes in Zero-Shot Enhancer–Gene Linking
Abstract
Sequence-to-function models offer a promising approach for zero-shot prediction of enhancer–gene regulatory interactions. However, their reported performance is confounded by differences in how their enhancer–gene scores are extracted. We introduce E2GScore, a unified enhancer–gene scoring framework that parameterizes the scoring algorithm, and benchmark Enformer, Borzoi, and AlphaGenome under a matched protocol on 10,353 K562 CRISPRi-validated enhancer–gene pairs. We show that critical parameters can affect performance as much as the model choice itself, identifying four failure modes that can substantially distort conclusions in this benchmark and may generalize to zero-shot enhancer–gene linking more broadly, namely: benchmark-support conflation, model–protocol entanglement, cross-gene scale sensitivity, and biological context mismatch. Our conclusion is that benchmarking these models on enhancer–gene linking should carefully account for protocol and evaluation design.