Decision-Aligned Evaluation Informs Evidence Integration for GRN Inference
Abstract
A gene regulatory network (GRN) represents regulatory relationships between transcription factors (TFs) and their target genes (targets). GRN inference methods score each candidate TF–target pair, so for a TF of experimental interest a researcher can shortlist its top-ranked candidate targets for perturbation experiments. Prevailing GRN benchmarks instead pool prediction scores across TFs into a single ranking. They also report performance without distinguishing whether a target appears in other labeled regulatory relationships during training. Two controls expose the resulting mismatches: TF frequency cannot rank targets for the same TF, yet scores highly under pooled evaluation; target in-degree ranks targets well when the training graph contains other regulatory relationships for them, but falls exactly to the no-skill baseline when facing targets unseen during training. We therefore compute ranking metrics within each TF and report results separately according to target history. Under the revised evaluation, the published methods we test show limited performance. Nevertheless, structural, expression, text, and target in-degree predictors exhibit complementary strengths. We therefore introduce SETD, an ensemble that learns evidence weights according to target history and adjusts them for each TF. At the primary TFs+500 configuration across seven cell-specific BEELINE datasets, SETD outperforms the strongest published method, improving macro per-TF AP by 8.6%–16.5% and EPR@10 by 31.3%–40.2%. These results show that decision-aligned evaluation can reveal evidence differences hidden by aggregate benchmark scores and turn them into a better experimental shortlist.