Agentic Discovery in RNA Editing
Abstract
State-of-the-art RNA-editing predictors achieve strong performance, yet deciding what to investigate next remains a scientific challenge. We study an agentic search that produced ADARPro, an improved RNA-editing predictor. The agent worked independently within human-defined objectives and constraints. A fixed benchmark and common references let the agent test candidate formulations on the same development data and use measured gains to guide its next action. Building on the investigators' earlier exploration of local windows and data handling, it developed effective formulations: pair-first sampling (substrates before sites) and local windows centered on the target and its predicted pairing partner. It implemented controlled comparisons and combined supported components, while closing tested alternatives or deferring directions without suitable measurements. The resulting ADARPro ensemble was frozen before final evaluation. On the held-out high/low editing benchmark, ADARPro outperformed reproduced per-endpoint models of AdarEdit, a state-of-the-art RNA-editing predictor. Mean gains were 2.77, 2.61, and 2.53 percentage points in area under the receiver operating characteristic curve (AUROC), area under the precision-recall curve (AUPRC), and F1, respectively. All 15 conditional 95% whole-pair bootstrap intervals were positive across four tissues and an overlapping pooled endpoint. The evidence map separates useful training and representation choices from failed formulations and biological questions that remain open.