An Auditable Agent for Adaptive Tissue-Expert Selection in Spatial Gene Expression Prediction
Abstract
An agent that selects among specialist models is only as trustworthy as the criterion it uses to select them. That criterion can reward apparent structure on validation data even when the selected specialist fails on a new patient. We study this problem in spatial gene-expression prediction from routine hematoxylin-and-eosin (H&E) histology, where spatial transcriptomics is scarce and expensive and a single global predictor must represent visually and molecularly diverse tissue regions. We introduce an adaptive tissue-expert agent that considers 12 residual expert candidates. The candidates include two biologically motivated families, organized by visual appearance or image-predicted expression, together with alternatives that vary niche count, regularization, routing, and sample support. For each held-out patient, the agent measures candidate gains on separate selection patients and challenges them with 19 within-section spatial-block randomizations. Every label-dependent stage is refitted, and a maximum-statistic test accounts for the search across candidates. The agent selects an eligible specialist or returns to the global predictor. We evaluate this decision policy on 119 tissue sections, 62 patients, and 208,122 spatial spots from brain, breast, heart, and kidney. Across five frozen image encoders and three seeds, screening changes 592 of 930 validation-only choices. Validation-only selection performs worse than the matched global predictor in 137 of 310 patient–encoder evaluations; screened selection does so in eight. Mean held-out-patient Pearson correlation is 0.17021 after screening, compared with 0.17037 for validation-only selection and 0.16979 for the global predictor. Thus, the screen removes most unreliable choices while preserving similar aggregate performance. The resulting contribution is an auditable agentic pipeline and a patient-level evaluation protocol that separates predictor performance from the quality of the policy used to select it.