Validation in the Wild: Evaluating Agentic Strategy Selection under Biological Distribution Shift
Abstract
Reliable model validation becomes difficult when biological deployment data differ systematically from the data available during model development. We investigate whether a scientific agent with access to analytical tools can autonomously select a validation strategy that provides a reliable estimate of deployment performance under structured distribution shift. Using pasture biomass estimation as a controlled biological case study, we construct 12 pseudo deployment episodes covering temporal, spatial, spatiotemporal, and in distribution control conditions. For each episode, five candidate validation protocols are evaluated using frozen DINOv2 image representations and a Ridge prediction model with leakage safe preprocessing. We compare the same local Qwen3-4B model in two settings: direct validation strategy selection and selection with access to metadata inspection, representation shift diagnostics, and validation evidence. All decisions are fixed before deployment labels are revealed. Tool access reduces the mean absolute validation to deployment R2 gap from 0.161 to 0.155 and breaks the protocol selection collapse observed in the LLM only baseline. However, the agent does not outperform strong fixed baselines overall, with date grouped cross validation achieving a lower gap of 0.110. Performance also depends strongly on the type of shift: the agent approaches oracle level validation reliability under combined spatiotemporal shift, yet fails to use relevant spatial evidence effectively under pure spatial shift. Both LLM settings also remain substantially overconfident. These results suggest that tool access can improve scientific agent adaptivity, but reliable methodological reasoning requires more than simply making relevant analytical tools available.