AgriManager: A Framework and Benchmark for LLM-RL Generalization in Agricultural Management
Chi Gui ⋅ Junwen Zheng ⋅ Ritwik Nigam ⋅ Qianlan Yang ⋅ Meagan Lang ⋅ Vardhan Dongre ⋅ Jeremiah Barr ⋅ Yu-Xiong Wang ⋅ Vikram Adve
Abstract
Agricultural management is heterogeneous along two axes. The first is distributional shift within a fixed task interface: the same sensors, action menu, and objective, but with weather, planted crops, and prices varying across years and locations. The second is structural shift across interfaces: different farms expose different observed variables, available actions, and management objectives. We formalize the second axis in MDP terms as \emph{schema shift}: variation in the policy-facing schema $\Sigma = (\Sigma_O, \Sigma_A, \Sigma_R)$ between training and deployment. Fixed-interface neural policies are structurally inapplicable when $\Sigma_O$ or $\Sigma_A$ shifts and require retraining when $\Sigma_R$ shifts, while prompt-conditioned LLM policies re-read $\Sigma$ from the prompt of each new episode. Whether LLM-RL actually generalizes across both axes has not been systematically tested. We address this question with \emph{AgriManager}, the first multi-simulator LLM-RL framework for agricultural management, unifying the DSSAT, WOFOST, and Cycles crop simulators under prompt-based schema specification with joint multi-objective, multi-environment training. On this platform we design a three-tier generalization ladder, each tier grounded in a real-world deployment scenario: Tier 1 within-schema distributional shift (weather, crop, price), Tier 2 single-axis schema shift along $\Sigma_O / \Sigma_A / \Sigma_R$, Tier 3 compositional schema shift across simulators. Within fixed schemas, prompt-conditioned LLM-RL with explicit reasoning matches or exceeds fixed-interface NN-RL on all three settings, with reasoning becoming necessary when the reward shape penalizes input use. Single-axis schema shift imposes a different burden along each axis: information sufficiency for $\Sigma_O$, joint action-menu coverage for $\Sigma_A$, and objective coverage for $\Sigma_R$. Under compositional schema shift, a unified policy with marginal per-platform coverage and explicit reasoning matches or exceeds per-source specialists on all three callback targets. Without that coverage, pairwise zero-shot transfer is bounded by transition dynamics, which the prompt cannot describe and which we leave as an open task. Together, these results provide an initial empirical baseline for unified prompt-conditioned control across heterogeneous agricultural deployments.
Chat is not available.
Successful Page Load