One Step is Enough: Multi-Agent Reinforcement Learning Based on One-Step Policy Optimization for Order Dispatch on Ride-Sharing Platforms
Abstract
Order dispatch is a critical task in ride-sharing systems with Autonomous Vehicles (AVs), directly influencing operational efficiency and profitability. While Multi-Agent Reinforcement Learning (MARL) offers a scalable paradigm by decomposing the massive state-action space, existing methods are heavily reliant on accurate value function estimation. In large-scale, highly stochastic urban environments, such estimation is notoriously prone to bias and instability, severely limiting training efficiency and policy quality. To overcome this limitation, we propose two novel policy optimization methods that completely bypass the need for critic networks. First, we establish a formal connection between the homogeneity of AV fleets and the short-memory mixing property of the transportation network, proving that agent value functions converge to a common time-varying baseline up to a negligible residual. Leveraging this insight, we introduce Single-Trajectory Group Relative Policy Optimization (ST-GRPO), an adaptation of Large Language Models (LLMs) post-training techniques to multi-agent trajectories, which replaces the traditional value baseline with the fleet-wide average reward-to-go. Inspired by this reduction, we further derive One-Step Policy Optimization (OSPO), demonstrating that under the established structural priors, an optimal policy can be learned using only immediate, group-normalized rewards—rendering long-horizon bootstrapping unnecessary. Experiments on real-world ride-hailing datasets from Manhattan and Queens demonstrate that both ST-GRPO and OSPO achieve promising performance, particularly in reducing pickup times and increasing order service rates. Remarkably, both methods operate efficiently using simple Multilayer Perceptron (MLP) networks with minimal GPU utilization, underscoring the practical impact of exploiting domain structure for scalable MARL. Our code, trained models, and processed data are provided at the anonymous repository: https://anonymous.4open.science/r/OSPO-2105 .