Learning Collision‑Free Dispatch Policies for Route‑Wise Decision‑Dependent Anomaly Detection
Abstract
Mobile sensing for real-time anomaly detection couples routing, sampling, and stopping: dispatch determines which locations are observed, while observations update the evidence used for future dispatch. We study this problem for a fleet of unmanned aerial vehicles (UAVs), where each UAV observes only its current location and anomalies may occur at unknown locations and times. The goal is to detect anomalies quickly while controlling false alarms under decision-dependent partial observations, local mobility, and collision-avoidance constraints. Existing quickest-detection methods typically ignore mobility and collision constraints, whereas learning-based routing methods scale to large fleets but are not designed for statistically calibrated detection. We formulate route-wise monitoring as a partially observable Markov decision process (POMDP) and propose DispatchPPO, a deep reinforcement learning method for long-horizon detection-aware dispatch. Its policy architecture, DispatchNet, uses autoregressive decoding with dynamic feasibility masking to generate collision-free joint UAV moves without enumerating the exponential joint action space. We establish theoretical properties for persistent coverage, evidence-based focusing, and detection delay through a resource-allocation information-rate bound. Experiments show that DispatchPPO reduces detection delay relative to heuristic and optimization-based baselines, while transferring effectively to real-robot and wildfire monitoring case studies.