QUEST: Q-Learning for Uncertainty-Guided Efficient Search Teams
Arsh Verma ⋅ Tejus Gupta ⋅ Jeff Schneider
Abstract
Robots deployed for active search must decide where to sense while uncertainty, teammates, and remaining mobility change online. In active search, each waypoint decision also commits the robot to a route through the map, making efficient localization a long-horizon problem over belief, topology, robot state, and traversal budget rather than a sequence of independent next-best sensing decisions. We present QUEST, a shared-belief graph neural framework for multi-agent active search with noisy sensors, a shared posterior, and finite decision and path budgets. Rather than learning a myopic policy that values only the next route, QUEST trains Q-functions whose targets include both reward collected along the executed path and the downstream value of the resulting belief, robot positions, team coverage, and remaining budget. The same formulation supports single robots, homogeneous robot teams, and heterogeneous UAV--UGV teams with different sensors and motion budgets. During execution, each robot acts independently using the shared posterior and knowledge of teammates' positions and search plans. Across simulated search environments, QUEST reaches $0.913$ F-score with four UGVs, versus $0.853$ for the strongest learned baseline, and keeps above $0.90$ F-score under communication outage and two mid-mission robot failures without retraining. On out-of-distribution UAV--UGV teams, it improves early F-score ($0.759$ vs.\ $0.728$) while using $10\%$ less path length and $12\%$ less duplicate coverage than the same baseline. The same policies transfer across map shape, variable team size, communication outage, and mid-mission robot failure, and show sensing-mode adaptation after UGV failure.
Chat is not available.
Successful Page Load