Query-Composition Sensitivity of Open-Vocabulary Detectors in Under-Canopy Search and Rescue
Abstract
In under-canopy search and rescue, visual evidence includes the missing person and personal belongings likely lost or left behind as critical search clues. Such belongings span categories too diverse to define exhaustively in advance and are highly variable in appearance, motivating open-vocabulary detectors (OVDs) rather than fixed closed-set detectors. However, in operational search settings, operators often construct text queries that name different combinations of search targets depending on their search intent. For example, operators often query a person alone or together with clothing or other belongings, and naming additional targets is likely to affect the predictions even for a search target named in every query. In this work, we investigate how text-query composition affects OVD predictions and detection performance in search and rescue scenarios. We adapt OVDs to under-canopy search and evaluate the adapted OVDs using a controlled text-query protocol that varies which person and belongings targets are named together and how each target is named. Experimental results show that predictions and detection performance for the same search target can differ depending on the other targets named in the query for fusion-based OVDs, while the alignment-based detector remains invariant. Adaptation also transfers unevenly across operationally equivalent target names. These findings demonstrate that query construction should be explicitly considered when evaluating and deploying OVDs. To support this study, we introduce \dataset, a benchmark for detecting personal belongings in forest environments, containing 31,897 under-canopy images with 39,831 annotations, publicly available at https://huggingface.co/datasets/anonreviewer2026/ForestBelongings.