Agentic Search for Deployment-Specific Configuration of Promptable Detectors
Abstract
Promptable detectors can accommodate new target classes without retraining, but they still require deployment-specific configuration. Deployment-specific configuration can span heterogeneous choices: prompt phrasing, visual exemplars, text and visual negatives, score calibration, box re-classification, low-rank detector updates, and tiled inference. Yet deployment guidance typically prescribes a fixed recipe, implicitly assuming that deployments fail alike. Our results indicate that they do not. We formulate configuration as a search problem and give it to an agent: from a labeled deployment sample and a target metric, it diagnoses the current configuration's errors, proposes targeted changes, and adopts one only when a paired bootstrap on held-out data supports it. Across three promptable detector backends the search consistently improves sealed-test performance, by +0.05 to +0.41 micro F_beta on the primary track, while selecting different recipes for the same data, so the appropriate configuration depends jointly on the dataset and the detector. On a twenty-dataset few-shot benchmark, substantial gains persist even under sparse and incompletely annotated supervision.