ConfDet: Learning Reliable Confidence for MLLM-based Detection
Abstract
MLLM-based detection has shown promising end-to-end detection ability in recent years, but its generated detection results still lack reliable confidence for evaluation and interpretation. Since detection data itself do not provide human-annotated confidence labels, existing MLLM-based detection methods mostly obtain confidence implicitly from prompted self-assessment, generation probabilities, or external proxy scores. While these scores are available for evaluation, they are not explicitly connected with the actual detection quality. Inspired by conventional detectors that optimize confidence through classification, objectness, or quality estimation branches, we formulate confidence in MLLM-based detection as an explicit modeling target to be learned, optimized and calibrated. Specifically, we present ConfDet, a practical confidence framework that first uses discrete confidence tokens based supervised fine-tuning to enable stable and parsable confidence generation, then applies GRPO with detection and separation joint rewards to optimize confidence score separation and preserve detection quality, and finally performs condition-aware post-hoc calibration to align confidence with empirical accuracy. Experiments across different in-distribution and out-of-distribution benchmarks show that, ConfDet greatly improves AP through better confidence ranking and reduces confidence miscalibration across benchmarks, providing a reliable confidence for MLLM-based detection.