Region-Normalized DPO for Medical Image Segmentation
Abstract
Direct Preference Optimization (DPO) has recently been applied to medical image segmentation, since comparing candidate masks is substantially cheaper than producing new dense pixel-level annotations, making pairwise preference signals an appealing source of supervision. However, existing work derives preferences from ground-truth masks, leaving open whether the approach is sound under the imperfect feedback available in practice. In our analysis of this reformulation, we identify that candidate masks tend to agree over most of the image and differ only in localized regions, yet standard DPO averages the preference signal over the full spatial domain. This couples update strength to disagreement area, diluting small correct refinements and amplifying large misranked differences. Crucially, this bias even degrades performance even under oracle preferences derived from ground truth. We propose Region-Normalized DPO (RN-DPO), which normalizes the likelihood ratio over the disagreement region between candidate masks, removing this coupling. We further provide a systematic empirical study of preference-based segmentation fine-tuning under controlled noisy judges, analyzing the effects of mining strategies, judge types, and reliability regimes. Across two medical segmentation benchmarks, multiple judge configurations, and two backbone architectures, RN-DPO consistently improves over vanilla DPO and robust DPO variants, with gains persisting under noiseless oracle preferences.