How Complete Should a Reference Be? A Benchmark Audit for Fluorescence Spot Detection
Abstract
Fluorescence spot-detection benchmarks in biology and medicine often rely on reference annotations that are incomplete, ambiguous, or inconsistently defined. In such settings, detector performance is not determined solely by model predictions; it can also depend on which faint or ambiguous signals are counted as reference positives. We formulate this issue as a reference-audit problem: instead of treating the evaluation reference as fixed ground truth, we test whether benchmark outcomes remain stable under plausible changes in reference definition and annotation completeness. Using the public FISH_spots dataset, a benchmark for fluorescence in situ hybridization (FISH) spot detection, we construct Raw, Lenient, and Strict audit references and evaluate diverse classical and learning-based detectors under a unified point-matching protocol. We further apply annotation-retention stress tests that simulate random and structured missing-label mechanisms. Across these audits, changing only the evaluation reference can alter measured F1 scores and model rankings, including the selected top-1 detector, even when aggregate rank correlation remains high. These findings suggest that model selection in fluorescence spot-detection benchmarks should be tested across plausible reference definitions, rather than based on a single reference annotation set.