Auditing Capsule Vision 2024: Within-Split Train-to-Validation Re-Exposure and a Kvasir-Channel Sensitivity Diagnostic
Abstract
Public validation sets in machine-learning challenges frequently become de facto test sets in downstream work, making train–validation source overlap a first-order concern for claim validity. We audit the released public split of the Capsule Vision 2024 Challenge (CV2024) using perceptual hashing (pHash+dHash, PDQ at 99.4%), pixel NCC, and DINOv2/ResNet feature checks. The released CV2024 public split contains 1,381 / 11,581 KVASIR validation rows (11.9%; 540 pixel-confirmed at NCC ≥ 0.99, i.e. 4.66%) with pHash-exact training matches — directly violating the organizers' "NOT have duplicates" preprocessing description. All 1,381 trace to shared Kvasir-Capsule video prefixes, consistent with frame-level (not video-level) random splitting. The full CV2024-KVASIR slice falls in the disclosed Kvasir-Capsule near-duplicate set (60.7% bit-identical; non-KVASIR controls ≤ 0.31%). Training on the Kvasir-origin-removed le6 split drops public-validation balanced accuracy by Δ = −0.213 (95% CI [−0.220, −0.206], n=10 paired seeds), quantifying the public score's dependence on the KVASIR channel; the magnitude is mechanically expected given the 71.8% KVASIR validation composition rather than a leakage-specific causal effect. le6 is not a hidden-test proxy (team-level Spearman against AIIMS: ρ=0.355 vs. 0.566 for the original score). We release le6, le6plusinternal, hash/NCC annotations, audit code, and a Croissant 1.1 descriptor as a within-public-pool sensitivity probe; no image bytes are redistributed. Detector portability: ISIC 2019 cross-source rate is 0.008% (2/25,331, both re-attributed as within-source metadata-gap pairs ⇒ genuine cross-source 0/25,331; Cassidy et al. 2022's recall-tuned multi-hash pipeline flags 14,310); HyperKvasir and Kvasir-SEG cross-benches against CV2024 return 0 pairs (procedure-disjoint negative controls). At the matched NCC ≥ 0.99 endpoint, the CV2024 rate of 4.66% (540/11,581) is ≈ 148–583× above the ISIC cross-source rate (Wilson-lower / Wilson-upper to point/point).