Dataset Collections: Challenges of Large-Scale Data Aggregation in 3D Medical Image Datasets
Abstract
The development of deep learning-based medical AI depends on high-quality annotated data. Yet, dataset creation remains constrained by the need for expert radiologist annotations, limiting curated datasets to relatively small cohorts. To overcome this limitation, recent efforts increasingly aggregate disparate public datasets into large-scale dataset collections. However, combining datasets at scale often obscures data provenance, introducing significant risks of reuse propagation, unintended leakage, bias amplification and distorted data distributions. In this work, we study these problems on a massive collection of 760k 3D radiological images across 1069 public datasets and provide practical resources to address it. Firstly, we formalize dataset collections as a distinct paradigm in medical image analysis and characterize their structural challenges, including reuse propagation, duplication, and provenance fragmentation. Secondly, we perform one of the first large-scale empirical studies of duplication on medical image collections and demonstrate extensive exact and near-duplicate reuse across public datasets via perceptual hashing and manual tracing on our 760k 3D radiological images across 1069 public datasets, with 118k near or exact duplicates. We make these hashes publicly available, thereby providing a standardized reference for identifying potential overlaps and shared provenance for researchers. Thirdly, owing to the insufficient robustness of hashing in isolation, we introduce MAP (Metadata for Aggregation and Provenance), a lightweight, portable metadata profile for documenting provenance, known overlaps, and deduplication decisions in dataset collections. Our contributions enable dataset collections to shift from coarse aggregations to reliable and provenance-aware data scaling mechanisms in medical image analysis. A webtool supporting our work is made available here: