View Count Is Not Enough: Composition and Selection in Reduced Multi-View Protocols
Abstract
Reducing a multi-view imaging protocol is often described by the number of images retained. However, two protocols with the same number of images may preserve different anatomical views, magnifications, or acquisition states and therefore retain different predictive information. Moreover, when many candidate protocols are searched, using the same data both to select a composition and to estimate its performance can yield an optimistic assessment of the protocol-selection procedure. We study both problems in a public Nigerian digital colposcopy cohort containing four pre-acetic and four post-acetic views per examination [1]. Our analysis distinguishes the performance of a specified protocol composition from the out-of-sample performance of a procedure that selects a composition. We use the 235-patient complete-view cohort for controlled protocol reduction. Each image is represented using frozen DINOv2 patch-mean features [2], pooled separately across retained pre- and post-acetic views, and classified at the patient level using logistic regression. Holding the representation, classifier family, patient splits, and candidate view positions fixed, we enumerate all 254 non-full subsets of the eight-view examination using patient-level stratified five-fold cross-validation repeated three times. We then evaluate protocol selection using nested cross-validation: candidate subsets are ranked only within the outer-training patients, the selected subset is refitted on those patients, and its performance is measured once on untouched outer-test patients. Protocol composition substantially affects performance even when view count is fixed. Among the 28 possible two-view protocols, AUROC ranges from 0.506 to 0.664 despite identical view count. Across protocols retaining one to seven views, selecting a candidate and estimating its performance from the same evaluation predictions produces AUROCs 0.023–0.071 higher than the corresponding nested-selection estimates. Protocol selection is also unstable: the same one-view protocol is selected in 11 of 15 outer folds, whereas 11–13 distinct subsets are selected across the 15 folds when three to five views are retained. These results identify two distinct issues in reduced multi-view learning. First, view count alone does not define a reduced protocol: which views are retained materially affects predictive performance. Second, a data-driven procedure for choosing a protocol should itself be evaluated out of sample rather than reporting the best candidate found on the evaluation data. The instability of selected subsets further suggests that protocols identified from a single finite dataset should be treated as candidates for external validation rather than definitive reduced protocols. Because this study uses one Nigerian cohort and predicts recorded expert colposcopic impression rather than histopathology, it does not establish clinical equivalence or deployment readiness. References [1] A. O. Ogunlaja et al., “A multi-magnification digital colposcopy image datasets for cervical cancer screening in Nigeria,” Data Brief, vol. 65, Art. no. 112633, 2026. [2] M. Oquab et al., “DINOv2: Learning robust visual features without supervision,” Transactions on Machine Learning Research, 2024.