Auditing Correlated Failures in Frozen-Feature Pretrained-Encoder Pools for Medical Segmentation
Eunseob Choi ⋅ Kyeonghun Kim ⋅ Hyuk-Jae Lee ⋅ Nam-Joon Kim
Abstract
Cross-architecture/backbone medical-segmentation pools are commonly motivated as complementary failure coverage when members make different mistakes; that assumption is rarely tested at the per-case level. We audit this assumption for a fixed pool of 11 pretrained-encoder bundles read out by a deliberately weak $1\times1$-convolution decoder (4 seeds), and for a small set of decoder-recipe perturbations on the same bundles. Across five primary task rows (Kvasir polyp, ACDC LV, BraTS foreground, RIGA Cup, RIGA Disc), this 11-bundle $1\times1$-conv readout pool yields per-case Dice correlation gaps of $\Delta=0.261\text{--}0.638$ between same-bundle and cross-bundle pairs after a Dice $\geq 0.30$ functional floor; all paired case and hierarchical bootstrap intervals exclude zero. The result is stable under floor sweeps, ICC(A,1), leave-one-out, subject aggregation, quality-controlled pair regression, and cross-fitted item-difficulty residualisation. Decoder-recipe diagnostics show that richer readouts attenuate the gap (UNet-skip; case-identical nnU-Net on RIGA Cup: $0.473\to 0.0079$, separate-split $0.034\text{--}0.047$); the reported magnitudes are recipe-specific. Additional probes (pixel error, joint-failure lift, same-architecture-family ViT, same-family cross-checkpoint, same-backbone) locate the dependence structure but do not causally separate architecture, pretraining, and recipe. The claim is scoped to 2D frozen/light-adaptation encoder pools with this readout family and to retrospective benchmark audits. We release per-case Dice traces, summary JSONs, scorer code, hashes, and metadata as the audit artifact.
Chat is not available.
Successful Page Load