Separating Architecture from Initialization: What a Controlled CNN-Transformer Benchmark Can and Cannot Attribute.
Abstract
Evaluation comparisons between Convolutional Models and Vision Transformer models often conflate architecture with initialization, making the reported score differences hard to interpret. We test whether the apparent Transformer advantage persists when both the CNN and ViT families are trained under identical conditions. We benchmark ten architectures across nine datasets, three data fractions, and two seeds, comparing scratch-trained CNNs, scratch-trained attention models, and ImageNet-1k fine-tuned attention models. When trained from scratch, every CNN outperforms every attention model in mean accuracy, and the CNN group leads on all nine datasets (paired Wilcoxon, p=0.004). The size of this gap varied substantially, from 8.4 to 18.3 percentage points depending on treatment of a non-converging model. Four benchmark conclusions, including the only significant corruption-robustness difference, were driven by a single non-converging architecture. Pretraining recovered 52-65% of available headroom on natural, fine-grained, and satellite datasets, but only 20% on medical data, while its apparent value depends on the baseline used for comparison. These results show that CNN-Transformer conclusions strongly depend on initialization and evaluation choices, motivating benchmarks that vary architecture and initialization independently.