KernelDNA: Cross-Layer Kernel Sharing via Decoupled Neural Adapters
Yadong Zhang ⋅ Haiduo Huang ⋅ Yinghui Xu ⋅ Tian Xia ⋅ Wenzhe zhao ⋅ Pengju Ren
Abstract
Dynamic convolution improves the representational flexibility of CNNs by generating input-adaptive kernels, yet existing methods either inflate parameters linearly with the number of base kernels or sacrifice inference throughput due to runtime kernel assembly. We revisit dynamic convolution from the angle of cross-layer redundancy: a CKA analysis across modern CNN families shows that within-stage convolutional kernels are highly correlated, suggesting that an entire stage can be reparameterized by a single shared kernel together with lightweight, layer-specific transformations. Building on this observation, we propose KernelDNA, a parameter-efficient framework in which multiple convolutional layers in a stage share one ''parent'' kernel, while layer-specific ''child'' kernels are produced by a \textit{decoupled} adapter that splits modulation into (i) an input-dependent dynamic channel gate and (ii) static spatial and filter modulations that can be pre-fused into the parent kernel before deployment. We provide a theoretical justification grounded in tensor approximation: when two kernels have CKA similarity $\ge 1 - \delta$, the multiplicative adapter incurs a Frobenius approximation error $O(\sqrt{\delta})$, and the implicit gradient consensus across child layers provably reduces the Rademacher complexity of the hypothesis class. Across ImageNet-1K and MS-COCO, KernelDNA achieves state-of-the-art accuracy-efficiency trade-offs over diverse backbones (ResNet18/50, MobileNetV2, ConvNeXt-Tiny), reducing parameters by $1.2$-$5\times$ versus dynamic convolution baselines while retaining $90$-$99$ % of the standard-convolution throughput---outperforming KernelWarehouse and FDConv across all backbones, and ODConv on $4$ of $5$ settings at a fraction of its parameter count.
Chat is not available.
Successful Page Load