Asymmetric Generalization in Deep CTR Models: A Block-wise Diagnosis
Abstract
Deep click-through rate (CTR) models commonly consist of sparse embedding blocks and feature-interaction blocks, which jointly learn predictive signals from high-dimensional categorical features. This naturally raises the question of whether different blocks in an end-to-end trained CTR model play the same role in generalization. To address this question, we investigate the generalization behavior of deep CTR models from a block-wise diagnostic perspective. Specifically, we formulate deep CTR models as two-block systems comprising an embedding block and a feature-interaction block, and introduce a retuning-based diagnostic protocol to separately assess the transferability of different blocks. Experiments across multiple datasets and architectures reveal a consistent asymmetric generalization pattern: the embedding block exhibits stronger training-set specificity, whereas the feature-interaction block preserves comparatively more transferable structure. Motivated by this diagnosis, we further develop a component-selective embedding aggregation method that averages multiple independently retuned embedding tables while keeping the feature-interaction block fixed. The resulting model retains the original single-path inference architecture and incurs no additional online inference cost. Experiments on Avazu, Criteo, and Taobao show that the proposed method improves representative deep CTR models over standard training and retuning baselines.