When More Cores Are Not Faster: Towards NUMA-Aware Extreme Multi-Label Classification
Kopal Rastogi
Abstract
Extreme multi-label classification (XMLC) is an important task in machine learning that addresses the task of predicting multiple relevant labels for each input instance, where each instance has multiple associated labels from a very large set (in the order of millions). While XML builds upon and generalizes multi-label learning principles, computational scalability is a major challenge. Methods such as DiSMEC and ProXML attempt parallel computation to train one-vs-all sparse models efficiently, making their scalability with increasing computational resources an important practical consideration. Such scalability is typically characterized in terms of the number of CPU cores available for parallel execution. However, core count alone does not capture the underlying hardware topology. While NUMA-aware optimization has been studied in other machine learning and high-performance computing workloads, its implications for the multicore scaling of sparse extreme multi-label classification methods remain underexplored. We conducted a preliminary empirical study of the multicore scaling behaviour of DiSMEC and ProXML on three benchmark XMLC datasets -- Eurlex-4k, Eurlex-4.3k, and Wiki10-31k using the data representations and benchmark splits from the Extreme Classification Repository. The experiments were performed on a 64-bit Linux system with a 40-core Intel Xeon Silver 4210R processor at 2.40\,GHz and 64\,GB RAM, using OpenMP for parallelization. Each method was executed using 1, 10, 20, 30, and 40 CPU cores. We measured training time, prediction time, and predictive performance using precision@k, nDCG@k, propensity-scored precision@k, and propensity-scored nDCG@k for $k \in \{1,3,5\}$ across configurations. The predictive performance remained unchanged, while test time showed only negligible variation. Training time generally decreased substantially as parallelism increased; however, the improvement was not consistently monotonic. The results clearly exhibited non-monotonic scaling between 20 and 30 cores; that is, training time decreased substantially up to 20 cores but increased when the configuration was expanded to 30 cores. The experiments were conducted on a 40-core system having two NUMA nodes, making this transition particularly interesting. Although these results do not establish this computer memory design as the definitive cause of the slowdown, the observed pattern motivates a systematic investigation of memory locality, thread placement, and cross-node latency in scalable XMLC. This work identifies hardware topology as a potentially overlooked factor in the scalability of extreme multi-label learning and motivates future NUMA-aware execution and scheduling strategies for one -vs-all sparse models such as DiSMEC, ProXML, and related XMLC methods.
Chat is not available.
Successful Page Load