Quantifying Centrality for Complex Data
Abstract
Analyzing complex data across a wide range of scientific fields often requires identifying central or typical elements, distinguishing them from atypical ones, and constructing central or "normal" regions at prespecified levels of centrality. We introduce centrality scores (C-scores) based on distance profiles for random objects taking values in general metric spaces. The proposed C-scores are defined through weighted optimal transport of distance profiles and provide a unified framework for constructing central or "normal" regions for complex data. We show that, for Euclidean data, the limiting behavior of the proposed C-scores is asymptotically equivalent to the underlying probability density function. The practical merits of the proposed method are illustrated using gene expression data, distributions of recorded temperatures, handwritten digit recognition, and taxi trip records. These examples demonstrate the effectiveness of C-scores in identifying central and peripheral elements and in yielding meaningful insights for non-Euclidean and high-dimensional data.