Cardinality-Decomposed Loss for Heterogeneous GNNs
Abstract
Graph Neural Networks trained on heterogenous bipartite graphs form a common basis in recommendation systems. These graphs often express relations that vary in cardinality, for example, user-item preferences are one-to-many and user-attribute features are one-to-one. Traditionally, a unique loss function is applied for all of the network components which is often Bayesian Personalized Ranking (BPR). While BPR works well for the recommendation task, we find that it causes attribute embeddings to collapse to near-random geometry — a silent failure that leaves standard ranking metrics largely unaffected and therefore invisible to conventional evaluation. This in turn pollutes user node embeddings, which are shaped by both edge types simultaneously, hurting downstream tasks like personalization, segmentation, etc. Here we propose a Cardinality-Decomposed Loss (CDL) that combines both Cross Entropy (CE) and BPR to enable the model to collectively optimize for relations across cardinalities. As we implement this loss, we also confirm the conflict between CE and BPR by showing that the two losses compete against each other in the shared encoder's parameter space. We evaluate CDL on five datasets spanning two structural configurations — one-to-one attributes on user nodes (MovieLens-1M, Last.fm-360K, PayPal Audience Factory, BookCrossing) and on item nodes (Yelp) — and find that CDL consistently improves discriminability in attribute embeddings. We also show that whenever these attributes contain meaningful preference signal, we also see improvement in the ranking task (measured by NDCG). On the other hand, when attributes are weakly correlated with preferences, there is an inherent tension between the two objectives. We use a lambda parameter to navigate this trade-off, and a lambda-sweep reveals that dataset behavior is governed by two graph properties — semantic alignment and topology leakage. Semantic alignment captures whether the one-to-one attribute is predictive of user preferences, while topology leakage captures whether message passing already encodes attribute structure implicitly through the graph's connectivity.