Busemannformer: Horospherical Self-Attention for Hyperbolic Graph Transformers
Youheng Yao ⋅ Ziyao Zeng ⋅ Wenbo Liao ⋅ Tianqi Wang
Abstract
Hyperbolic neural networks excel on hierarchical graph data, yet existing hyperbolic Transformers lift standard attention into curved space without using the ideal boundary, the structure most directly tied to tree hierarchy. We propose \textbf{Horospherical Self-Attention} (HSA): each query selects an ideal-boundary direction and scores keys by their Busemann depth along that direction, yielding a token-to-token hyperbolic attention kernel. We prove flat-limit consistency (HSA recovers dot-product attention as $K\to 0^-$), characterise score level sets as horospheres, and bound key-side gradients. HSA is integrated into \textbf{Busemannformer}, a full hyperbolic graph Transformer. Busemannformer-LD reaches \textbf{91.78\%} F1 on Disease-NC; in a controlled score ablation, its log-depth Busemann score achieves the best Disease-NC result among tested scores, reaching \textbf{93.29\%} F1 and performance comparable to the current SOTA Hypformer result with a smaller training budget. For link prediction, we find a score-symmetry principle: Busemannformer-LD with a symmetric score achieves \textbf{90.71\%} AUC on PubMed-LP ($+$10.3 pp over the asymmetric variant), while the asymmetric score performs best on Disease-LP among our variants and outperforms all Euclidean baselines. Ablations identify log-depth Busemann scoring as the key component and reveal a clear failure mode on Airport, where labels track centrality rather than tree depth.
Chat is not available.
Successful Page Load