PVFormer: Proper Velocity Transformer for Stable and Scalable Hyperbolic Representation Learning
Abstract
Hyperbolic representations have shown remarkable success on hierarchical and tree-like data, and Transformers have become a central architecture for representation learning. However, a principled construction of hyperbolic Transformers remains challenging because the self-attention layer decomposes into three coupled primitives, transformation, similarity, and aggregation, that should all respect the underlying geometry. Existing hyperbolic Transformer designs are typically built in the Poincar\'e or Lorentz models, where bounded domains, time-like constraints, tangent-space detours, or projection steps make it difficult to keep the full block intrinsic and scalable. To address this gap, we propose \PVFormer{}, to the best of our knowledge, the first Transformer architecture that operates natively in proper velocity (\PV) space. The proper velocity model provides an unconstrained representation of hyperbolic geometry, and \PVFormer{} exploits this structure to redesign the Transformer block in \PV{} space. On the transformation side, we use \PV{} homomorphism transformations to update query and value representations. On the similarity side, we generalize Euclidean inner-product attention with a \PV{} Busemann score inspired by Busemann-based hyperbolic learning and a \PV{} point-to-hyperplane key transformation. On the aggregation side, we replace Euclidean averaging with \PV{} midpoint aggregation and further derive a scalable \PV{} linear attention variant for large graphs. Extensive experiments support the effectiveness of \PVFormer{} across graph, text, and vision benchmarks. The code will be open-sourced once accepted.