Why don't Graph Transformers attend to the same token?
Alexandre Bismuth
Abstract
In large language models, attention concentrates heavily on the first token, an `attention sink' believed to protect deep models against over-mixing. Graph Transformers face the same pressure, yet a state-of-the-art graph transformer (GRIT) forms almost no sinks and sinks less as graphs grow, the reverse of the language-model trend. The graph is not the reason: we prove that symmetry and locality cap received attention, trained models obey the symmetry cap exactly on symmetric graphs, and both caps are loose on real molecules. The scoring function decides instead: replacing GRIT's bounded gated score by an unbounded dot product raises its sink rate from $0.09$ to $0.53$ on identical graphs, and a virtual node, exempt from the symmetry cap, sinks strongly under dot-product scoring while GRIT declines it even when sinking is rewarded.
Chat is not available.
Successful Page Load