Patch Hierarchical Attention Transformers for Efficient Particle Jet Tagging
Abstract
Real-time jet tagging is critical for identifying short-lived particle decays in the high-throughput detectors of the Large Hadron Collider, where real-time trigger systems which are responsible for deciding which collision events to store impose strict latency and accuracy constraints. Transformer architectures achieve the highest jet tagging accuracy when compute is unconstrained, but their quadratic self-attention cost places them orders of magnitude beyond the trigger budget. Existing efficient variants reduce this cost by compressing the attention matrix or restricting it to ordered local windows, at the price of the explicit particle-particle interactions that drive substructure identification. To address this limitation, we introduce the Patch Hierarchical Attention Transformer (PHAT-JeT), which combines two mechanisms: a physics-inspired geometric message-passing module that encodes local detector-plane structure, and a hierarchical patch-based attention scheme that computes exact attention within small particle groups while preserving global context through lightweight patch-token communication. Within this compute budget, PHAT-JeT achieves state-of-the-art accuracy and background rejection among resource-constrained jet tagging models on four benchmarks (hls4ml, JetClass, Top Tagging, and Quark-Gluon). Our code is available at https://anonymous.4open.science/r/PHAT-JeT-540B/README.md