RATS! Patches Talk Through Registers: Emergent Parts in Register Attention Transformers
Timing Yang ⋅ Predrag Neskovic ⋅ Jansen Seheult ⋅ Wenchao Han ⋅ Anand Bhattad ⋅ Alan Yuille ⋅ Feng Wang
Abstract
When humans see a bird, they recognize far more than just ``bird'' --- they see a head, wings, and talons, a structured assembly of reusable parts that can be identified across every bird they have ever seen. We ask whether a self-supervised visual model can discover the same compositional structure on its own. To this end, we propose $RATS$ ($R$egister $A$ttention $T$ransformer$s$), which decomposes the classification token into $N$ learnable register tokens that route patch information through an $L{\to}N{\to}N{\to}L$ bottleneck. The $N$ registers are hard-partitioned across $H$ attention heads, structurally isolating each subset in an independent projection subspace. Without auxiliary losses or part annotations, each register spontaneously specializes into a semantically coherent visual part. RATS surpasses all baselines by an average of +12 mIoU on five segmentation benchmarks, and demonstrates stronger dense prediction on ADE20K (+1.11 mIoU) and COCO (+0.2 AP$^{\text{m}}$). The visual dictionary extracted from the trained registers also shows signs of part-level consistency and semantic proximity across related categories. Our results suggest that RATS may provide a useful architectural prior for structured and interpretable visual representation learning.
Chat is not available.
Successful Page Load