DAGA: Dynamic Attention-Guided Adaptation for Self-Supervised Vision Transformers
Abstract
Self-supervised Vision Transformers in the DINO family encode rich semantic structure within their self-attention maps, yet existing parameter-efficient fine-tuning (PEFT) methods apply static, content-agnostic transformations that ignore this internal signal. We present DAGA (Dynamic Attention-Guided Adaptation), a PEFT module that repurposes a frozen backbone's emergent attention maps as instance-specific guidance: per-image attention is converted into channel and spatial gates that condition a lightweight bottleneck adapter. With only 1.4\% additional parameters, DAGA reaches 85.9\% top-1 on ImageNet-1K with DINOv3-ViT-B and surpasses recent PEFT baselines, with strong gains on dense prediction (+12.1\% mIoU on ADE20K) and fine-grained classification (+1.6 / +3.5 on CUB-200 / FGVC-Aircraft). DAGA's effectiveness is tied to the semantic quality of the backbone's attention: gains are large on DINOv2/v3 and iBOT but small on MAE, CLIP, and DeiT, positioning it as the first PEFT module to operationalize emergent SSL attention as in-loop guidance and complementing self-distillation methods that refine attention itself. Anonymized code: https://anonymous.4open.science/r/ID1680.