SD-LoRA: Training-Time Structural Distillation into LoRA for Few-Shot Vision-Language Adaptation
Abstract
Few-shot adaptation of vision--language models must balance two competing goals: learning fine-grained visual distinctions from very limited supervision, and preserving the efficiency that makes frozen CLIP-like models attractive for deployment. Existing lightweight adaptation methods, including prompt tuning, LoRA, and cache-based adapters, typically supervise global image embeddings and therefore underuse token-level structure. Conversely, structure-aware methods can exploit local evidence but often introduce additional inference-time computation. We propose SD-LoRA, a structural distillation framework that uses token-level graph reasoning only during training and absorbs its effect into a lightweight LoRA-adapted CLIP model. During training, a heterogeneous graph teacher models patch-patch, global-local, and visual-text relations over high-resolution ViT tokens. Rather than keeping this graph at test time, SD-LoRA transfers its structural bias into the shared CLIP-LoRA backbone through implicit distillation and Prototype Predictive Alignment, which aligns graph-conditioned representations with support-derived class prototypes. At inference, the entire graph teacher is discarded; prediction uses only a single CLIP-LoRA forward pass and a cache-based classifier. Across 11 few-shot benchmarks, SD-LoRA improves over a strong CLIP-LoRA baseline by 1.3/1.2/2.4/1.9/2.0% in the 1/2/4/8/16-shot settings, while adding no graph computation at test time. Further results on base-to-novel generalization, cross-dataset transfer, OOD robustness, and calibration show that training-time structural supervision improves not only accuracy but also transfer reliability.