CURE: Visual Reprogramming of Vision-Language Models under Limited Supervision
Abstract
Visual reprogramming (VR) efficiently repurposes pre-trained vision-language models for new image classification tasks by adding trainable patterns to inputs without modifying the backbone. Nevertheless, existing VR methods heavily depend on labeled data and thus suffer substantial performance degradation under limited supervision (i.e., scarce labeled images). We introduce semi-supervised learning (SSL) into VR to exploit abundant unlabeled images and enhance data efficiency. However, directly applying generic SSL techniques often amplifies biases in unlabeled posteriors and destabilizes VR training. To address these challenges, we propose CURE (Consistency and Unlabeled Recalibration with Confidence-Margin Enhancement), a semi-supervised VR framework comprising two key components: (i) Recalibreted Gaussian Confidence-Margin Soft Weighting to dynamically adjust sample importance; (ii) Dual-attribute Consistency Regularization to ensure consistency across attribute prompts. Extensive experiments show that CURE consistently outperforms both existing VR methods and direct SSL extensions, improving average classification accuracy by 5.06 and 6.67 percentage points across twelve widely-used benchmarks for limited supervision tasks. Our code is available at https://anonymous.4open.science/r/CURE-93E2/.