A Guarantee That Costs One Bit: Certified Few-Shot Prompt Adaptation for Deployable Vision–Language Models
Abstract
Before a specialised model is put to work, someone has to say how often it will be wrong. Prompt tuning does not answer that: it adapts a frozen, compact vision–language model for the price of 2048 parameters and a handful of labelled examples per class, and stops there. Certifying the fit is hard not because the fit is poor but because a prior, a posterior, the search over both and a numerical evaluation of the risk must all be stated—and charged for—in that same 2048-dimensional space. The accounting, not the fitting, is what fails to scale. TaskLine-PB certifies not the fitted prompt itself but a choice between it and zero-shot CLIP. One split of the task's own examples learns the full 2048-dimensional update, kept intact rather than compressed or projected; an independent split fixes only the odds with which a stochastic predictor draws it rather than the pretrained prompt, and that predictor's class-balanced risk carries the guarantee. Relative entropy is unchanged by the map from the coordinate to the prompt, so the bound pays for the prompt that is deployed and not for a proxy; the best certificate over every posterior on the line has a closed form, which at the two named ends is a formula in two error counts; and the whole adaptation costs one bit, with nothing left to estimate and no parameter to tune. However badly the fitting split is corrupted, the certificate cannot rise above the branch that never reads it—a theorem, not an observation. Sixteen labelled examples per class put every certificate on eight CLIP tasks below random guessing, at a cost of 0.0066 on average and at most 0.0144 against a certificate of the fitted prompt alone; to our knowledge this is the first non-vacuous PAC–Bayes risk certificate for few-shot soft-prompt adaptation of a frozen vision–language model without auxiliary-task checkpoints, an upstream task collection, or a learned generator.