What Limited Experience Warrants: A One-Bit Certificate for Few-Shot Revision of a Frozen Vision-Language Model
Abstract
A system that already knows a great deal is given a handful of labelled examples and produces a revised state. How much of that revision does the experience warrant? Held-out accuracy answers only by spending examples that could have been learned from, and spends them again whenever anything is chosen on them. TaskLine-PB answers with a guarantee instead. Few-shot soft-prompt tuning of a frozen vision-language model here moves 2048 parameters; we keep that update whole and split the task's own examples, letting the first half fit the prompt and allowing the second a single decision: the probability with which a stochastic predictor draws the fitted prompt rather than the pretrained one. That probability is the warranted weight; it is that predictor's class-balanced risk that carries the guarantee. The two prompts lie on a line through prompt space, and relative entropy is unchanged by the map from a coordinate on that line to the prompt it names, so the bound pays for the prompt that is deployed and not for a proxy; the best certificate over every posterior on that line has a closed form, which at the two named ends reduces to a formula in two error counts; and the whole revision costs a single bit, with nothing left to estimate: no Monte-Carlo term, no quadrature, no free parameter. Because the pretrained prompt stays in the hypothesis class as a named alternative rather than as a starting point left behind, the certificate cannot rise above the branch that never looks at the fitting split, however badly that split is corrupted — a theorem rather than an observation. A sweep over six budgets leaves all 48 of its cells non-vacuous, two labelled examples per class included — one to fit the direction and one to certify it — and at sixteen the two-atom prior costs 0.0066 on average and 0.0144 at worst against a certificate of the fitted prompt alone; to our knowledge this is the first non-vacuous PAC-Bayes risk certificate for few-shot soft-prompt adaptation of a frozen vision-language model without auxiliary-task checkpoints, an upstream task collection, or a learned generator. On a ladder of corrupted fitting splits that carries real human annotation errors, the warranted weight falls below one half in 55 of 63 conditions. Nothing about the revision was made smaller; only the question the certificate has to answer was.