What Does Held-Out Accuracy License? A Stated-Inference Evaluation Protocol for Few-Shot Prompt Tuning
Keisuke Yokota
Abstract
A few-shot adaptation of a frozen vision–language model is reported through a held-out accuracy computed on a split drawn from the same scarce budget that produced the adaptation, and spent again whenever anything is chosen on it. Such a number does not state what it is an estimate of, what sampling law identifies that quantity, what the selection it funded costs, or with what confidence any of it holds. We audit that protocol on one widely used pipeline—CoOp-style soft-prompt tuning of a frozen CLIP—and put in its place one that states all four and charges for the selection. A fitting split produces the full 2048-dimensional prompt update, kept intact rather than compressed or projected; an independent calibration split then fixes only the odds with which a stochastic predictor draws the fitted prompt rather than the pretrained one, and it is that predictor's class-balanced risk that is reported, under a sampling law we state rather than leave to the word "split". No search stands between the counts and the number: the best certificate over every posterior on the line has a closed form, which at the two named ends reduces to a formula in two error counts, so no grid is resolved, no draw is taken and no hyperparameter is set at the certification step. Sixteen labelled examples per class put every certificate on eight CLIP tasks below random guessing, at a cost of 0.0066 on average and at most 0.0144 against the certificate that reports the fitted prompt alone—the price of the construction, given in place of the smaller number it replaces. Three findings concern the protocol rather than the method: what a task can conclude is set by the calibration sample $n = Km_2$ and not by the per-class shot count that few-shot benchmarks standardise on, so a hundred-class task concludes from two examples per class what a ten-class task needs twenty-four for; reading a published certified number beside ours takes five agreements, and two differences that matter survive them; and a corruption ladder on which our certificate moves by 0.064 where the fitted-endpoint one moves by 0.743 is measuring how much of the fit survives rather than whether the bound does, because a theorem settles in advance what such a ladder would otherwise be asked to settle.
Chat is not available.
Successful Page Load