Vision-TabPFN: An Efficient Tabular Prior for Few-Shot Vision
Abstract
A model trained only on synthetic tables can, it turns out, classify images and does so best precisely where vision-specific few-shot methods break down. We adapt TABPFN, a 10.7M-parameter tabular in-context learner, to few-shot image classification by mapping frozen CLIP embeddings through a learned projection head into TABPFN’s input space, then fine-tuning only its last 2–4 transformer layers (2.3M parameters, 22% of the model). The result is competitive with much larger baselines on standard benchmarks (CIFAR-100, miniImageNet, CUB) and substantially stronger under domain shift: 93.8% on BloodMNIST 5-shot (+18.8 pp over the next-best method) and 96.0% on EuroSAT 5-shot, versus 43.3% and 76.9% for a 300M-parameter in-context classifier (CAML) on the same tasks. TABPFN’s flexible class interface also lets the same checkpoint run at N>5 without retraining, unlike CAML or SNAIL. Tabular in-context inference is therefore not modality- bound: a small amount of episodic adaptation on top of a frozen tabular prior is a cheap, domain-agnostic alternative to training vision-scale few-shot models from scratch.