The Density of States of a Post-Training Update: A One-Bit Certificate and a Ceiling No Fitting Rule Can Breach
Abstract
The posterior minimising a linear PAC-Bayes bound is a tempered Gibbs measure and its value is a free energy; that much is classical. What follows once the energy is an error count on a held-out calibration sample is that the value depends on the prior only through the law of that count — its density of states — in any measurable parameter space. The obstacle to certifying what a post-training run produced is therefore not sampling the Gibbs measure but knowing its density of states, which is immediate for a prior on finitely many models and not for one whose atoms are the model actually deployed. TaskLine-PB gives one: the run keeps the full update, and the prior is the two checkpoints that update already names. The density of states is then two error counts, the Gibbs measure a two-state system charged at most one bit however many parameters moved, and — because one state never reads the fitting split — the free energy inherits a ceiling no fitting rule can breach. On a frozen CLIP ViT-B/32 post-trained from sixteen labelled examples per class, every certificate on eight tasks falls below random guessing, the two-state prior costing 0.0066 on average and 0.0144 at worst against certifying the post-trained prompt alone, with nothing estimated numerically: no draw, no quadrature, no free parameter.