Small Models Are Views, Not Artifacts: The Case for Random-Access Compressed Model Storage
Abstract
Deploying a small model into a regulated or safety-relevant setting requires answering a question the current ecosystem cannot answer: is this artifact the one that was evaluated? Every base model is republished distilled, quantized to several precisions, and fine-tuned into thousands of variants, each an independent file with no verifiable relationship to its parent — only 25.8% of Hugging Face repositories carry usable lineage metadata at all. We argue that this matrix should instead be a single self-describing, block-indexed, multi-precision store, and that a small model should be a lossless random-access read over it. Losslessness is what makes this a trustworthiness argument rather than a storage one: it separates is this the audited artifact? from does the evaluation transfer to this precision?, answering the first with a bit-exact guarantee and per-block cryptographic attestation, and leaving only the second open. The obstacle has never been that argument but the belief that random access destroys the compression which makes such a store worthwhile. We show it does not. Every primitive already exists in the recent literature; what is missing is the artifact composing them. A measurement study on Qwen2.5-0.5B finds that self-describing random-access blocks match or beat a monolithic stream at 64 rows per block and coarser, a third precision level saves a further 13.7% of stored bits, per-block attestation costs 0.004 bits per weight, 49% of the file yields a model within 0.06% of full-precision perplexity, and the correct block granularity is analytically predictable from entropies alone (ρ = 0.90).