How Small Can We Go? Calibrating Activation-Cache Compression for SAE Training
Abstract
To train a sparse autoencoder (SAE), you need a massive cache of model activations. We look into the effects that compression has on an SAE in two ways: we keep the original fp32-trained SAE and pass the compressed cache into the same SAE to see how many feature activations stay the same (the frozen-SAE test), and we retrain an SAE on the compressed cache and see how many of the original features still exist (the retraining test). We found that in the retraining test, there is no clear scale for retained features. On a GPT-2 cache, rounding everything to fp16 keeps 60.6% of features, but changing one value in a million by the smallest step fp32 allows (a change eight orders of magnitude smaller) keeps only 70.5%. Meanwhile, changing the random seed keeps 23.0%. For the frozen-SAE test, the percentage of features from the codecs that match the original increases in order of how precise the codecs are, but plain L2 distance can tell us the same result. We also see the same frozen reading correspond to different percentages retained at different sites. So, what is the correct measurement? We build one by calibrating: retraining on identical bytes sets the upper bound of reasonable variance, and a reseed sets the floor. We then measure the effect of each compressed codec against its own site's floor. With TopK SAEs on GPT-2 small and Pythia-1.4B at research-scale budgets, naive per-token int4 loses more features than a reseed everywhere we tested, as do PCA and random projection on GPT-2. Standardizing each channel before quantizing brings int4 back above the floor in terms of retained features, at the same file size. A bf16 forward pass, which many pipelines use for speed, loses slightly more features than rotated 8-bit storage. Calibrating a site costs four training runs and two evaluation passes.