How Much Does the Wrong SAE Cost? Calibrating Quality Metrics Against Board-Game Ground Truth
Sofia Sartori
Abstract
Sparse autoencoders (SAEs) are widely used to extract interpretable features from black-box models. To use them for discovering new knowledge, it is first necessary to select a dictionary, and this is normally done through proxy metrics (e.g., reconstruction error, $L_0$). In most domains, it is impossible to verify this step, because the knowledge to be extracted is not previously known, which makes it very difficult to tell whether the model did not learn the necessary information or whether the dictionary failed to extract it. Therefore, we used a scenario where it is possible to break this circularity. Using 80 SAEs trained on OthelloGPT, we found that (i) proxy metrics can distinguish healthy dictionaries from collapsed ones with no errors, the collapsed ones recovering the board 41 points of balanced accuracy below the rest, and (ii) the healthy set has a performance range only 5.8 points wide, and although the metrics are not able to rank them perfectly, the shortfall is at most 1.0 point relative to the best available SAE, and three dictionaries drawn at random land within 1.6 points. We argue that rank agreement is the wrong quantity to validate a selection metric, and that the ideal approach is a collapse screen and a shortfall metric rather than a pure ranking. Code and results, anonymized for review: https://anonymous.4open.science/r/interp4discovery-2026-submission-9941
Chat is not available.
Successful Page Load