ToyBench: Ground-Truth Evaluation of SAEs Across Feature Distributions
Oliver Sieweke ⋅ Niclas Küpper ⋅ Kaushik Reddy ⋅ Kola Ayonrinde
Abstract
Sparse autoencoders (SAEs) do not fully recover features from superposition. Each SAE architecture carries assumptions about how concepts are structured and may fail when these are violated. We introduce ToyBench, a benchmark of eight synthetic feature distributions with the ground truth that is unavailable in LLMs, each isolating properties of real-world concepts such as correlations, hierarchy, and manifold geometry. A toy model of superposition compresses features from each distribution into a low-dimensional embedding; we then train SAEs to recover them. Comparing four common SAE architectures, we find none dominates: BatchTopK leads on average, yet even ReLU SAEs rank first on some distributions. Feature detection never exceeds $F_1 = 0.84$, well short of idealized SAEs constructed with perfect knowledge of the true features, which sit above $0.95$. ToyBench offers a cheap diagnostic of architecture-specific trade-offs, and tmos, the library behind it, supports evaluating SAEs on custom distributions.
Chat is not available.
Successful Page Load