SAEScientist-Bench: Can AI Agents Conduct Autonomous SAE Interpretability Research?
Abstract
Understanding how trained models work and how their behavior can be changed is an important part of AI research and development. Sparse autoencoders (SAEs) support this research by exposing features within language models that can be examined and manipulated. We introduce SAEScientist-Bench to evaluate whether AI agents can autonomously design experiments to discover SAE features for model interpretation and steering. Given a concept, an agent designs probe texts, compares candidates, and selects a feature from a Gemma Scope dictionary of 131K+ features in Gemma-2-9B-IT. The evaluator measures activation rank within the SAE dictionary, separation of relevant and control texts, and changes in generated answers, with Expert features providing a baseline. Across ten agent configurations and 20 tasks, different agents lead Rank, Activation, and Steering. Their investigations reveal different approaches to probe design, candidate comparison, and feature selection, including language–topic contrasts and tests across examples. Generation tests connect these selections to changes in answers. Some selections change an answer's language while preserving its content, whereas others replace the requested content with the target topic. SAEScientist-Bench makes model analysis and intervention a measurable part of automated AI R&D.