Concept-Based Mechanistic Interpretability Needs a Concrete Evaluation Paradigm
Abstract
This position paper argues that concept-based mechanistic interpretability lacks a concrete evaluation paradigm and that the field must converge on one before continuing to scale methods whose effectiveness remains unverified. Sparse Dictionary Learning (SDL), encompassing sparse autoencoders, transcoders, crosscoders, and their variants, has become the dominant methodology for concept-based mechanistic interpretability, supported by a growing evaluation suite: reconstruction loss, automated interpretability scoring, human evaluation, feature absorption rates, sparse probing, and model editing benchmarks. Despite the significant empirical success SDL has achieved, we argue that none of these evaluation metrics directly measures what SDL is actually supposed to achieve: recovering the ground-truth features underlying observed representations. This insufficiency is structural, not incidental, and we support this perspective with both theoretical and empirical evidence. On the theoretical side, we build on the most recent formal analyses of SDL in mechanistic interpretability to show that the optimization landscape provably admits failure modes that existing metrics cannot detect. On the empirical side, we present three lines of evidence demonstrating that methods spanning a wide range of ground-truth recovery performance appear indistinguishable under current evaluation practice. Together, these results call for a new evaluation paradigm. We propose a set of principles that an ideal evaluation benchmark for concept-based mechanistic interpretability should satisfy, and argue for the community to develop and adopt more benchmarks grounded in these principles before continuing to scale methods whose effectiveness remains unverified.