Agreement-Tested Coordinates: When Named Features Can Support Discovery, and When They Cannot
Tong Wang ⋅ Yiqing Xu ⋅ Leo Y. Yang
Abstract
Interpretability is increasingly used as a discovery tool: an analyst reads a model's named coordinates and treats them as candidate findings about the domain. This inference is only licensed if the names mean something reproducible. We propose two operational criteria that a named coordinate must satisfy before it can support a discovery claim: \textbf{conceptual clarity}, that independent annotators applying the written definition agree at chance-adjusted $\kappa \geq 0.70$, and \textbf{label disentanglement}, that the coordinate does not merely paraphrase the prediction target. We instantiate them in LLM-assisted Feature Discovery (LFD), which proposes lexical and semantic features from outcome-opposed contrastive pairs, screens candidates by cross-LLM $\kappa$, and selects by residual held-out gain. Across ten tasks in seven corpora, LFD has nearly identical observed mean accuracy to a strong text-bottleneck baseline ($0.708$ vs.\ $0.707$ balanced accuracy) while producing markedly more reproducible coordinates. Human evaluation ($276$ raters across the clarity and leakage audits) shows that the clarity gap transfers to people: on a task chosen because its correlation profile \emph{favors} the baseline, independent raters reproduce LFD definitions at feature-median $\kappa = 0.70$ versus $0.11$ for baseline concepts, with two baseline concepts falling below chance. We also report where the criteria fail: compositional tasks where no compact named basis captures the signal, rare features whose $\kappa$ is undefined or unstable, and a measured accuracy cost when disentanglement is certified rather than merely observed. The negative cases delimit when interpretability output is safe to read as knowledge.
Chat is not available.
Successful Page Load