Concept frustration: Aligning human concepts and machine representations
Abstract
Aligning human-interpretable concepts with the internal representations of modern machine learning systems remains a central challenge for interpretable AI. We introduce a geometric framework to compare supervised human concepts with unsupervised representations derived from foundation models. We formalise concept frustration: a mismatch that arises when an unobserved concept induces relationships between known concepts that cannot be made consistent within an existing ontology. We develop task-aligned similarity measures that detect this phenomenon, and show that frustration is identifiable in task-aligned geometry while conventional Euclidean comparisons fail. Under a linear–Gaussian generative model, we derive a closed-form expression for Bayes-optimal concept-based classifier accuracy, decomposing predictive signal into known and unknown contributions and identifying where frustration impacts performance. Experiments on synthetic data and real language and vision tasks demonstrate that frustration is present in foundation model representations and that incorporating missing concepts reorganises learned representations to better align human and machine reasoning. These results provide a principled framework for diagnosing incomplete concept ontologies and improving alignment in interpretable AI systems.