AI Construct Lexis: An Ontology of the Hidden Assumptions in AI Evaluation
Abstract
There are now thousands of measurement instruments used to evaluate language models, and the resulting scores are routinely used to support consequential claims about capabilities and risks. Yet, the evidence linking benchmark scores to constructs (i.e., abstract concepts such as reasoning ability or safety) is often implicit, incomplete, or absent. We introduce AI Construct Lexis, a versioned ontology and interactive artifact that extracts and organizes relationships among constructs, behaviors, and measurement instruments as claimed in the literature for LLM evaluation. Drawing on over 8,337 published papers from selected top AI venues, we use language models to extract constructs, definitions, behavioral indicators, and measurement instruments, yielding 977 distinct constructs, 1,694 measurement instruments, 830 behavioral indicators, and 5,900 relationships. The Lexis captures relationships claimed in the published literature, not necessarily validated relationships; valid score interpretation still requires empirical evidence that the instruments capture the constructs they are used to support. By exposing the field’s implicit measurement assumptions, the Lexis enables their inspection, testing, falsification, and refinement into evidence-based relationships. On expert-annotated samples, the extraction pipeline achieves high agreement on extracted concepts while revealing substantial ambiguity in the boundaries of constructs and behaviors. By making these relationships explicit, AI Construct Lexis provides infrastructure for a more rigorous and cumulative science of AI evaluation.