NovaMart: A Causally Consistent Simulated Enterprise for Measuring Tribal Knowledge Extraction
Abstract
AI agents are moving into the daily work of enterprise engineers and analysts: building pipelines, answering metric questions, changing production code. Doing any of this well requires the company's tribal knowledge: which undocumented filters go into a metric, which of several views is canonical, which half-finished migration decided whose data moved. Agents that lack it fail silently, shipping numbers that look right while disagreeing with every existing report. Evaluating this capability is hard because the knowledge is spread across adjacent sources (codebase, commit history, databases, deployment logs, historical queries, dashboards), artifact classes companies share, if at all, only as stripped metadata; as of our survey (July 2026), no benchmark covers them together. We present NovaMart: a simulated e-commerce company covering all of these sources mentioned earlier, causally connected because its entire history is executed rather than authored on real behavioral traffic; the full estate is released with 51 audited tribal-knowledge claims. Evaluating three widely used agent systems, we find what separates them is excavation, knowledge computed from artifacts rather than read from narration, with the claims that require runtime logs separating the systems most sharply.