EnterpriseFact: Where Frontier Agents Fail in Enterprise Fact-Finding
Abstract
Enterprise LLM agents answer business questions across structured databases and unstructured documents, but aggregate evaluation obscures where multi-fact answers fail. We introduce EnterpriseFact, an atomic-fact diagnostic benchmark spanning 120 questions, 651 facts, two synthetic enterprise environments, and 11 leading models over 3,951 attempts. The strongest model recovers 70.4\% of required facts on average but achieves a full pass on only 24.7\% of attempts. Recovery is higher for direct and absence facts than for joined, policy/rule, and conflict facts. Missing and conflicting evidence produce distinct failures: agents may recognize unavailable information yet introduce unsupported values, while conflicts are often collapsed to one source. For the relocatable facts studied, a controlled counterfactual shows that source heterogeneity alone does not explain the observed cross-modal gap. Fact-level diagnostics reveal where failures emerge as enterprise agents retrieve, connect, and represent information.