EMERGE: A Benchmark for Updating Knowledge Graphs with Emerging Textual Knowledge
Abstract
Knowledge Graphs (KGs) are structured knowledge repositories containing entities and relations between them. In this paper, we study the problem of automatically updating KGs over time in response to evolving knowledge in unstructured textual sources. Addressing this problem requires identifying a wide range of update operations based on the state of an existing KG at a given time and the information extracted from text. This contrasts with traditional information extraction pipelines, which extract knowledge from text independently of the current state of a KG. To address this challenge, we propose a generic and extensible pipeline that pairs textual passages with the KG edit operations they induce on a given KG snapshot. We instantiate this pipeline on the Wikidata knowledge graph and the English Wikipedia corpus to create EMERGE, a dataset of 233K Wikipedia passages associated with a total of 1.19 million KG edits across seven yearly Wikidata snapshots from 2019 to 2025. Our experimental results highlight key challenges in updating KG snapshots based on emerging textual knowledge, particularly in integrating knowledge expressed in text with the existing KG structure. These findings position the dataset as a valuable benchmark for future research. The code and dataset are available at https://github.com/klimzaporojets/emerge.