Can Machines Foresee the Future of AI Research? GapBench-3: Predicting the Next AI Experiment on a New 7.3-Million-Edge Map
Damià Vicens Ramis ⋅ Nikolas Rieger ⋅ Sparsh Tyagi ⋅ Samir Jusufi
Abstract
Artificial intelligence (AI) research is the first scientific field whose own literature is large, structured and fast-moving enough for machines to attempt a genuinely new task, predicting which experiment the field will run next. We present GapBench-3, the first benchmark that forecasts scientific experiments, defined as (method, dataset, task) configurations rather than pairwise links. It rests on a map of the field that we collected, linked and release: 680,502 papers, 7.29 million dated typed connections and 5.07 million resolved citations, joining the frozen Papers-with-Code archive to a delta we harvested ourselves. The map's final window is a real, unobservable future, created by that archive's shutdown. The experiment level turns out not to be the closure of the pair level: only 29% of new configurations complete a triangle whose three pairs already exist, and on the sealed future only 11% do. Benchmarks at this level also silently reward memorisation, since a trivial counter of already-present pairs outranks a strong untrained text embedding under standard negative sampling, so we contribute the closure-matched negative family that removes this artefact. The deepest result is a law only a sealed future could reveal: predictability is a property of the horizon before it is a property of the model. Under a preregistered, hash-committed seal opened exactly once, the advantage of training, worth $+0.19$ mean reciprocal rank inside the visible horizon, vanishes on the real one, where a frozen 2020 embedding sets the mark that every learned system, ours and fourteen reproduced families, has yet to pass. We identify the mechanism, quantify it, release the instrument, and freeze a capsule of 200 predictions that the future itself will grade in July 2027.
Chat is not available.
Successful Page Load