Can Machines Foresee the Future of AI Research? GapBench-3: Predicting the Next AI Experiment on a New 7.3-Million-Edge Map
Damià Vicens Ramis ⋅ Nikolas Rieger ⋅ Sparsh Tyagi ⋅ Samir Jusufi
Abstract
Artificial intelligence (AI) research is the first scientific field whose own literature is large, structured and fast-moving enough for machines to attempt a genuinely new task, predicting which experiment the field will run next. We present GapBench-3, the first benchmark that forecasts scientific experiments, defined as (method, dataset, task) configurations rather than pairwise links. It rests on an analysis-ready map of the field that we collected, linked, audited and release: $680{,}502$ papers, $7.29$ million dated typed connections and $5.07$ million resolved citations, joining the frozen Papers-with-Code archive to a delta we harvested ourselves. The map's final window is a real, unobservable future, created by that archive's shutdown. The experiment level turns out not to be the closure of the pair level: only $29\%$ of new configurations complete a triangle whose three pairs already exist, and on the sealed future only $11\%$ do. Benchmarks at this level also silently reward memorisation, since a trivial counter of already-present pairs outranks a strong untrained text embedding under standard negative sampling, so we contribute the closure-matched negative family that removes this artefact. The deepest result is a law only a sealed future could reveal: predictability is a property of the horizon before it is a property of the model. Under a preregistered, hash-committed seal opened exactly once, the advantage of training, worth $+0.19$ mean reciprocal rank inside the visible horizon, vanishes on the real one, where a frozen 2020 embedding sets the mark that every learned system, ours and fourteen reproduced families, has yet to pass. We identify the mechanism, quantify it, release the instrument with its full provenance, and freeze a capsule of $200$ predictions that the future itself will grade in July 2027.
Chat is not available.
Successful Page Load