UNIARTIFACT: Benchmarking Unified Cross-Artifact Retrieval
Abstract
A research paper is often released with a poster, slides, figures, tables, and code. These artifacts describe the same work, yet retrieval systems search them separately, and no single retriever spans all of their relations. We present UniArtifact, a framework and public benchmark for unified cross-artifact retrieval: a query can be a caption, poster panel, table region, README, source file, or any observed unit, and the task is to find the matching unit of a requested artifact type under one instruction-conditioned scoring function over a shared candidate pool. We use the paper as a hub for its artifacts. Official links and document structure provide direct pairs, such as a figure with its caption or a repository with its paper, which provide supervision; two direct relations composed through the paper give derived pairs, such as an abstract with its poster or a poster with its repository, which we hold out to test whether a model learns a shared retrieval space rather than only memorising which artifacts belong to each paper. The benchmark contains 1,500 license-checked papers from ICLR, ICML, and NeurIPS with a binding graph of 138,186 edges, including 15,372 figure-caption and 14,706 table-caption pairs, 989 posters, 418 slide decks, 1,016 repositories, 1,315 models, and 932 datasets. We benchmark eight open encoders together with BM25 and a per-direction oracle to measure how hard the task is: lexical matching (BM25) is the strongest method on the text-bearing directions (macro nDCG@10 0.65, up to 0.99 on abstract-to-full-text), yet purely visual and layout-bound directions have no method above chance (worst-direction nDCG@10 approximately 0), showing that the difficulty of UniArtifact is concentrated in cross-modal retrieval that current systems cannot handle.