SplitWitness-G: Measuring Train-Visible Evidence in Link-Prediction Benchmarks
Abstract
Link-prediction benchmarks are often treated as tests of whether models can infer missing relations. This interpretation becomes unreliable when positive and negative candidates differ in simpler evidence already visible in the training graph, such as repeated interactions, inverse edges, short paths, popularity, or recency. We introduce SplitWitness-G, an edge-level framework for auditing this hidden evaluation advantage. It represents every test candidate using an interpretable profile of training-visible evidence and applies complementary controls to the negative and positive sides of the benchmark. Negative matching asks whether labels remain distinguishable after positives are paired with valid negatives exposing similar evidence. Positive repair asks whether evidence-heavy positives can be replaced without altering the benchmark’s relation, temporal, and degree structure. Across static, temporal, and knowledge-graph settings, these controls reveal markedly different outcomes: some candidate sets can be balanced effectively, others retain substantial separability, and some cannot be repaired without changing the task itself. SplitWitness-G therefore does not produce a universally corrected leaderboard. Instead, it makes the assumptions behind a benchmark score explicit by showing what evidence candidate construction exposes, what remains after control, and where stronger intervention would compromise evaluation validity.