Beyond Document Relevance: Evaluating Citation-Set Structure in Citation Recommendation
makan farhoodimoghadam ⋅ Avery Louis ⋅ Joshua Bloom ⋅ Karthik Ram
Abstract
When authors of a scientific paper cite several papers at a single point in a text, they are often not aiming to cite the exact same claim repeatedly. Citation sets include a high degree of author judgement --- they are often chosen together to cover a conceptual space, reference both a problem and its main methods, or pair a method with a result it is usually compared to. Despite this, citation recommendation, as it is usually posed, ignores this set-level structure, and instead scores each target in isolation to maximize its independent semantic similarity to the citing sentence. We show that in the case of recommending a citation set, this may be the wrong approach. We introduce LOCUS, an evaluation that determines whether or not a retriever can tell which of a paper's own citations belong at one sentence rather than another one in the text, and that looks beyond the use of topic-level similarity. A simple, training-free co-citation reranker built from past citation structure improves Mean Reciprocal Rank (MRR) by $+0.102$ to $+0.121$ across four retrievers, three of which are citation-trained. For SPECTER2, the improvement is $+0.118$. The gain is more concentrated in the 61.2% of pools with target--anchor graph support. Further, once one citation is known, the retriever is $3.4\times$ more effective at recovering the remaining members at Recall@$1$.
Chat is not available.
Successful Page Load