Citation Receipts: A Protocol for Verifiable AI Citation Counts
Abstract
Search-enabled language models can cite papers that did not exist when the model was trained, yet publishers and authors cannot verify how often those papers appear in generated answers. We present the Citation Receipt Protocol (CRP), in which a model provider signs one receipt for each visible citation backed by an authenticated fetch. A receipt binds the source identifier and version, exact content hashes, license and fetch references, a unique generation event, and a commitment to the rendered response. Hourly receipt sets are committed to an append-only Merkle log. Auditors submit queries that look like ordinary traffic; under probe indistinguishability and independent suppression at rate q, the chance that n probed citations all pass is (1 - q)^n. Thus, 299 clean probes exclude suppression of at least 1% with a one-sided 95% bound. The implementation passes 372 tests. In a controlled experiment with 20 arXiv papers first submitted in July 2026, after Claude Sonnet 5's documented January 2026 training-data cutoff, the model cited no target paper in 100 calls without retrieval. With controlled retrieval it cited the target in 95 of 100 natural-retrieval calls and 95 of 100 oracle calls. CRP issued and verified all 190 eligible receipts, with no out-of-context receipts or event-attribution failures. CRP does not attribute knowledge stored in model weights or judge whether a citation supports a claim. It supplies evidence from which a paper-level "cited by AI" count can be built.