Repair or Recall? Measuring Document Faithfulness with Counterfactual PDF Perturbations
Kenan Ak ⋅ Gwang Gook Lee ⋅ Jay Mohta ⋅ Yan Xu ⋅ Dimitrios Dimitriadis
Abstract
Vision-language models may answer document questions either by reading the page they are given or by recalling a familiar fact from training data, yet standard benchmarks score both behaviours identically. We introduce WikiPDF-VQA, a counterfactual benchmark that perturbs answer-bearing table cells in Wikipedia articles rendered as multi-page PDFs and scores models against the displayed document. Across 8 open-weight models on randomly sampled articles, corrupting an answer into a non-word reduces accuracy from $57.1\%$ to $35.5\%$, whereas replacing it with a different real value leaves accuracy at $58.6\%$. To separate memory from a linguistic repair prior, we probe each fact closed-book, across both our random and our popular article sets. Even when a model cannot recall a fact without the document, it replaces a typo with the original value $16.6\%$ of the time. When the model can recall the fact, this rate rises to $38.6\%$, showing that both the repair prior and memorised knowledge can override the document. In the one model we test at two visual-token budgets, higher resolution reduces memory-driven overrides but leaves knowledge-free repair unchanged. Improving document faithfulness therefore requires suppressing both memory-driven overrides and knowledge-free repair.
Chat is not available.
Successful Page Load