GT-Free OCR Metrics: A Reference-Free Evaluation Framework for Document OCR Systems
Kshitij Singh
Abstract
Evaluating document OCR systems traditionally requires ground-truth (GT) text annotations aligned to each recognised element. Producing such annotations demands expert human labour, scales poorly with corpus size, varies between annotators on ambiguous content, ages quickly as document distributions and OCR systems evolve, and is often infeasible for sensitive or proprietary material. These constraints prevent continuous OCR quality monitoring at production scale and force practitioners to rely on small audited samples or proxy heuristics. To overcome these inherent limitations of ground-truth-based evaluations, we present a new class of reference-free OCR metrics based on render-and-compare that do not require any ground-truth annotations: OCR output is rendered back into a page image and the result is compared visually against the masked original using no-reference or learned metrics. The underlying principle is straightforward: the original page image already serves as a pixel-level ground truth for the content the OCR system tries to extract (text, formulas, and tables), so visual agreement between the original and a re-rendering of the OCR output is itself a proxy for recognition fidelity. To our knowledge, no prior work has systematically benchmarked reference-free visual metrics as proxies for whole-document OCR quality at scale. We conduct the first such study, evaluating 147 method implementations spanning 15 broad categories of techniques on 1,355 OmniDocBench pages across 5 OCR-output variants (text-only, formula-only, table-only, combined, combined without image masking). We measure every method by Spearman rank correlation against ground-truth-based metrics (edit distance for text and formulae, TEDS for tables), quantifying whether a reference-free metric correctly orders pages by OCR quality without any annotations. A naïve CLIP-cosine full-page similarity baseline achieves Spearman correlation $\rho = 0.339$; our best composite method reaches $\rho = 0.494$ (mean across the 5 variants; significance $p < 0.001$, computed over $N = 1{,}355$ pages), with a per-variant peak of $\rho = 0.605$ on the formula-only variant. Both numbers indicate that reference-free OCR quality estimation is genuinely viable: a meaningful page-level ranking signal can be recovered without any ground-truth annotations, and the gap to the ground-truth-based ranking is consistent with the inherent noise of a multi-domain document corpus where OCR errors are heterogeneous across text, formula, and table content. We also discuss inherent limitations that bound what any reference-free metric can achieve in this setting. To support reproducibility and for future research in this direction we are releasing three public HuggingFace datasets (rendered page pairs, per-element OCR log-probabilities, document-similarity triplets), and a trained document-similarity (DocSim) LoRA head (a DreamSim-style recipe adapted to document page reconstructions) along with the associated code.
Chat is not available.
Successful Page Load