Plausible Is Not Ready: Stage-Resolved Validation of Synthetic Training Data for a Low-Resource Heritage Corpus
Abstract
Synthetic imagery is a standard answer to scarce annotated data, and it is usually validated by how it looks. We test that practice on Paleo-Hebrew stamp seals, where only 307 annotated photographs exist and six off-the-shelf Hebrew OCR systems produce zero exact matches. We release PaleoHebrew-Seals, a real benchmark of 2,912 localised sign boxes aligned to scholarly transcriptions, together with 200,000 synthetic images generated in two separable stages, exact procedural rendering and then diffusion stylisation under structural conditioning that preserves the inherited labels. Evaluating both stages across seven detectors and sixteen classifier backbones on a leakage-audited real split, we find that the same corpus improves detection for five architectures by up to 0.440 mAP50-95 and degrades two by up to 0.107, while classifier effects range from -3.23 to +65.75 points of macro-F1. Readiness is therefore a joint property of a corpus, a downstream portfolio and a real split, and it cannot be read off appearance or off a single model. We release the metadata, splits and audits that make it checkable.