A Matched-Budget Audit Framework for Recaptioned Image-Text Supervision Distributions
Giyeong Oh ⋅ Junghun Park ⋅ Yuhan Bae ⋅ Youngjae Yu
Abstract
Recaptioned image-text corpora are now standard for text-to-image (T2I) training, with VLM captioners replacing sparse alt-text by dense descriptions. Yet a recaptioned corpus is not only a collection of captions: it is a supervision distribution induced by a documented captioning policy ($ \pi $), captioner ($ V_c $ ), and source corpus ($ C $). Length-correlated proxies miss caption-register artifacts and downstream T2I benchmarks entangle the corpus with training choices, so this distribution is hard to audit at corpus scale. We introduce a reusable matched-budget audit framework for recaptioned supervision distributions $ D_{\pi,V_c,C} $: at a fixed $ B = 64 $ lexical-unit window it reports a five-axis profile spanning prompt-side coverage, image-conditioned faithfulness, and caption-surface health, with claimed controllable basic units (CBU) as the shared denominator. We instantiate the framework on seven paired comparisons over five public source corpora. Across the four cross-corpus pairs, the released surface raises supported CBU per caption by +3.40 to +6.35 under both Qwen and Gemma Judges, and on CC12M the same framework exposes a long-vs-dense frontier that is consistent across both judges and four budgets. We release the audited multi-source recap corpus together with the audit-artifact bundle.
Chat is not available.
Successful Page Load