When “High Quality” Tracks “In Domain”: A Data-Readiness Case Study in Technical Retrieval
Abstract
Data-readiness pipelines increasingly use language models to score whether a corpus is fit for training, then schedule the resulting quality tiers across adaptation stages. We re-analyse an industrial retrieval study in Korean power engineering, a specialized technical domain with limited in-language text, in which a 1B-token corpus was scored by an LLM-assisted rubric, split into low, medium and high tiers, and used for staged continued pretraining followed by staged contrastive fine-tuning of a multilingual retriever. Recomputing the corpus composition from the reported tables, we find that the readiness tiers are substantially confounded with source domain: power-engineering abstracts make up 34.0\% of the corpus but 55.2\% of the high tier, while general-purpose sources are depleted by half. Recent work at web scale reports that quality filters track stylistic and domain similarity as much as quality in any absolute sense. Here the confound is present by construction, because the rubric carries an explicit domain-relevance criterion, and it can be related to a retrieval outcome. Consistent with this, the high tier alone reaches 0.92 HitRate@10 against 0.93 for the full staged schedule, at 40\% of the pretraining tokens, and pseudo-perplexity falls monotonically across the pretraining stages while HitRate@10 does not. We report these as case-specific observations, bounded by an experimental record that gives neither seeds nor per-query outcomes.