Coordinated, Not Independent: Dependence and Averaging Choices in AI-Prevalence Evaluation
Aram Bagdasarian
Abstract
A growing class of AI evaluations reports a single population-level estimate of the fraction of text that is AI-generated—for peer review, social platforms, and, under proposed U.S. legislation, federal public comments. Because these pipelines lack held-out ground truth, their credibility depends primarily on the evaluation protocol. We audit one such pipeline end to end using five corpora from three independent settings—including two dated slices of the same federal docket—and PREVALBENCH, a 19-condition benchmark with known ground truth. Our contribution is not to identify new statistical phenomena, but to measure their practical magnitude in a deployed evaluation and specify more defensible reporting practices. First, coordinated campaigns create severe dependence: one campaign constitutes $46%$ of our pooled rulemaking stratum, making submission-level analyses pseudo-replicated. Across the observational corpora, naive confidence intervals are $1.8\times$ to $21\times$ narrower than campaign-clustered intervals; on controlled data, their coverage falls to $13%$ for a nominal $95%$ interval. A campaign-clustered bootstrap substantially improves coverage under extreme concentration, from $13%$ to $73%$, and is conservative elsewhere. A design-effect calculation based only on campaign sizes anticipates the scale of the problem, to within an order of magnitude, before any scores are computed. Second, aggregation defines the estimand: weighting the same document-level scores by words, submissions, or campaigns yields $0.00$, $0.19$, and $0.31$ on one corpus. These values answer different questions rather than estimate the same quantity. The document-level fraction readers may infer is not identifiable from aggregate lexical signals alone, and the mixture model required to interpret it as such is rejected on our data; we therefore abstain from reporting it. Third, coordination and lexical position are distinct signals: a federal docket that is $95%$ coordinated lies near the human-reference end of our lexical axis both before and after ChatGPT’s release, whereas campaigns in an unrelated 2026 docket lie closer to the AI-reference end. We do not claim that this score validly measures AI use. We treat it as a descriptive, reference-relative lexical coordinate and propose four reporting practices that require practitioners to name and bound what their pipelines already compute. PREVALBENCH and our code will be released upon acceptance.
Chat is not available.
Successful Page Load