Under-reported Harm: A Claim-level Audit of Abstract Reporting on AI Assistance and Human Learning
Abstract
AI assistance is rapidly entering clinical training and practice, and evidence of agents' benefits and harms to human learning is required to develop evidence-based guidelines for governance in academic medical centers. To assess how current evidence is being reported, this study extracted 6,063 atomic claims from the abstracts of 1,402 papers (2015–2026) and coded each into a 19-field predicate–argument schema. Within the same paper, harm claims are hedged at about twice the odds of benefit claims (OR 2.02, 95% CI 1.49–2.72) and marked statistically significant at about one-seventh the odds (OR 0.14, 0.08–0.27), and both differences hold within generative-AI studies. Across the corpus, 82.1% of harm claims report no statistical test in the abstract, against 70.8% of benefit claims. The share of harm claims is inversely correlated with quantitative reporting (Spearman ρ = −0.46, 95% CI −0.59 to −0.11, across research areas): among claims coded beneficial or harmful, the harm share is 39.8% for behavioral-strategy outcomes (8.4% of the corpus) and 37.4% under longitudinal measurement (2.9% of the corpus). These results indicate that abstract-level under-reporting of harms is pervasive in studies of AI assistance in learning.