Text-Based AI Tools for Research Integrity Must Be Audited on Linguistic Fairness Before Deployment
Shuai Shao ⋅ Yongkang Wan ⋅ Daoyin Dang ⋅ Lanyun Zhu ⋅ Di Yang ⋅ Yutong Bai ⋅ Yan WANG ⋅ Jiangtao Wang
Abstract
Research integrity is the cornerstone of scientific progress. The academic community has increasingly adopted text-based Artificial Intelligence (AI) tools to automate research integrity detection, yet whether these tools are fair to researchers from different linguistic backgrounds has rarely been examined. If detection tools themselves harbor linguistic bias, false positives will damage researchers' academic careers and subject entire research communities to unfair treatment. To expose this risk, we conduct a case study of a paper mill (organizations that mass-produce fraudulent academic manuscripts for sale) detection model published in \textit{The BMJ} (impact factor $=$ 43) in January 2026. Through independent reproduction and controlled experiments, we find that the model's predictions are substantially influenced by linguistic style: legitimate non-native English papers from high-impact journals receive a mean paper mill probability of 42.0\%, versus 2.0\% for native English papers. Large Language Model (LLM) based controlled experiments further show that switching writing style alone can flip predictions from negative to positive, and the reverse never occurs. Through further analysis, we argue that these biases are not an isolated case: unavoidable geographic bias in training data, combined with the inherent tendency of language models to encode linguistic style, makes linguistic bias an intrinsic risk for all such tools. Based on these findings, our position is that \textbf{text-based AI tools for research integrity must undergo mandatory fairness auditing for linguistic bias before deployment}. We propose concrete auditing standards and call on the AI community to establish norms that balance detection effectiveness with linguistic equity.
Chat is not available.
Successful Page Load