Pref-DetectGPT: Unveiling Machine-Generated Text via Preference-Aware Curvature Measurement
Jiahao Wang ⋅ Feifei Kou ⋅ Zhongbao Zhang ⋅ Jiwei Zhang ⋅ Lei Shi ⋅ Pengfei Zhang ⋅ Suguo Zhu ⋅ Mingying Xu
Abstract
The widespread deployment of large language models (LLMs) has made reliable detection of machine-generated text increasingly critical. Recent curvature-based detectors introduce perturbations to candidate text and compute curvature as the log probability difference between original and perturbed variants using advanced surrogate models. However, such curvature measurements only capture weak discrepancies between machine-generated and human-written text. We propose $\textbf{Pref-DetectGPT}$, a training-free method that introduces preference-aware curvature measurement by leveraging the implicit preference capabilities of preference-optimized surrogate models. Specifically, we construct an implicit reward model from the preference-optimized policy to score both original and perturbed text, and redefine curvature as their normalized deviation, which serves as a significantly stronger signal for distinguishing machine-generated text from human-written text. Empirical evaluations across multiple public benchmark datasets demonstrate that Pref-DetectGPT achieves state-of-the-art detection performance with relative improvements of 3.49\%, 3.13\%, and 10.66\% in AUROC, AUPR, and TPR5\%, while exhibiting strong robustness against various adversarial attacks and input lengths.
Chat is not available.
Successful Page Load