Provenance and Perceived Humanness in Social Media Political Discourse
Abstract
Is this post written by a human? As generative AI becomes embedded in everyday writing, that question is increasingly difficult to answer from text alone. We find that frontier models can recover the source of short social-media posts accurately when confident, but remain confident on only a small fraction of content. We therefore ask a different question: how human does a post seem? Using pairwise human judgments, we model perceived humanness across four major social-media platforms surrounding the U.S.\ presidential debate. Lightweight models trained on extreme-PHI examples maintain strong ranking performance while requiring substantially less labeled data. Across platforms, more human-like content is consistently more readable, uses shorter words, and is more negative in sentiment, alongside platform-specific linguistic patterns. Together, our results provide a scalable framework for studying perceived humanness in social media discourse: when provenance cannot be known, perception can still be measured.