Paper deep dive
Can Humans Tell? A Dual-Axis Study of Human Perception of LLM-Generated News
Alexander Loth, Martin Kappes, Marc-Oliver Pahl
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 97%
Last extracted: 4/10/2026, 2:17:30 AM
Summary
This study investigates human ability to distinguish between human-written and LLM-generated news articles using a dual-axis assessment platform called JudgeGPT. Based on 2,318 judgments from 1,054 participants, the authors find that humans cannot reliably distinguish between machine and human text, regardless of the LLM model used. The study identifies two response strategies ('Skeptics' vs. 'Believers') and notes that while domain expertise improves accuracy, cognitive fatigue degrades performance after approximately 30 evaluations. The findings suggest that user-side detection is ineffective, advocating for system-level countermeasures like cryptographic content provenance.
Entities (5)
Relation Signals (3)
JudgeGPT â generatesdatafor â Human Perception Study
confidence 100% ¡ JudgeGPT is a web-based study platform for collecting human perception data on AI-generated text
RogueGPT â producesstimulifor â JudgeGPT
confidence 100% ¡ stimuli are generated by RogueGPT... which orchestrates six LLMs
Alexander Loth â authored â JudgeGPT
confidence 95% ¡ we developed JudgeGPT (Loth et al., 2026b)
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Can humans tell whether a news article was written by a person or a large language model (LLM)? We investigate this question using JudgeGPT, a study platform that independently measures source attribution (human vs. machine) and authenticity judgment (legitimate vs. fake) on continuous scales. From 2,318 judgments collected from 1,054 participants across content generated by six LLMs, we report five findings: (1) participants cannot reliably distinguish machine-generated from human-written text (p > .05, Welch's t-test); (2) this inability holds across all tested models, including open-weight models with as few as 7B parameters; (3) self-reported domain expertise predicts judgment accuracy (r = .35, p < .001) whereas political orientation does not (r = -.10, n.s.); (4) clustering reveals distinct response strategies ("Skeptics" vs. "Believers"); and (5) accuracy degrades after approximately 30 sequential evaluations due to cognitive fatigue. The answer, in short, is no: humans cannot reliably tell. These results indicate that user-side detection is not a viable defense and motivate system-level countermeasures such as cryptographic content provenance.
Tags
Links
- Source: https://arxiv.org/abs/2604.03755v1
- Canonical: https://arxiv.org/abs/2604.03755v1
Trouble viewing inline? Open PDF directly â
Full Text
22,278 characters extracted from source content.
Expand or collapse full text
by Can Humans Tell? A Dual-Axis Study of Human Perception of LLM-Generated News Alexander Loth alexander.loth@stud.fra-uas.de 0009-0003-9327-6865 Frankfurt University of Applied SciencesFrankfurt am MainGermany , Martin Kappes kappes@fra-uas.de 0000-0002-8768-8359 Frankfurt University of Applied SciencesFrankfurt am MainGermany and Marc-Oliver Pahl marc-oliver.pahl@imt-atlantique.fr 0000-0001-5241-3809 IMT Atlantique, UMR IRISA, Chaire Cyber CNIRennesFrance (2026) Abstract. Can humans tell whether a news article was written by a person or a large language model (LLM)? We investigate this question using JudgeGPT, a study platform that independently measures source attribution (human vs. machine) and authenticity judgment (legitimate vs. fake) on continuous scales. From 2,318 judgments collected from 1,054 participants across content generated by six LLMs, we report five findings: (1) participants cannot reliably distinguish machine-generated from human-written text (p>.05p>.05, Welchâs t-test); (2) this inability holds across all tested models, including open-weight models with as few as 7B parameters; (3) self-reported domain expertise predicts judgment accuracy (r=.35r=.35, p<.001p<.001) whereas political orientation does not (r=â.10r=-.10, n.s.); (4) clustering reveals distinct response strategies (âSkepticsâ vs. âBelieversâ); and (5) accuracy degrades after approximately 30 sequential evaluations due to cognitive fatigue. The answer, in short, is no: humans cannot reliably tell. These results indicate that user-side detection is not a viable defense and motivate system-level countermeasures such as cryptographic content provenance. LLM-Generated Text, Human Perception, Misinformation Detection, Dual-Axis Assessment, Fake News, Cognitive Fatigue, Content Provenance, Web Trust â journalyear: 2026â copyright: câ doi: 10.1145/3795513.3807431â conference: 18th ACM Web Science Conference; May 26â29, 2026; Braunschweig, Germanyâ booktitle: 18th ACM Web Science Conference (WebSci Companion â26), May 26â29, 2026, Braunschweig, Germanyâ isbn: 979-8-4007-2492-3/2026/05â ccs: Human-centered computing Human computer interaction (HCI)â ccs: Information systems Web miningâ ccs: Security and privacy Social aspects of security and privacy 1. Introduction Text has long been the webâs default trust carrier, yet LLMs now produce synthetic text that humans rate as no less credible than human-written content (Kreps et al., 2022; Clark et al., 2021). As false and misleading content already spreads faster than corrections online (Vosoughi et al., 2018), the ability to generate convincing fabrications at marginal cost amplifies what Wardle and Derakhshan term âinformation disorderâ (Wardle and Derakhshan, 2017) across an already fragile ecosystem (Ferrara, 2024; Lazer et al., 2018; Tandoc et al., 2018). Automated detection approaches face a moving-target problem: classifiers trained on one model generation degrade rapidly on the next (Chen and Shu, 2023; Su et al., 2024). A complementary line of research therefore focuses on the human side, investigating cognitive inoculation (Kozyreva et al., 2023; Roozenbeek et al., 2022) and digital literacy interventions (Lewandowsky and Van Der Linden, 2021). Designing such interventions, however, requires empirical data on how and where human judgment fails. Building on an earlier survey of the research landscape (Loth et al., 2024), we developed JudgeGPT (Loth et al., 2026b), a web-based study platform with two methodological innovations. First, it decouples authenticity (legitimate vs. fake) from source (human vs. machine) via independent continuous scales, avoiding the conflation inherent in binary real/fake tasks. Second, stimuli are generated by RogueGPT (Loth et al., 2026c, b), a controlled multi-model framework that enables systematic comparison across LLM families. This poster reports five empirical findings from 2,318 judgments and discusses their implications for web system design. 2. Related Work Research on AI-generated misinformation spans two complementary directions. On the detection side, neural classifiers such as Grover (Zellers et al., 2019) and more recent transformer-based detectors (Chen and Shu, 2023) initially achieved high accuracy but degrade as models evolve (Ferrara, 2024); comprehensive surveys confirm this trend (Zhou and Zafarani, 2020). On the human side, controlled experiments consistently show that participants perform near chance when distinguishing machine-generated from human-written text (Jakesch et al., 2023; Clark et al., 2021). Kreps et al. demonstrated that GPT-generated political messages are rated as credible as journalist-written ones (Kreps et al., 2022), while Su et al. showed that detection difficulty increases with model scale (Su et al., 2024). Most human-subject studies, however, employ binary classification tasks (real vs. fake) that conflate source and veracity: a machine-generated paraphrase of a true story is not misinformation, yet binary framing treats it as such. Our dual-axis method addresses this limitation. We additionally build on work in psychological inoculation (Roozenbeek et al., 2022; Kozyreva et al., 2023) and analytic thinking as a predictor of misinformation resilience (Pennycook and Rand, 2021), while connecting to emerging provenance standards (C2PA) as system-level alternatives (Loth et al., 2026e). 3. Methodology JudgeGPT is a web-based study platform for collecting human perception data on AI-generated text (Loth et al., 2026b). Its design was informed by a systematic gap analysis of prior work (Loth et al., 2024). Dual-Axis Assessment. For each news fragment, participants provide three independent ratings on continuous 0â100 sliders (Figure 1): (1) Source judgment (0 = certainly machine, 100 = certainly human), (2) Authenticity judgment (0 = certainly fake, 100 = certainly legitimate), and (3) Topic familiarity. Continuous scales were chosen over Likert items to capture degree of certainty and to enable parametric analyses including Pearson correlation and clustering. Figure 1. The JudgeGPT interface. Participants rate each fragment on three continuous axes: source attribution, authenticity, and topic familiarity. Stimulus Design. Stimuli follow a 2Ă22Ă 2 factorial design crossing origin (human, machine) with veracity (real, fake). Machine-generated fragments are produced by RogueGPT (Loth et al., 2026c, b), which orchestrates six LLMs (GPT-4 (Bubeck et al., 2023), GPT-3.5, GPT-4o, LLaMA-2 13B, Gemma 7B, Mistral 7B) with persona-based prompting. Human-sourced fragments are drawn from established news outlets and known misinformation databases. AI-generated texts are grounded in real news topics and verified for factual consistency to avoid conflating perception of style with detection of hallucinations. The stimulus set is intentionally skewed toward machine-origin fragments (âź 98%), with human-origin items serving as calibration anchors. This design choice reflects the studyâs focus on within-AI variation (across models) rather than human-vs.-AI comparison; participants are not informed of the base rate, and the near-chance detection results (Finding 1) hold when analyzed on the human-origin subset alone. Data Collection. Participants provide informed consent and complete a demographic questionnaire (age, education, political orientation, AI familiarity) before evaluating a sequence of fragments. Each participant evaluates between 5 and 87 fragments (median = 12, IQR = 6â22), with presentation order randomized and model assignment balanced across participants. The platform records the three slider ratings, response latency, and an anonymous participant identifier for linking judgments to covariates. The sample skews toward educated European demographics (68% university-educated, 74% European); we discuss generalizability limitations in Section 5. 4. Findings The dataset comprises 2,318 judgments from 1,054 unique participants. Table 1 provides an overview; each finding is detailed below. Table 1. Summary of empirical findings (N=2,318N=2,318 judgments from 1,0541,054 participants; median 12 items per participant, range 5â87). Finding Statistic Implication F1: Undetectable tâ(2316)=0.87t(2316)=0.87, p=.38p=.38 Intuition fails as defense F2: Model-agnostic All 6 LLMs at xÂŻâ0.50 xâ 0.50 Low barrier to deceptive text F3: Expertise >> Politics r=.35r=.35 vs. r=â.10r=-.10 Literacy-based interventions F4: User personas k=2k=2 clusters (sil. =.41=.41) Adaptive designs needed F5: Fatigue Decline after âź 30 items Sustained vigilance infeasible Finding 1: Machine-Generated Text Is Not Detectable by Participants. Figure 2 shows near-complete overlap between human- and machine-generated content on both axes. A Welchâs t-test yields no significant difference in source scores between conditions (tâ(2316)=0.87t(2316)=0.87, p=.38p=.38, Cohenâs d=0.04d=0.04), confirming that participantsâ intuitions do not reliably discriminate AI-generated from human-written text. Figure 2. Source score (left) and authenticity score (right) distributions for machine- vs. human-written fragments. Overlapping distributions and a non-significant t-test indicate participants cannot distinguish the two conditions. Finding 2: Detection Failure Is Model-Agnostic. Figure 3 disaggregates source scores by generating model. Mean scores for all six LLMs fall within the 0.44â0.55 range, clustering around the 0.5 chance level. A one-way ANOVA reveals no significant between-model effect (Fâ(5,2349)=1.92F(5,2349)=1.92, p=.09p=.09). Open-weight models with as few as 7B parameters produce text rated no differently from GPT-4o output, indicating that the capability to generate human-indistinguishable text is no longer restricted to frontier models. Figure 3. Mean source and authenticity scores per LLM (Âą SE). The dashed line marks chance level (0.5). No model is reliably identified as machine-generated. Finding 3: Domain Expertise Predicts Accuracy; Political Orientation Does Not. Pearson correlations (Figure 4) show that self-reported fake news familiarity is positively associated with source judgment accuracy (r=.35r=.35, p<.001p<.001) and authenticity accuracy (r=.29r=.29, p<.001p<.001). Political extremity shows only a weak, non-significant negative correlation (r=â.10r=-.10, p=.12p=.12). While the political orientation measure is coarse (a single self-report item) and the stimulus set was not designed to control for politically sensitive content, the pattern suggests that learned analytical skills may be a stronger predictor of detection performance than ideological predisposition. A dedicated study with politically balanced stimuli would be needed to draw causal conclusions about the role of political orientation. Figure 4. Relationships between participant covariates and judgment scores. Top: political view (weak slope). Bottom: fake news familiarity (positive slope). Regression lines with 95% CI shown. Finding 4: Clustering Reveals Distinct Response Strategies. K-means clustering (k=2k=2, selected via silhouette analysis, silhouette coefficient =.41=.41) on participant-level mean scores identifies two groups (Figure 5): âSkeptics,â who assign low trust across all content regardless of origin, and âBelievers,â who maintain high baseline trust. These distinct response strategies imply that uniform interventions will be suboptimal. Figure 5. Pairplot of participant-level mean scores colored by cluster. The two groups show clearly separated distributions on both axes. Finding 5: Initial Learning Effect Followed by Cognitive Fatigue. A rolling-window analysis of sequential judgments (Figure 6) reveals improved accuracy during the first 15â20 evaluations, consistent with a learning effect. Beyond approximately 30 items, accuracy declines and participants increasingly default to âfakeâ classifications, indicating cognitive fatigue. This temporal pattern limits the practical duration of detection-based interventions. Figure 6. Rolling mean of judgment scores across sequential evaluations. Scores improve initially before declining after âź 30 items, consistent with a learning-then-fatigue pattern. 5. Discussion and Implications Three design implications follow from our findings: ⢠Provenance over detection. The failure of human judgment across all tested models (F1, F2) implies that detection cannot be offloaded to users. Cryptographic provenance frameworks such as C2PA, which attach verifiable content histories at the point of creation, offer a more scalable alternative (Loth et al., 2026e, f), complemented by source-level credibility signals (Loth et al., 2026a). ⢠Bounded inoculation. The learning effect (F5, first 15â20 items) supports the efficacy of inoculation-based interventions (Roozenbeek et al., 2022; Kozyreva et al., 2023), but the subsequent fatigue effect constrains their practical deployment to short, focused sessions rather than continuous monitoring. ⢠Persona-aware design. The Skeptic/Believer distinction (F4) suggests that web platforms should adapt trust indicators to user disposition rather than applying uniform labels. The expertise effect (F3) further implies that interventions building analytical skill may be more productive than those targeting partisan bias, though this finding should be validated with politically balanced stimuli in future work. These findings are part of a broader empirical program on AI-driven disinformation encompassing generation infrastructure (Loth et al., 2026c), human perception (Loth et al., 2026b), expert assessment (Loth et al., 2026d), provenance-based countermeasures (Loth et al., 2026e, f), domain credibility infrastructure (Loth et al., 2026a), and a unified dissertation framework formalizing the âIndistinguishability Thresholdâ (Loth, 2026). Limitations. The participant sample skews toward educated European demographics (68% university-educated, 74% European), limiting generalizability to broader populations. The controlled evaluation format does not capture the social and contextual cues present in real-world social media encounters. The stimulus set is intentionally skewed toward machine-generated content (âź 98%); while this design serves the studyâs focus on within-AI variation across models, a balanced design would strengthen claims about human-origin anchoring effects and control for potential base-rate adaptation by participants. The political orientation finding (F3) is preliminary: the single-item measure and lack of politically controlled stimuli preclude causal claims. Future work will address these limitations through a longitudinal browser-extension study with a more diverse participant pool and balanced stimulus design. 6. Conclusion Can humans tell? Our dual-axis study of 2,318 judgments across six LLM families provides a clear empirical answer: they cannot. Machine-generated text is indistinguishable from human writing regardless of model size or family, domain expertise predicts detection accuracy more strongly than political orientation, participants adopt distinct trust strategies, and cognitive fatigue limits sustained detection. These findings support a shift from user-level detection toward system-level countermeasures, including content provenance, adaptive trust indicators, and bounded inoculation interventions. Data Availability. All datasets are archived on Zenodo under restricted academic access: RogueGPT stimulus corpus (DOI: 10.5281/zenodo.18703138), JudgeGPT human perception data (DOI: 10.5281/zenodo.18703385), and the CRED-1 domain credibility dataset (DOI: 10.5281/zenodo.19355308). The JudgeGPT platform111https://github.com/aloth/JudgeGPT and RogueGPT framework222https://github.com/aloth/RogueGPT are open-source. Acknowledgements.We thank all participants of the JudgeGPT study. References (1) Bubeck et al. (2023) SĂŠbastien Bubeck, Varun Chandrasekaran, Ronen Eldan, Johannes Gehrke, Eric Horvitz, Ece Kamar, Peter Lee, Yin Tat Lee, Yuanzhi Li, Scott Lundberg, Harsha Nori, Hamid Palangi, Marco Tulio Ribeiro, and Yi Zhang. 2023. Sparks of Artificial General Intelligence: Early experiments with GPT-4. arXiv:2303.12712 [cs.CL] Chen and Shu (2023) Canyu Chen and Kai Shu. 2023. Combating Misinformation in the Age of LLMs: Opportunities and Challenges. arXiv:2311.05656 [cs.CY] Clark et al. (2021) Elizabeth Clark, Tal August, Sofia Serber, Nithum Haber, Asli Celikyilmaz, and Noah A. Smith. 2021. All Thatâs âHumanâ Is Not Gold: Evaluating Human Evaluation of Generated Text. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics. ACL, 7282â7296. Ferrara (2024) Emilio Ferrara. 2024. GenAI against humanity: nefarious applications of generative artificial intelligence and large language models. Journal of Computational Social Science (22 Feb 2024). doi:10.1007/s42001-024-00250-1 Jakesch et al. (2023) Maurice Jakesch, Jeffrey T. Hancock, and Mor Naaman. 2023. Human Heuristics for AI-Generated Language Are Flawed. Proceedings of the National Academy of Sciences 120, 11 (2023), e2208839120. doi:10.1073/pnas.2208839120 Kozyreva et al. (2023) Anastasia Kozyreva, Sam Wineburg, Stephan Lewandowsky, and Ralph Hertwig. 2023. Critical Ignoring as a Core Competence for Digital Citizens. Current Directions in Psychological Science 32, 1 (2023), 39â43. doi:10.1177/09637214221121570 Kreps et al. (2022) Sarah Kreps, R. Miles McCain, and Miles Brundage. 2022. All the News Thatâs Fit to Fabricate: AI-Generated Text as a Tool of Media Misinformation. Journal of Experimental Political Science 9, 1 (2022), 104â117. Lazer et al. (2018) David MJ Lazer, Matthew A Baum, Yochai Benkler, Adam J Berinsky, Kelly M Greenhill, Filippo Menczer, Miriam J Metzger, Brendan Nyhan, Gordon Pennycook, David Rothschild, et al. 2018. The science of fake news. Science 359, 6380 (2018), 1094â1096. Lewandowsky and Van Der Linden (2021) Stephan Lewandowsky and Sander Van Der Linden. 2021. Countering misinformation and fake news through inoculation and prebunking. European Review of Social Psychology 32, 2 (2021), 348â384. Loth (2026) Alexander Loth. 2026. The Indistinguishability Threshold: Measuring Cognitive Vulnerabilities to AI-Generated Disinformation. In 18th ACM Web Science Conference (WebSci Companion â26), May 26â29, 2026, Braunschweig, Germany, PhD Symposium (Braunschweig, Germany). ACM, New York, NY, USA. doi:10.1145/3795513.3807421 Loth et al. (2024) Alexander Loth, Martin Kappes, and Marc-Oliver Pahl. 2024. Blessing or curse? A survey on the Impact of Generative AI on Fake News. arXiv:2404.03021 [cs.CL] Loth et al. (2026a) Alexander Loth, Martin Kappes, and Marc-Oliver Pahl. 2026a. CRED-1: An Open Multi-Signal Domain Credibility Dataset for Automated Pre-Bunking of Online Misinformation. (2026). doi:10.2139/ssrn.6448466 Preprint available at SSRN. Loth et al. (2026b) Alexander Loth, Martin Kappes, and Marc-Oliver Pahl. 2026b. Eroding the Truth-Default: A Causal Analysis of Human Susceptibility to Foundation Model Hallucinations and Disinformation in the Wild. In Companion Proceedings of the ACM Web Conference 2026 (W â26 Companion) (Dubai, United Arab Emirates). ACM, New York, NY, USA. doi:10.1145/3774905.3795832 To appear. Loth et al. (2026c) Alexander Loth, Martin Kappes, and Marc-Oliver Pahl. 2026c. Industrialized Deception: The Collateral Effects of LLM-Generated Misinformation on Digital Ecosystems. In Companion Proceedings of the ACM Web Conference 2026 (W â26 Companion) (Dubai, United Arab Emirates). ACM, New York, NY, USA. doi:10.1145/3774905.3795471 To appear. Loth et al. (2026d) Alexander Loth, Martin Kappes, and Marc-Oliver Pahl. 2026d. The Verification Crisis: Expert Perceptions of GenAI Disinformation and the Case for Reproducible Provenance. In Companion Proceedings of the ACM Web Conference 2026 (W â26 Companion) (Dubai, United Arab Emirates). ACM, New York, NY, USA. doi:10.1145/3774905.3795484 To appear. Loth et al. (2026e) Alexander Loth, Dominique Conceicao Rosario, Peter Ebinger, Martin Kappes, and Marc-Oliver Pahl. 2026e. Origin Lens: A Privacy-First Mobile Framework for Cryptographic Image Provenance and AI Detection. In Companion Proceedings of the ACM Web Conference 2026 (W â26 Companion) (Dubai, United Arab Emirates). ACM, New York, NY, USA. To appear. Loth et al. (2026f) Alexander Loth, Dominique Conceicao Rosario, Peter Ebinger, Martin Kappes, and Marc-Oliver Pahl. 2026f. Origin Lens: Reclaiming Trust on the AI-Mediated Web Through On-Device Image Provenance Verification. In 18th ACM Web Science Conference (WebSci Companion â26), May 26â29, 2026, Braunschweig, Germany (Braunschweig, Germany). ACM, New York, NY, USA. doi:10.1145/3795513.3806658 Pennycook and Rand (2021) Gordon Pennycook and David G Rand. 2021. The psychology of fake news. Trends in cognitive sciences 25, 5 (2021), 388â402. Roozenbeek et al. (2022) Jon Roozenbeek, Sander van der Linden, Beth Goldberg, Steve Rathje, and Stephan Lewandowsky. 2022. Psychological Inoculation Improves Resilience Against Misinformation on Social Media. Science Advances 8, 34. Su et al. (2024) Jinyan Su, Claire Cardie, and Preslav Nakov. 2024. Adapting Fake News Detection to the Era of Large Language Models. arXiv:2311.04917 [cs.AI] Tandoc et al. (2018) Edson C Tandoc, Zheng Wei Lim, and Richard Ling. 2018. Defining âfake newsâ: A typology of scholarly definitions. Digital journalism 6, 2 (2018), 137â153. Vosoughi et al. (2018) Soroush Vosoughi, Deb Roy, and Sinan Aral. 2018. The Spread of True and False News Online. Science 359, 6380 (2018), 1146â1151. Wardle and Derakhshan (2017) Claire Wardle and Hossein Derakhshan. 2017. Information Disorder: Toward an Interdisciplinary Framework for Research and Policy Making. Council of Europe Report (2017). Zellers et al. (2019) Rowan Zellers, Ari Holtzman, Hannah Rashkin, Yonatan Bisk, Ali Farhadi, Franziska Roesner, and Yejin Choi. 2019. Defending Against Neural Fake News. In Advances in Neural Information Processing Systems, Vol. 32. Zhou and Zafarani (2020) Xinyi Zhou and Reza Zafarani. 2020. A Survey of Fake News: Fundamental Theories, Detection Methods, and Opportunities. ACM Comput. Surv. 53, 5, Article 109 (sep 2020), 40 pages. doi:10.1145/3395046