Source-linked AI summary
Human heuristics for AI-generated language are flawed
Maurice Jakesch, Jeffrey Hancock, Mor Naaman
TL;DR
People increasingly encounter AI-generated language that can be mistaken for human writing, raising concerns about deception and manipulation. Across six experiments, this paper tests detection of AI-generated self-presentations and finds that flawed human heuristics leave judgments near chance and can be exploited to produce text perceived as more human than human.
Problem
AI-generated language is increasingly mixed into human communication, but evidence is limited on whether people can detect it in consequential verbal self-presentations.
Method
Six experiments tested judgments of AI-generated self-presentations across professional, romantic, and hospitality contexts using qualitative, quantitative, and computational analyses of detection heuristics.
Results
Participants remained near chance at detecting AI-generated self-presentations, while analyses showed that flawed cues hindered judgment and enabled optimized AI text to appear more human than human-written text.
Takeaways & Limitations
Human intuition is vulnerable to predictable manipulation because AI systems can exploit flawed heuristics to generate self-presentations perceived as more human than human.
Takeaways & Limitations
The findings are limited to the current generation of language models and people’s current heuristics, which technology and culture may change.
Abstract
from arXiv · showhide
Human communication is increasingly intermixed with language generated by AI. Across chat, email, and social media, AI systems suggest words, complete sentences, or produce entire conversations. AI-generated language is often not identified as such but presented as language written by humans, raising concerns about novel forms of deception and manipulation. Here, we study how humans discern whether verbal self-presentations, one of the most personal and consequential forms of language, were generated by AI. In six experiments, participants (N = 4,600) were unable to detect self-presentations generated by state-of-the-art AI language models in professional, hospitality, and dating contexts. A computational analysis of language features shows that human judgments of AI-generated language are hindered by intuitive but flawed heuristics such as associating first-person pronouns, use of contractions, or family topics with human-written language. We experimentally demonstrate that these heuristics make human judgment of AI-generated language predictable and manipulable, allowing AI systems to produce text perceived as "more human than human." We discuss solutions, such as AI accents, to reduce the deceptive potential of language generated by AI, limiting the subversion of human intuition.
Introduction
AI-generated language is increasingly embedded in human communication, creating risks when people mistake it for human writing. This study examines whether people can detect AI-generated verbal self-presentations and why their judgments fail.
- AI systems can generate coherent writing and entire conversations at scale, reducing effort while enabling plagiarism, manipulation, and deception when mistaken for human language.
- Verbal self-presentation is consequential because self-descriptions help shape impressions and establish trust in professional, hospitality, and dating interactions.
- AI-generated self-presentations may undermine cues such as tone and compositional skill that people use to assess others.
- Using qualitative, quantitative, and computational methods, the study reconstructs detection heuristics, tests whether they help or hinder judgment, and examines whether AI can manipulate perceived humanity.
Results
Across six experiments, people’s judgments of AI-generated self-presentations remained near chance despite incentives and feedback. Shared heuristics partly explained this failure, and AI-generated text optimized around them was judged more human than human-written text.
- Main experiments: 52.2% of the time, participants correctly identified the source of hospitality self-presentations, while accuracy remained near chance across tested contexts.
- Main experiments: 51.6% accuracy with monetary incentives and 51.2% after feedback showed that added effort and training did not materially improve judgments.
- Shared heuristics: Fleiss’ kappa = 0.07 indicated agreement significantly above chance despite near-chance accuracy, consistent with shared but flawed heuristics.
- Shared heuristics: 40% of explanations cited family or life-experience content, 28% grammatical cues, and 24% tone when judging whether self-presentations were human-written.
- Feature analysis: 58.8% accuracy was achieved by a classifier using nonsensical or repetitive labels, compared with 51.7% for participants directly judging source.
- Feature analysis: Participants treated grammatical issues, long words, and rare bigrams as AI cues even though these features were more indicative of human-written language in the experiments.
- Validation studies: 65.7% of optimized AI self-presentations were rated human, versus 51.6% of regular generated and 51.7% of human-written presentations.
Discussion
Humans remained near chance at detecting AI-generated self-presentations, partly because misleading intuitive heuristics offset useful cues. These heuristics also allow AI systems to manipulate judgments and produce text perceived as more human than human-written language.
- Discussion: Across contexts, demographics, effort, and expertise, human discernment of AI-generated self-presentations remained close to chance.The authors attribute failure either to limited reliable cues or reliance on flawed heuristics.
- Discussion: AI-generated self-presentations contained detectable nonsensical and repetitive features, yet direct source judgments still remained near chance.Participants identified these features more often in a separate labeling task than in human-written text.
- Discussion: Useful cues could have yielded 58.8% accuracy, but misleading cues about grammar, rare bigrams, long words, family topics, and first-person pronouns reduced accuracy to chance.Some cues associated with human-written language were equally present in AI-generated and human-written self-presentations.
- Discussion: Human-like AI text does not necessarily indicate greater machine intelligence because systems can increase perceived humanity by emphasizing family topics.The authors interpret human detection failure as vulnerability to dysfunctional heuristics rather than evidence of intelligence.
- Discussion: Validation experiments showed that AI systems can exploit flawed heuristics to generate self-presentations perceived as more human than human-written language.The authors connect this vulnerability to risks involving disclosure of private information and adherence to recommendations from entities perceived as human.
- Discussion: Improving detection through education and technical tools may be limited, while future model adaptations could invalidate learned heuristics.Transparent, context-appropriate disclosure mechanisms remain an open research problem.
- Discussion: The authors propose self-disclosing AI language and dedicated AI accents that preserve communication flow while supporting intuitive source judgments.Such systems would avoid language wrongly associated with humanity, including informal and colloquial speech.
Materials and Methods
The study combined simplified Turing-test judgments with labeling tasks across hospitality, dating, and professional contexts. Researchers generated self-presentations with language models, modeled heuristic cues, and selected AI text optimized for perceived humanity.
- Experiment design: Six experiments asked participants to judge whether online-profile-style self-presentations were human-written or AI-generated in hospitality, dating, and professional scenarios.Participants rated multiple presentations after comprehension checks.
- Experiment design: Experiments varied context, presentation length, and interventions including incentives, feedback, and training to examine effects on detection accuracy.Dating and professional studies used longer presentations and fewer ratings to keep duration comparable.
- Data collection: Researchers collected real-world profile data and used subsets for evaluation and full datasets to train state-of-the-art language models.Different models were used as more powerful systems became available during the research.
- Data collection: The hospitality study collected 28,890 Airbnb host self-presentations and generated 1,500 AI presentations by fine-tuning a 774M-parameter GPT-2 model.The source texts contained 30–60 words, and nucleus sampling used P = 0.95.
- Data preparation: The researchers found no substantial plagiarism after checking duplicates and training-data sentence overlap.Ninety-five percent of sentences in AI-generated texts were absent from the training data.
- Predicting responses: About 180 computational and manually developed language features were used to model participants’ perceptions and actual AI-generation status.The feature design combined automatically extracted measures with measures motivated by qualitative labeling.
- Predicting responses: Lasso feature selection reduced predictors to 15 before logistic regression on 4,900 self-presentations, alongside a language-model sequence classifier.The selected-feature process used 20% of presentations, while regression used a separate 70% subset.
- Generating optimized language: Validation samples included 100 human-written, 100 regular AI-generated, and 100 AI presentations optimized for perceived humanity.Classifiers selected generated texts predicted to be perceived as human, using different classifiers across contexts.
Data sharing
The analyses and materials are openly available through an Open Science Framework repository.
- Data sharing: Data and code for the three main studies and three validation experiments are publicly available through an Open Science Framework repository.The work also used previously published data.
Figures and Tables
The figures show near-chance human detection and successful optimization of AI text for perceived humanity, while Table 1 separates functional from misleading heuristics. Supplementary materials document recruitment, explanations, robustness checks, and labeling-task differences.
- Figures: Figure 1 reports near-chance detection across three social contexts despite monetary incentives or correctness feedback.Error bars represent 95% CIs for 6,000–16,000 judgments of 2,000–3,000 self-presentations per bar.
- Figures: Figure 2 compares human-written, regular AI-generated, and optimized AI-generated self-presentations, with optimized text rated as more human.Error bars represent 95% CIs for 350–450 judgments of 100 self-presentations per bar.
- Tables: Table 1 models perceived AI-generation and actual AI-generation, identifying nonsense, repetition, and conversational words as functional cues.Other participant heuristics were inversely related or unrelated to the actual source.
- Supplementary analyses: Supplementary analyses found slightly higher accuracy among participants older than 50, with no gender or ethnic group performing better.Older participants achieved 53% accuracy versus 51% for younger participants.
- Supplementary analyses: Supplementary qualitative analyses coded participants’ explanations to summarize recurring self-reported heuristics.Two researchers independently coded a sample of responses.
- Supplementary analyses: Accuracy did not change after participants explained their judgments in open-ended responses across any of the three contexts.This analysis addressed whether explanation might alter subsequent evaluations or performance.
- Labeling task: In a separate labeling task, crowdworkers rated AI-generated text as nonsensical more often than human-written text: 13.6% versus 9.6%.The difference was especially pronounced in hospitality, where the older GPT-2 model was used.
- Labeling task: Crowdworkers rated AI-generated text as more repetitive but as having fewer grammatical issues than human-written text.Repetition was 12.7% versus 7.1%, while grammatical issues were 14.8% versus 19.6%.
SI References
The supplementary materials document the experiments, self-presentation examples, participant-judgment heuristics, language features, and analyses of detection accuracy. They also report that no participant group performed much above chance and describe auxiliary tests of effort, feedback, and demographics.
- Experiment overview: Table S1 catalogs the main and validation experiments across hospitality, dating, and professional contexts, including variation in presentation type, bonus payments, and feedback.The experiments also used collected self-presentations and generated text from state-of-the-art language models.
- Self-presentation examples: Table S2 provides human-written and AI-generated self-presentation examples from hospitality, dating, and professional settings, including optimized generated examples.The examples include both ordinary and optimized AI-generated self-presentations.