Source-linked AI summary
AI model GPT-3 (dis)informs us better than humans
Giovanni Spitale, Nikola Biller-Andorno, Federico Germani
TL;DR
The paper examines whether people can distinguish accurate information from disinformation and human-written tweets from GPT-3-generated tweets. It finds that GPT-3 produces easier-to-understand accurate information, more compelling disinformation, and text humans cannot reliably identify as synthetic.
Problem
The study addresses limited evidence on whether people can distinguish accurate information from disinformation and human-written tweets from GPT-3-generated tweets.
Method
The study compares how recruited participants evaluate organic and GPT-3-generated tweets containing reliable or false information.
Results
GPT-3-generated reliable tweets were recognized as true better and faster, while false synthetic tweets were recognized as false worse than corresponding organic tweets.
Takeaways & Limitations
GPT-3 can improve the understandability of accurate information while increasing the persuasiveness of disinformation, creating risks for information dissemination.
Takeaways & Limitations
The recruitment strategy could not produce a representative sample upfront, so representativeness was assessed through sequential demographic targeting.
Abstract
from arXiv · showhide
Artificial intelligence is changing the way we create and evaluate information, and this is happening during an infodemic, which has been having dramatic effects on global health. In this paper we evaluate whether recruited individuals can distinguish disinformation from accurate information, structured in the form of tweets, and determine whether a tweet is organic or synthetic, i.e., whether it has been written by a Twitter user or by the AI model GPT-3. Our results show that GPT-3 is a double-edge sword, which, in comparison with humans, can produce accurate information that is easier to understand, but can also produce more compelling disinformation. We also show that humans cannot distinguish tweets generated by GPT-3 from tweets written by human users. Starting from our results, we reflect on the dangers of AI for disinformation, and on how we can improve information campaigns to benefit global health.
Discussion
GPT-3 can inform more effectively than organic tweets while producing more compelling disinformation, and people struggle to identify synthetic text. The findings highlight training-data dependence, uncertainty about machine-generated content, and the need for transparency and regulation.
- How to communicate and evaluate information: GPT-3 synthetic tweets with reliable information are recognized as true better and faster, while false synthetic tweets are recognized as false worse than comparable organic tweets.GPT-3 does not outperform humans in recognizing information and disinformation.
- “Disobedience”, training datasets, and error propagation: GPT-3 is less likely to produce misinformation on some topics, including vaccines and autism, depending on the composition of its training datasets.The paper attributes this topic-specific “disobedience” to statistical patterns in the data used to train GPT-3.
- ‘As human as humans’: synthetic text identification and impersonation: Both human respondents and GPT-3 struggle to differentiate organic from synthetic tweets, although training courses based on linguistic markers, grammar, and syntax might improve recognition.The availability of ChatGPT makes synthetic-text identification and impersonation an ongoing concern.
- Resignation theory: Exposure to synthetic and organic texts decreases respondents’ confidence in distinguishing them, likely because GPT-3 mimics human writing styles and language patterns.Participants may also become more sceptical of both synthetic and organic information after recognizing GPT-3’s ability to generate human-like disinformation.
- Beyond Twitter: 5% of Twitter users are bots, yet bots account for 20% - 29% of posted Twitter content, motivating the study’s focus on tweets and its concern with broader information dissemination.Twitter users consume mostly news and political information, and its simple API facilitates unsupervised bots.
- The genie is out of the bottle: Regulating the training datasets used to develop advanced AI text generators is crucial for transparency, truthful outputs, and limiting misuse to generate deceiving information.The paper presents advanced AI text generators as capable of affecting information dissemination positively and negatively.
Secondary endpoints hypotheses
Secondary endpoints showed that respondents recognized synthetic accurate tweets more successfully than organic accurate tweets, while confidence changed differently for recognizing disinformation versus synthetic content.
- Secondary Endpoint 1: 0.78 for synthetic accurate information versus 0.64 for organic accurate information.Scores range from 0 to 1 and indicate performance in recognizing accurate tweets.
- Secondary Endpoint 2: 0.315 for recognizing synthetic tweets regardless of truthfulness versus 0.59 for recognizing organic tweets regardless of truthfulness.Scores range from 0 to 1 and combine accurate information and disinformation.
- Secondary Endpoint 3: Pre-confidence in recognizing disinformation was 2.932271, while post-confidence was 3.319149.Confidence scores range from 1 to 5.
- Secondary Endpoint 4: Pre-confidence in recognizing synthetic versus organic contents was 2.703557, while post-confidence was 1.75.Confidence scores range from 1 to 5.
Supplementary Results
Supplementary analyses found small demographic associations with recognition scores, while score changes in confidence related to disinformation recognition. Survey completion time was unrelated to either recognition score.
- OS score and demographics: Age showed a small association with OS Score, with younger respondents performing slightly better at distinguishing synthetic from human tweets than older respondents.Respondents aged 18–41 performed slightly better than those aged 16–17 and especially those aged 42+.
- TF score and demographics: Age and education level each correlated with TF score with small effect sizes.The TF-score distribution across ages was generally uniform, despite 42–57-year-olds performing slightly better than respondents aged 58–76.
- TF score and demographics: Higher education was consistently associated with higher TF scores, with doctorate/PhD, Master’s, and Bachelor’s participants ordered from higher to lower scores.The passage describes a stepwise pattern across education levels, with doctorate/PhD participants scoring higher than Master’s participants, who scored higher than Bachelor’s participants.
- OS / TF self-confidence delta and OS / TF score: A small but significant correlation was found between TF Delta and TF Score, whereas OS Delta and OS Score were not correlated.TF Delta measured the change in confidence about recognizing disinformation; OS Delta measured the change in confidence about recognizing AI-generated text.
- Duration and OS / TF scores: Survey duration had no significant correlation with either OS Score or TF Score.The analysis separately tested duration against both recognition scores and found no significant association.