Source-linked AI summary

Combating Disinformation in a Social Media Age

Kai Shu, Amrita Bhattacharjee, Faisal Alatawi, Tahora Nazer, Kaize Ding, Mansooreh Karami, Huan Liu

arXiv:2007.07388v1cs.SIcs.IR

TL;DR

Disinformation on social media is difficult to detect and consequential because it spreads rapidly through accessible platforms and can be reinforced by users’ beliefs and networks. This paper surveys its forms, spread factors, detection and mitigation approaches, educational efforts, and future directions, including threat modeling and psychological research. It concludes that early detection remains unresolved and that domain dependence and scarce data limit models across contexts.

  • Problem

    Social media enables widespread disinformation, while belief formation, bot activity, echo chambers, and limited early evidence make detection and mitigation challenging.

  • Method

    The paper provides a comprehensive overview of disinformation forms, history, spread factors, detection techniques, mitigation approaches, education, and future research.

  • Results

    The review identifies threat modeling, user reactions, echo chambers, and transfer learning as promising directions for improving understanding and detection.

  • Takeaways & Limitations

    Future work should investigate disinformation generators, psychological factors, echo chambers, and domain-adaptive methods for early detection.

  • Takeaways & Limitations

    Early detection remains unsolved, and topic dependence plus scarce data can cause models trained in one domain or event set to perform poorly elsewhere.

Abstract

from arXiv · show

The creation, dissemination, and consumption of disinformation and fabricated content on social media is a growing concern, especially with the ease of access to such sources, and the lack of awareness of the existence of such false information. In this paper, we present an overview of the techniques explored to date for the combating of disinformation with various forms. We introduce different forms of disinformation, discuss factors related to the spread of disinformation, elaborate on the inherent challenges in detecting disinformation, and show some approaches to mitigating disinformation via education, research, and collaboration. Looking ahead, we present some promising future research directions on disinformation.

1 Introduction

Social media makes information fast and widely accessible, but this reach also enables disinformation to influence opinions and actions. The paper surveys disinformation terminology, history, spread, detection challenges, mitigation, education, and future research.

  • Motivation: Social media enables rapid information diffusion to millions, especially during time-sensitive crises, while making news veracity important to verify.A considerable fraction of people use social media for news, allowing information to affect their opinions and actions.
  • Research context: Disinformation research examines the problem from multiple perspectives, including detection and related fields.
  • Paper scope: The paper reviews disinformation definitions, history, characteristics, spread factors, detection challenges, mitigation techniques, education, and future research.
  • Definitions: Disinformation is false information deliberately spread to mislead, while fake news is intentionally and verifiably false content that can mislead readers.The paper uses “disinformation” and “fake news” interchangeably.
  • Definitions: Hoaxes are deliberate false messages intended to persuade or manipulate many people through threats or deception.
  • Spread and impact: Social media and internet connectivity facilitate rapid propagation of rumors, hoaxes, propaganda, and other false claims that can shape public sentiment and actions.The paper situates contemporary disinformation within a longer history of political deception and fabricated information.
  • Detection challenges: Detection is difficult because sensationalized content attracts interaction, automated sources are inexpensive to create, and early detectors lack later user reactions or fact-checking knowledge.
  • Detection challenges: Filter bubbles and echo chambers reinforce users’ existing beliefs, making later correction and disinformation mitigation especially difficult.

2 Disinformation in Different Forms

Disinformation appears across text, images, videos, and multimodal posts, and can be generated by humans, machines, or both. The reviewed detection approaches therefore exploit artifacts, temporal inconsistencies, event-invariant features, and combined content and user-response signals.

  • 2 Disinformation in Different Forms: False content spans text, images, videos, and multimodal combinations, while human and machine generation complicate detection assumptions.
  • 2.1 GAN Generated Fake Images: GAN-generated and doctored images can alter semantic meaning and become difficult for humans to distinguish from authentic images.
  • 2.1 GAN Generated Fake Images: Image detectors have been evaluated on CycleGAN-generated and compressed images, addressing distortions introduced by social-media processing.
  • 2.1 GAN Generated Fake Images: Assuming one known GAN limits real-world applicability because attackers’ generation models are usually unavailable.
  • 2.1 GAN Generated Fake Images: Other image detectors use GAN checkerboard artifacts or color-channel co-occurrence matrices as discriminative features.
  • 2.2 Fake Videos and Deepfakes: Deepfake detection exploits across-frame inconsistencies caused when face-swapping models process video frames independently.
  • 2.2 Fake Videos and Deepfakes: Additional video approaches examine abnormal eye blinking and discrepancies in 3D head pose and facial landmarks.
  • 2.3 Multimodal Content: Multimodal fake posts combine images, text, and comments to appear more believable, motivating models that learn event-invariant features and combine modalities.

3 Factors behind the Spread of Disinformation

Disinformation spreads through uncertainty, emotion, repetition, ideological alignment, and low-cost amplification by bots and alternative media. These factors exploit users’ difficulty recognizing falsehood and the speed and reach of social platforms.

  • Emotional and cognitive factors: Users are less able to spot falsehoods when emotionally affected or encountering claims consistent with their values, while repeated exposure increases belief and sharing.Malicious actors exploit these vulnerabilities by presenting the same disinformation repeatedly and through multiple sources.
  • Media and platform conditions: Low-cost online publishing and broad platform reach have enabled alternative media sources to spread false or highly biased claims.These sources may resemble mainstream media while lacking the same journalistic integrity.
  • Emotional and situational factors: Uncertainty, anxiety, personal importance, and lack of control make users more prone to propagating unproved claims and rumors.During crises, people may form unofficial information networks when official information is unavailable.
  • Bots and mitigation: Removing a small percentage of malicious bots can virtually eliminate the spread of low-credibility content.The paper presents bot reduction as a promising way to slow disinformation propagation, while noting bots are not the only source of spread.
  • Bots and amplification: Social bots mimic humans, amplify newly created content, target influential users, and increase the perceived credibility and exposure of fabricated articles.Bots can also use political messaging, disaster hashtags, misdirection, and smoke screening to shape public attention.
  • Bots and amplification: In 2017, between 9% and 15% of Twitter users exhibited bot behaviors, while Twitter suspended 70 million suspicious accounts in 2018.These figures illustrate the scale of bot activity reported across the platform.
  • Bots and political disinformation: During the 2016 US presidential election, bots produced 1.7 billion tweets and outnumbered tweets supporting one candidate by a factor of four.The paper reports a comparable 4-to-1 imbalance in Facebook stories supporting the same candidate.

4 Detecting Disinformation

The paper organizes disinformation detection around targeted users, content, and network propagation, while reviewing machine-learning, crowd-based, fact-checking, and structured-data approaches. It emphasizes that early detection and changing events remain difficult because useful reactions, knowledge, and transferable features may be unavailable.

  • Detection dimensions: Disinformation detection can use features from targeted users, the content itself, and the way information spreads through networks.The paper argues that effective detectors should combine these areas rather than focus on only one.
  • User-based detection: User profiles, sharing behavior, crowd flags, sentiment, and conversational stance provide signals for identifying disinformation.Responses can be classified into support, deny, query, or comment, while emotional responses may distinguish disinformation from real news.
  • Crowd and expert signals: Detective uses Bayesian inference to identify fake news while learning users’ flagging behavior and referring selected articles to experts for verification.The method aims to reduce the effect of adversarial users over time.
  • Mitigation through fact-checking: Personalized recommendations of fact-checking URLs are proposed to encourage users to engage with debunking content.The recommendations are based on users’ interests.
  • Early detection: Early detection is difficult because detectors may lack user responses and fact-checking knowledge bases during the initial dissemination of a claim.Existing responses to previously propagated articles may provide latent user information for improving detection.
  • Content and event generalization: Event Adversarial Neural Network derives event-invariant features to address poor transfer from previously observed events to newly emerged ones.Conventional strategies often learn event-specific features that do not transfer well to unseen events.
  • AI-generated disinformation: Grover generates news articles from headlines and identifies content generated by itself with around 90% accuracy, framing AI-generated news as a threat-modeling problem.The paper also notes that adversaries could use similar technologies to generate articles or comments.
  • Content representation: Detection trained on neural-generated summaries performed best on the reported dataset, suggesting summarization can act as a feature generator.The comparison used entire news text, headlines, and abstractive summaries as inputs.

5 Mitigating Disinformation via Education, Research, and Collaboration

The paper surveys how education, interdisciplinary research, and computational interventions can mitigate disinformation. These approaches target public media literacy, cognitive and social factors, malicious sources, network diffusion, and content verification.

  • Education: Over 3,400 U.S. high school students were evaluated on six social-media news-literacy tasks, and the majority failed.The assessment tested distinguishing fact from fiction and judging source credibility.
  • Education: Schools and governments are expanding media-literacy education to teach citizens how to evaluate online information before sharing it.Examples include curricula across all 50 U.S. states and Finnish instruction on assessing article authenticity.
  • Research and collaboration: Disinformation research spans cognitive science, journalism, political science, and computational mitigation.Researchers study belief formation, media reporting and elite discourse, while platforms develop computational interventions.
  • Computational mitigation: Computational mitigation includes identifying malicious sources, limiting network diffusion, and flagging posts for fact-checking and possible removal.Source identification uses provenance paths and influential-user analysis; network methods include influence minimization and counter-campaign cascades.
  • Cognitive science: Cognitive studies link belief in disinformation to limited open-minded and analytic thinking more than partisan bias.People may revise their assessments when shown the truth, while cognitive ability influences the degree of change.
  • Journalism and political science: Disinformation exposure is concentrated among a small group of users who tend to be conservative-leaning, older, and highly engaged with political news.Confirmation bias and social influence are associated with echo chambers where users share disinformation.
  • Journalism and political science: Experts recommend consuming news from verified accounts and avoiding rhetoric that attacks legitimate news organizations as fake.They also caution that elite discussions of disinformation can reduce trust in traditional media and hinder identification of real information.

6 Conclusion and Future Work

The paper consolidates research on disinformation forms, detection, mitigation, datasets, and social-science perspectives, then identifies future work across threat modeling, deepfakes, user behavior, and transfer learning. It emphasizes unresolved early detection and the need for models and data that generalize across forms and domains.

  • Conclusion: The paper provides a comprehensive overview of disinformation definitions, forms, history, detection techniques, mitigation efforts, datasets, and tools.Its stated goal is to characterize the current research landscape surrounding disinformation detection.
  • Future directions: Threat modeling could improve disinformation detectors, but existing work is limited and does not cover the full range of text, image, and video forms.The paper identifies only one known threat-modeling study and notes that it is limited to textual data.
  • Future directions: More video threat modeling and annotated deepfake datasets are needed because fabricated videos are easy to generate and publicly labeled datasets are scarce.The proposed use of stock video databases is described as inexpensive and quick to overwhelm detectors.
  • Future directions: User reactions to novel news may support detection because people respond differently to real and false information, while echo chambers may enable earlier identification.The paper connects susceptibility to disinformation, sharing among like-minded users, and possible early-detection signals.
  • Future directions: Early fake-news detection remains unsolved, and topic-dependent models may fail in new contexts; transfer learning is proposed to adapt models to newly available data.The proposed adaptation uses models that have learned linguistic differences between fake and real news.

A Datasets and Tools Available

The paper notes that ongoing research has produced a considerable number of datasets for disinformation detection. These resources support researchers developing computational detection methods.

  • Dataset landscape: A considerable number of datasets are available for researchers studying fake-news and disinformation detection.The paper presents these datasets as resources accompanying the growing volume of research in the field.

A.1 Datasets

The reviewed datasets cover multimedia tweets, rumors, credibility, political statements, death hoaxes, war-related articles, and social context. They vary substantially in scale, labels, domains, and included metadata.

  • Multimedia and rumor datasets: Twitter Media Corpus contains multimedia tweets labeled fake or real according to the authenticity of supporting images.It is designed for studying fake multimedia content on Twitter.
  • Multimedia and rumor datasets: PHEME contains tweets from nine events labeled True, False, or Unverified, covering both rumors and non-rumors.Its labels represent event-related rumor status.
  • Credibility and truthfulness datasets: CREDBANK contains over 60 million tweets from 1,049 real-world events with human-annotated credibility scores.Scores range from Certainly Accurate to Certainly Inaccurate across five credibility categories.
  • Credibility and truthfulness datasets: LIAR contains 12.8K labeled statements from PolitiFact and Facebook, using six fine-grained truthfulness labels and speaker credit histories.The statements span contexts and speakers with different political affiliations.
  • Hoax and domain-specific datasets: Twitter Death Hoax contains over 4,000 death reports posted from 2012 through 2014, including 2,031 verified real deaths.Wikipedia was used to verify the reported deaths.
  • Hoax and domain-specific datasets: FA-KES contains Syrian War news articles labeled fake or credible against ground truth from the Syrian Violations Documentation Center.The dataset is domain-specific and focuses on war-related reporting.
  • Social-context datasets: FakeNewsNet combines political and celebrity news with social-context data such as user profiles, following, and follower relationships.Its sources include PolitiFact and GossipCop.

A.2 Tools

Researchers and practitioners have made publicly available tools to support disinformation-detection research. These include tools for generating fake news, verifying social-media content, and collecting and classifying Facebook posts.

  • Publicly available tools support ongoing research on fake-news and disinformation detection.The section presents selected tools made available by researchers and practitioners.
  • Big Bird automatically generates fake news articles that can subsequently be edited by human agents.Generated articles are published on the notrealnews.net website.
  • Computational-Verification identifies fake versus real Twitter content using tweet-based and user-based features in a two-level classification model.Features include word and retweet counts, friend-follower ratio, and tweet counts.
  • Facebook Hoax collects Facebook page posts from a user-defined date and supports hoax classification using user likes.A dataset of over 15,000 posts was classified with logistic regression and harmonic boolean label crowdsourcing.
Loading 2007.07388v1…