Source-linked AI summary
False Information on Web and Social Media: A Survey
Srijan Kumar, Neil Shah
TL;DR
False information spreads easily online and can produce substantial real-world impact, while readers often struggle to recognize deceptive content. This survey unifies research on its forms, spread, impact, characteristics, detection methods, and future directions. It concludes that false information has distinctive measurable patterns, but detection remains constrained by class imbalance, masquerading, labeling difficulty, and limited standardized comparison.
Problem
False information’s rapid online spread and real-world effects make understanding its propagation, deceptive success, and detection an important research problem.
Method
The paper presents a comprehensive survey and unified framework covering actors, deception rationale, impact, characteristics, detection algorithms, and future research directions.
Results
False information shows distinctive textual, temporal, user, network, spreading, and mitigation characteristics, and detection methods comprise feature-based, graph-based, and propagation-modeling approaches.
Takeaways & Limitations
The survey’s synthesis supports using false information’s multifaceted characteristics to develop detection tools and identifies semantic dissonance detection as an open research avenue.
Takeaways & Limitations
Detection research is constrained by severe class imbalance, difficulty obtaining labels, and the absence of standardized comparisons across available datasets.
Abstract
from arXiv · showhide
False information can be created and spread easily through the web and social media platforms, resulting in widespread real-world impact. Characterizing how false information proliferates on social platforms and why it succeeds in deceiving readers are critical to develop efficient detection algorithms and tools for early detection. A recent surge of research in this area has aimed to address the key issues using methods based on feature engineering, graph mining, and information modeling. Majority of the research has primarily focused on two broad categories of false information: opinion-based (e.g., fake reviews), and fact-based (e.g., false news and hoaxes). Therefore, in this work, we present a comprehensive survey spanning diverse aspects of false information, namely (i) the actors involved in spreading false information, (ii) rationale behind successfully deceiving readers, (iii) quantifying the impact of false information, (iv) measuring its characteristics across different dimensions, and finally, (iv) algorithms developed to detect false information. In doing so, we create a unified framework to describe these recent methods and highlight a number of important directions for future research.
1 INTRODUCTION
False information spreads rapidly and cheaply through highly connected web and social platforms, producing substantial real-world effects. This survey organizes research on its forms, actors, impact, characteristics, and detection algorithms.
- Web platforms let users distribute information to millions within minutes at little or no cost, increasing the visibility of both true and false information.False information has affected stock markets, disaster responses, and terrorist attacks, while social media increasingly serves as a news source.
- The survey focuses on fake reviews, hoaxes, and fake news across e-commerce, collaborative, and social media platforms.It distinguishes these as major forms studied in prior research.
- False information is categorized by intent into misinformation, created without intent to mislead, and disinformation, created to deceive readers.The survey also considers whether information is opinion-based or fact-based.
- A small fraction of false information can receive greater engagement than true information, generate deeper reshare cascades, survive longer, and spread across the web.These measures include likes, shares, comments, cascade depth, and pre-removal lifetime.
- False information exhibits distinctive textual, temporal, user, network, spreading, and mitigation characteristics that researchers use to predict veracity.Fake reviews may be longer and more exaggerated, while fake reviews and news can appear in short bursts and spread rapidly after release.
- Detection algorithms are broadly grouped into feature-based, graph-based, and propagation-modeling approaches.The survey presents a comprehensive overview and categorizes prior research by platform.
2 TYPES OF FALSE INFORMATION
The survey classifies false information along two dimensions: the author’s intent to deceive and whether the content concerns opinion or a single ground truth. It applies these distinctions to misinformation, disinformation, fake reviews, hoaxes, and fake news.
- Unified categorization: The survey’s categorization combines intent and knowledge content to describe different forms of false information.Figure 1 presents these dimensions as organizing criteria.
- Categorization by intent: Misinformation is created without intent to deceive, whereas disinformation is created with the intent to mislead and deceive readers.Misinformation may arise from misunderstanding, inattention, or cognitive biases; disinformation often aims to sway opinion or drive traffic for money.
- Categorization by knowledge: Opinion-based false information expresses an opinion without an absolute ground truth, while fact-based false information contradicts, fabricates, or conflates a single-valued ground truth.Fake product reviews exemplify opinion-based false information.
3 ACTORS AND RATIONALE OF SUCCESSFUL DECEPTION BY FALSE INFORMATION
False information succeeds through coordinated accounts, automated amplification, human susceptibility, and socially reinforced exposure. Experiments and platform studies show that readers often struggle to distinguish deceptive content from genuine information.
- Bad actors: Sockpuppet and sybil accounts fabricate apparent agreement by coordinating similar reviews or comments from a single underlying source.Readers may not realize that an apparently broad discussion originates from one source.
- Bad actors: Bots spread identical content quickly, inflate users’ social status, and target influential real users who may reshare false messages.Studies reported that bots produced almost one-fifth of Twitter political chatter and 25% of false-information tweets.
- Bad actors: On Twitter, humans—not bots—were responsible for false-information spread in an analysis of over 126,000 cascades across 11 years.Bots accelerated true and false information roughly equally, while false information still spread farther, deeper, faster, and broader than true information.
- Social reinforcement: Echo chambers expose users primarily to content aligned with their beliefs, leaving opposing groups mostly disconnected and potentially encouraging false-information spread.The #beefban retweet graph visualizes this separation between groups with opposing beliefs.
- Human susceptibility: Humans often perform near random when identifying deceptive reviews and hoaxes, including accuracies of 53.1%–61.9% for fake reviews and 66% for Wikipedia hoaxes.These findings include both casual and trained readers and extend across multiple false-information types.
- Human susceptibility: Humans can be deceived by intelligently crafted false information whether it is produced manually or by machines.A deep neural network generated restaurant reviews that workers evaluated alongside real reviews.
4 IMPACT OF FALSE INFORMATION
False information is usually ineffective, but a small fraction attracts disproportionate engagement, spreads rapidly, and reaches large audiences. Its impact is measured through views, shares, cascade structure, survival time, and cross-platform spread.
- A 12-hour average delay between false-information spread and debunking gives unverified rumors time to become viral.False information spreads rapidly during its initial phase, before verification or debunking.
- About 1% of Wikipedia hoaxes survive for over one year without detection, despite 90% being identified within an hour of approval.Hoax impact was assessed using viewcount, survival duration, and spread across the web.
- At least 5 distinct links led readers to 7% of Wikipedia hoaxes, with each hoax receiving 1.1 such links on average.Traffic came from search engines, social networks, and Wikipedia itself.
- False-information cascades on Facebook were deeper than reference cascades at greater reshare depths, across 16,672 cascades and 62,497,651 shares.Reference cascades were more common near the original post, whereas false cascades extended farther.
- Across over 126,000 Twitter rumors, false information spread farther, faster, deeper, and more broadly than true information in every studied category.The top 1% of false tweets reached over 1,000 users, a level true-information tweets rarely achieved.
- A small fraction of false information attracts more attention than true information, spreads widely and quickly, and reaches large populations.Most false information is ineffective, while the most impactful pieces generate unusually high engagement.
5 CHARACTERISTICS OF FALSE INFORMATION
The survey organizes detectable characteristics of false information across textual, temporal, rating, user, graph, and propagation dimensions. It presents these characteristics as tell-tale signs for distinguishing false from true information.
- False-information characteristics are studied across textual content, time, ratings, graph structure, creator properties, and propagation behavior.The survey separately discusses opinion-based and fact-based false information along these axes.
- Table 2 categorizes prior research according to the features used to analyze false information.The broad feature groups are text, user, graph, rating score, time, and propagation-based features.
5.1 Opinion-based false information
Opinion-based false information, especially fake reviews, exhibits distinctive textual, rating, temporal, and structural patterns. Fraudsters often produce extreme or bursty activity and coordinate accounts in lockstep.
- Textual characteristics: Duplicate reviews provide a fraud signal: 6% of reviewers with multiple reviews had a maximum inter-review similarity score of 1.Fraudulent duplicates were often posted by the same user across different products.
- Rating characteristics: Fraudulent reviewers commonly have strongly positive or negative rating distributions rather than the typical J-shaped aggregate pattern.Fraudulent and non-fraudulent users can both show bimodal posting distributions, but fraudsters are more polarized in ratings.
- Textual characteristics: Fake reviews tend to be shorter, more exaggerated, less readable, more polarized, and more sentimental than truthful reviews.They also show distinctive lexical, stylistic, psycholinguistic, and sentiment patterns.
- Temporal characteristics: Reviews posted at far-away places within very short interarrival times are a distinguishing characteristic of fraud.Researchers use interarrival-time distributions between successive ratings or reviews to detect spammers.
- Temporal characteristics: Fraudulent users are more bursty than non-fraudulent users, although both groups can alternate between short bursts and longer inactive periods.Non-fraudulent users may become active after inactivity to summarize recent experiences.
- Graph and coordination characteristics: Fraudsters repeatedly target the same products, form dense or temporally coherent groups, and coordinate reviews in lockstep.Large temporally coherent bipartite cores are highly suggestive of fraud, while fraudster groups can dominate a product’s reviewers.
5.2 Characteristics of fact-based false information
Fact-based false information—hoaxes, rumors, and fake news—has distinctive textual, user, network, propagation, and debunking characteristics. Across platforms, it tends to spread rapidly and deeply before corrections take effect, while a small number of users often drive its reach.
- Textual and user characteristics: Hoaxes are typically longer than non-hoaxes but contain fewer web and internal Wikipedia references.Their creators also tend to use newer accounts with less editing experience, although inexperienced editors can produce non-hoaxes that are mistakenly judged fraudulent.
- User and network characteristics: About one-fifth of election-related rumor content was created and spread by bots, which disseminated rumors in short bursts.False-information creators and spreading accounts can also form tightly connected groups, including clusters of alternate-media domains and coordinated bot accounts.
- Propagation characteristics: Top 30 users generated 90% of fake-image retweets during Hurricane Sandy, indicating that false-information spread is often dominated by a handful of highly active users.Fact-checking was more grassroots and conversational, whereas false-news spread was concentrated among very active users.
- Propagation characteristics: False information spreads deeper, faster, farther, and more broadly than true information, especially before verification or debunking.It also crosses platforms more often than true information: 18% of false-information pieces appeared on multiple platforms versus 11% of true pieces.
- Debunking characteristics: Debunking typically lags initial spread by 10–20 hours, but corrective links increase false-information deletion probability 4.4 times.After debunking, denial tweets become more common than supporting tweets, and corrective links are especially effective when shared shortly after the original post.
6 DETECTION OF FALSE INFORMATION
The survey organizes false-information detection into feature-based, graph-based, and propagation-modeling approaches, while emphasizing severe class imbalance, deceptive content, and costly labeling. Across reviews, news, rumors, and hoaxes, combining textual, user, network, temporal, and propagation signals supports effective detection, though content alone is increasingly insufficient.
- Detection approaches: Detection algorithms are broadly categorized as feature-based, graph-based, and propagation-modeling approaches.Feature-based methods derive discriminative signals from observed differences between true and false information; graph-based methods identify dense structures in user-information networks.
- Challenges: False-information detection is constrained by fewer than 10% positive instances, deceptive presentation, and the considerable effort required to obtain reliable labels.Labels are commonly produced by experts, trained volunteers, or Mechanical Turk workers, who may miss misinformation.
- Opinion-based false information: Text-based fake-review detection commonly combines engineered linguistic features with duplicate-review analysis and, where available, user, graph, score, and temporal information.Graph-based methods can identify fraudulent users, but user-level fraud and review-level falsity do not always coincide.
- Graph-based detection: Belief propagation transfers prior beliefs across rating-network neighbors until convergence, and SpEagle combines this process with node and edge features.SpEagle achieved area under the ROC curve scores around 0.78 on average across three Yelp fake-review datasets.
- Feature-based detection: 71%–78% accuracy was achieved when distinguishing fake from real news using text, but satire was harder to separate and titles were more informative than article bodies.The survey argues that newer content-based results indicate content alone is increasingly inadequate as malicious creators produce more genuine-looking false information.
- Propagation modeling: Propagation models learn from real true- and false-information cascades, predict false information, and can also support rumor-mitigation strategies.Reported approaches use propagation structure and timing for early detection and simulate interventions that spread corrective information alongside rumors.
7 DISCUSSIONS AND OPEN CHALLENGES
The survey synthesizes false-information mechanisms, impacts, characteristics, and detection across fake reviews, hoaxes, and fake news, while identifying major limitations and research directions. Progress is constrained by the lack of standardized large-scale datasets and by increasingly convincing machine-generated content.
- The survey presents a comprehensive view of mechanisms, rationale, impact, characteristics, and detection for fake reviews, hoaxes, and fake news.
- Large-scale publicly available datasets spanning these categories are lacking, preventing benchmark comparisons between detection-algorithm categories.
- Existing datasets including BuzzFeed, LIAR, CREDBANK, and FakeNewsNet have not yet received standardized comparisons of existing algorithms.
- Advances in machine learning can produce genuine-looking text, audio, images, and videos, making false information increasingly difficult for readers to distinguish from true information.
- Open research avenues include semantic-dissonance detection and fact-checking against knowledge bases, alongside work across multiple computational fields.