Source-linked AI summary
The COVID-19 Infodemic: Twitter versus Facebook
Kai-Cheng Yang, Francesco Pierri, Pik-Mai Hui, David Axelrod, Christopher Torres-Lugo, John Bryden, Filippo Menczer
TL;DR
The paper asks how COVID-19 misinformation spreads across Twitter and Facebook and compares prevalence, diffusion, influencers, coordination, and automation. Using cross-platform URL and account analyses, it finds divergent misinformation ecosystems, influential verified superspreaders, and coordinated sharing on both platforms. The authors conclude that the phenomenon is overt but that inconsistent data access limits direct comparison and study of harmful manipulation.
Problem
The study addresses limited cross-platform evidence on COVID-19 misinformation prevalence, diffusion, and the roles of different account groups.
Method
The study analyzes Twitter and Facebook links using credibility-domain lists, suspicious YouTube availability, and account co-domain similarity to examine prevalence and coordination.
Results
Low-credibility ecosystems diverge across platforms, while high-profile verified accounts and coordinated sharing strongly contribute to Infodemic dissemination and automated accounts do not appear to play a strong role.
Takeaways & Limitations
The Infodemic is characterized as overt rather than covert, supporting societal-level responses alongside platform mitigation strategies.
Takeaways & Limitations
Differences in platform data availability, sampling and selection biases, and keyword-based collection make direct and fair comparisons impossible in many cases.
Abstract
from arXiv · showhide
The global spread of the novel coronavirus is affected by the spread of related misinformation -- the so-called COVID-19 Infodemic -- that makes populations more vulnerable to the disease through resistance to mitigation efforts. Here we analyze the prevalence and diffusion of links to low-credibility content about the pandemic across two major social media platforms, Twitter and Facebook. We characterize cross-platform similarities and differences in popular sources, diffusion patterns, influencers, coordination, and automation. Comparing the two platforms, we find divergence among the prevalence of popular low-credibility sources and suspicious videos. A minority of accounts and pages exert a strong influence on each platform. These misinformation "superspreaders" are often associated with the low-credibility sources and tend to be verified by the platforms. On both platforms, there is evidence of coordinated sharing of Infodemic content. The overt nature of this manipulation points to the need for societal-level solutions in addition to mitigation strategies within the platforms. However, we highlight limits imposed by inconsistent data-access policies on our capability to study harmful manipulations of information ecosystems.
1 Introduction
The study examines how COVID-19 misinformation spreads across Twitter and Facebook, addressing gaps in cross-platform prevalence, diffusion, influence, coordination, and automation. It identifies low-credibility content through source-domain matching and suspicious YouTube availability status.
- The study compares COVID-19 low-credibility content across Twitter and Facebook, where platform-specific vulnerabilities and concurrent propagation remain insufficiently understood.
- Diffusion analysis asks whether a few accounts, pages, or groups disproportionately amplify misinformation and whether sharing patterns indicate coordination.
- Links are classified as low-credibility by matching domains to an independent source corpus, while YouTube videos are labeled suspicious when banned or unavailable.
- The research questions cover prevalence, source-sharing similarities, influential accounts, verified accounts, coordinated behavior, and Twitter-bot amplification.
- The paper is organized around literature review, methodology, results addressing the research questions, and discussion of limitations and mitigation implications.
2 Literature review
Prior research documents health misinformation across social media but often relies on single-platform, small-scale, or methodologically inconsistent studies. The literature therefore leaves cross-platform spreading patterns and the roles of different account groups insufficiently understood.
- Health misinformation research spans web quality assessment, infodemiology, epidemic studies, and social-media analyses across multiple platforms.
- Existing studies commonly analyze sampled posts, images, or videos, but datasets were usually limited to hundreds or thousands of items because of access and manual-analysis constraints.
- Reported COVID-19 misinformation prevalence estimates range from 1% to 70%, reflecting differences in experimental design, platform access, and misinformation definitions.
- The review identifies two gaps: rare multi-platform comparisons and limited understanding of how account groups contribute to misinformation dissemination.
3 Methods
The methods collect Twitter and Facebook data using common keywords, classify linked content through predefined credibility-domain lists, and identify suspicious YouTube videos by availability status.
- Twitter and Facebook data are collected using the same keyword list to support a shared analytical framework.
- Posts are automatically classified as low- or high-credibility by tracking URLs that link to domains in predefined lists.
- Suspicious YouTube videos are identified through their availability status, including whether the uploads remain publicly accessible.
3.1 Identification of low-credibility information
Low-credibility information is operationalized at the source level using an external reliability corpus, producing a list of 674 domains for URL matching.
- The study identifies low-credibility news links by matching URLs to domains rather than evaluating individual articles.
- The domain list uses Media Bias/Fact Check ratings, including sources labeled “Very Low,” “Low,” “Questionable,” or “Conspiracy-Pseudoscience.”
- 674 low-credibility domains are included, while sources rated “Mostly-Factual,” “High,” or “Very High” are excluded.
3.2 High-credibility sources
The study curates 20 higher-credibility information sources as a benchmark for interpreting low-credibility content prevalence across the platforms.
- The benchmark contains 20 popular news outlets spanning the full U.S. political spectrum.The selected outlets have an MBFC factual-reporting level of “Mixed” or higher.
- The benchmark also includes the CDC and WHO as authoritative COVID-19 information sources.
3.3 Data collection
The study collects COVID-19-related Twitter and Facebook data, preserving different platform structures and interaction information while accounting for Facebook coverage limitations.
- Platform data structure: Twitter records original tweets, retweets, and all participating accounts, whereas Facebook records original posts and publishing pages or groups.Facebook provides aggregate reshares, comments, and reactions but no identities for users responsible for those interactions.
- Twitter data: Twitter data comprise over 53M English tweets from about 12M users collected between Jan. 1 and Oct. 31, 2020.The data come from the Decahose, a 10% random sample of public tweets, which is biased toward more active users because it samples tweets rather than users.
- Facebook data: Facebook data comprise over 37M English posts from over 140k public pages and groups collected during the same period.The collection used the CrowdTangle posts/search endpoint and filtered posts using generic COVID-19 keywords.
- Facebook data: CrowdTangle coverage may bias Facebook data because inclusion depends partly on audience size and researcher requests.Researcher requests focused on low-credibility pages or groups could over-represent such content.
- Facebook data: Facebook comments and reactions are highly correlated with reshares, so the study focuses on reshares.
- Cross-platform comparison: The study compares Twitter roots with Facebook roots, original tweets with original posts, and retweets with reshares.Prevalence is defined as original tweets plus retweets on Twitter and original posts plus reshares on Facebook.
- YouTube data: For YouTube, the study examines popular videos shared on both platforms and uses removal or privatization as a proxy for low-credibility content.It selects 16,669 videos shared on both platforms, of which 1,828 (11%) were removed or private, followed by manual inspection of about 3%.
- Research compliance: Data collection and analysis complied with platform terms of service and received IRB-exemption determinations.
3.4 Link extraction
The study extracts and expands social-media URLs before matching them to low- and high-credibility domain lists, revealing different amplification ratios across platforms.
- The study identified 49 frequently occurring URL-shortening services and expanded their links through HTTP requests.The services appeared at least 50 times in the datasets before expansion and domain matching.
- Expanded and extracted URLs were matched against lists of low- and high-credibility domains.
- 2.7:1 is the ratio of retweets to tweets for low-credibility content, compared with 68:1 for reshares to posts.The discrepancy reflects platform traffic differences, the 10% Twitter sample, and sampling bias.
4 Infodemic prevalence
Low-credibility COVID-19 content surged on Twitter and Facebook during the initial pandemic wave, with strongly correlated temporal patterns but platform-specific source popularity. Twitter had a higher average low-to-high-credibility ratio, while suspicious YouTube videos showed little cross-platform rank alignment.
- 4 Infodemic prevalence: Prevalence estimates are lower bounds because deleted content was excluded from the data.Twitter counts included tweets and retweets; Facebook counts included original posts and reshares.
- 4.1 Prevalence trends: Twitter and Facebook low-credibility link volumes were strongly correlated (Pearson r = 0.87, p < 0.01) and grew sharply during March.Both declined toward summer and then stabilized at relatively low levels.
- 4.1 Prevalence trends: The Infodemic surge roughly coincided with general pandemic attention, while worldwide hospitalization peaks trailed by a few weeks.The comparison used overall pandemic-related tweet volume and worldwide hospitalization rates.
- 4.1 Prevalence trends: 32% vs. 21% on average: Twitter had a higher ratio of low-credibility to high-credibility information than Facebook.The ratios remained relatively stable across the observation period, suggesting that the decline tracked broader pandemic attention.
- 4.2 Prevalence of specific domains: Low-credibility content combined surpassed every individual high-credibility domain in volume on both platforms, although most individual low-credibility domains were less prevalent.The Gateway Pundit and Breitbart were notable exceptions.
- 4.2 Prevalence of specific domains: Cross-platform source ranks were only moderately aligned (Spearman r = 0.57, p < 0.01), with some domains much more popular or exclusive to one platform.Figure 5 compares domain prevalence ranks and highlights sources with large discrepancies.
- 4.3 Suspicious YouTube videos: Suspicious YouTube videos had no qualitative cross-platform popularity correlation, despite being linked 6–980 times on Twitter and 39–64,257 times on Facebook.The analysis covered unavailable videos ranked within the top 500 on both platforms; repeated or re-edited uploads may share content but have different IDs.
- 4.3 Suspicious YouTube videos: Higher-prevalence videos were more often unavailable on both platforms, with the trend stronger on Twitter than on Facebook.This pattern suggests YouTube moderation may disproportionately affect videos attracting more attention.
5 Infodemic spreaders
Infodemic dissemination was highly concentrated around a small set of influential accounts despite many roots contributing content. Verified accounts and official source-associated accounts generated a disproportionate share of retweets and reshares on both platforms.
- 5.1 Concentration of influence: The concentration measure uses inverse normalized entropy, ranging from 0 for equal root contributions to 1 when one root supplies all content.The measure is set to Cs = 1 when Ns = 1.
- 5.1 Concentration of influence: Popularity was significantly more concentrated around root accounts than original activity on both platforms (p < 0.001), indicating Infodemic superspreaders.A diverse set of roots published low-credibility links, but messages from a small influential group were shared extensively.
- 5.2 Who are the Infodemic superspreaders?: Only 19% of Twitter and 21% of Facebook top spreaders for individual low-credibility sources were verified.Verification was overrepresented among superspreaders but did not characterize all of them.
- 5.2 Who are the Infodemic superspreaders?: Official source-associated accounts were the top spreaders in 16 of 21 Twitter cases and 18 of 23 Facebook cases.Most top low-credibility sources had official accounts on both platforms, and these accounts were often verified: 71.4% on Twitter and 65.2% on Facebook.
6 Infodemic manipulation
The study identifies coordinated sharing and limited bot amplification of low-credibility COVID-19 content across Twitter and Facebook, while emphasizing that platform comparisons are constrained by different data and sampling conditions.
- Coordinated amplification of low-credibility content: Twitter coordination used a 0.99 similarity threshold and at least 10 low-credibility-link tweets, whereas Facebook used 0.95 and at least 5 linked posts.The thresholds were selected by manually inspecting the outputs.
- Coordinated amplification of low-credibility content: Accounts and pages formed densely connected clusters by sharing unusually similar sets of low-credibility domains on both platforms.The analysis represents accounts as TF-IDF-weighted domain vectors and links pairs using cosine similarity.
- Coordinated amplification of low-credibility content: The suspicious clusters were predominantly right-leaning and U.S.-centric, with Twitter clusters concentrated on leading low-credibility sources and Facebook clusters linking to more varied sources.Clusters on both platforms also shared Russian state-affiliated media and an Indian right-wing magazine.
- Coordinated amplification of low-credibility content: Some Facebook clusters involved organizations with reach beyond the platform or pages and groups made to appear credible through verification.Examples included pages associated with The Answer radio stations.
- Automated amplification: BotometerLite assigned Twitter accounts scores from 0 to 1, using 0.5 to classify likely humans and likely bots; Facebook automation detection was unavailable.The analysis therefore limited automation estimates to Twitter.
- Automated amplification: The fitted power-law exponent was γ ≈1.04, suggesting a weak 4% level of bot amplification on Twitter.The relationship compares total original tweets plus retweets authored by likely humans and bots for each domain.
- Automated amplification: Sources with more Twitter bot activity were equally shared across Twitter and Facebook, despite asymmetric platform prevalence for some low-credibility sources and suspicious videos.The cross-platform rank comparison used source-rank differences to color the Figure 10 points.
7 Discussion
The discussion reports cross-platform differences in low-credibility content, coordinated amplification, and account influence, while cautioning that platform data biases limit direct comparisons. It concludes that the Infodemic appears overt and remains difficult to address because high-status accounts participate in spreading it.
- Main findings: High-profile, official, and verified accounts tend to be primary drivers of low-credibility COVID-19 information, alongside coordination on both platforms.The authors characterize the Infodemic as overt rather than covert because automated accounts do not appear to play a strong amplifying role.
- Cross-platform prevalence: Low-credibility content as a whole had higher prevalence than content from any single high-credibility source, while websites and suspicious videos differed in prevalence between platforms.The discussion attributes the discrepancy potentially to platform-specific source supply and user demographics.
- Temporal patterns: Similar surges of low-credibility content occurred on Twitter and Facebook during the pandemic’s first months.The strong correlation between low- and high-credibility timelines suggests peaks were likely driven by public attention rather than bursts of malicious content.
- Platform comparison: Facebook had a lower low-to-high-credibility information ratio than Twitter, but verified accounts played a stronger role in spreading low-credibility content on Facebook.The authors caution that different data-collection biases affect the accuracy of these comparisons.
- Platform comparison: Asymmetric suspicious-video prevalence may reflect duplicated uploads, different audience sharing, or faster removal after Twitter users flag videos.The proposed explanations are presented as possibilities rather than established causes.
- Limitations: Direct and fair platform comparisons are often impossible because Twitter and Facebook data differ in availability, sampling, selection, keyword coverage, and source-level labeling.Twitter data overrepresent active users, while Facebook data emphasize popular pages and public groups.
- Implications: Because high-status accounts play an important role, addressing pandemic misinformation is difficult amid moderation, political-bias, free-speech, and censorship concerns.The authors frame these issues as important unresolved parts of improving the information ecosystem.