Source-linked AI summary
An Exploratory Study of COVID-19 Misinformation on Twitter
Gautam Kishore Shahi, Anne Dirkson, Tim A. Majchrzak
TL;DR
COVID-19 misinformation was spreading rapidly on social media, while existing Twitter studies relied on small samples or disputed source-reliability proxies. This exploratory study analyzed fact-checked COVID-19 tweets and found that false tweets propagated faster than partially false tweets, while offering early insights and research gaps for crisis response.
Problem
Existing Twitter studies of COVID-19 misinformation were limited by small samples or source-based classification whose reliability remains disputed.
Method
The study used social media analytics to explore COVID-19 misinformation in tweets identified through fact-checking articles, collecting 3,053 tweet IDs from 7,623 articles and analyzing 1,565 deduplicated tweets.
Results
False tweets propagated faster than partially false tweets, with propagation speed highest during peak tweet periods; misinformation authors also appeared more driven by preventing harm than by rewards or money.
Takeaways & Limitations
The exploratory findings provide early insights, lessons learned, practitioner recommendations, and open questions about COVID-19 misinformation during an unfolding crisis.
Takeaways & Limitations
The dataset may exclude less viral rumours because it includes only claims investigated by fact-checking organisations, and the analysis is limited to Twitter.
Abstract
from arXiv · showhide
During the COVID-19 pandemic, social media has become a home ground for misinformation. To tackle this infodemic, scientific oversight, as well as a better understanding by practitioners in crisis management, is needed. We have conducted an exploratory study into the propagation, authors and content of misinformation on Twitter around the topic of COVID-19 in order to gain early insights. We have collected all tweets mentioned in the verdicts of fact-checked claims related to COVID-19 by over 92 professional fact-checking organisations between January and mid-July 2020 and share this corpus with the community. This resulted in 1 500 tweets relating to 1 274 false and 276 partially false claims, respectively. Exploratory analysis of author accounts revealed that the verified twitter handle(including Organisation/celebrity) are also involved in either creating (new tweets) or spreading (retweet) the misinformation. Additionally, we found that false claims propagate faster than partially false claims. Compare to a background corpus of COVID-19 tweets, tweets with misinformation are more often concerned with discrediting other information on social media. Authors use less tentative language and appear to be more driven by concerns of potential harm to others. Our results enable us to suggest gaps in the current scientific coverage of the topic as well as propose actions for authorities and social media users to counter misinformation.
1. Introduction
COVID-19 misinformation posed an urgent infodemic challenge, while existing Twitter studies left gaps in understanding its origins, spread, language, and account involvement. This exploratory study addresses those gaps and develops recommendations for practitioners.
- Motivation: COVID-19 misinformation spread rapidly during a global health crisis, making information quality important to the public response.The paper frames misinformation as part of the COVID-19 infodemic and notes that individual actions depend on available information.
- Research gap: Existing Twitter research relied partly on small manually annotated samples or disputed source-reliability proxies.The authors identify limitations in both small-scale manual annotation and source-based automatic classification.
- Research aims: The study explores which accounts create or spread COVID-19 misinformation, how it propagates, and what its content distinguishes from other COVID-19 tweets.The exploratory design reflects limited prior knowledge about the topic.
- Contributions: The paper synthesizes social-media analytics techniques, contributes early findings about misinformation, and proposes recommendations for authorities, crisis managers, and social-media listeners.These are presented as three explicit contributions.
- Study positioning: The article prioritizes rapid dissemination of early findings while retaining academic rigor.The authors characterize the work as uncommon because it combines crisis-era timeliness with academic oversight.
2. Background
The background distinguishes misinformation from related concepts and reviews approaches to sampling, propagation, and correction. It motivates comparing completely and partially false claims because their spread and believability may differ.
- Definitions: Misinformation is broadly defined here as circulating information that is false, without claims about whether dissemination was accidental or deliberate.The study therefore groups false information regardless of intent and avoids the term “fake news.”
- Definitions: Partially false claims contain elements of truth but become misleading through features such as miscaptioning or omitted background information.The paper treats falsity as a scale rather than a strict true-or-false boundary.
- Sampling and classification: The study uses manual fact-checker evaluations to classify claims as false or partially false.This distinction supports comparisons of how completely and partially false information spreads.
- Sampling: Top-down sampling uses already fact-checked rumours but misses rumours absent from fact-checking databases, whereas bottom-up sampling covers more rumours but requires manual annotation.This study adopts the top-down strategy using Snopes and more than 91 organisations.
- Propagation: Twitter propagation research commonly uses epidemiological models or retweet trees to quantify spread and network structure.Retweet-tree measures include depth, size, and breadth, while structural virality captures peer-to-peer and broadcast mechanisms.
- Correction: Fact-checking and community corrections can reduce sharing temporarily but may also produce mixed or delayed effects.The reviewed evidence includes persistent resharing after debunking, slow self-correction, and possible backfire effects.
3. Data collection & preprocessing
The study constructs a fact-checking-based misinformation corpus and a same-period COVID-19 background corpus, then retrieves, filters, normalizes, and preprocesses tweets for analysis. Retweets and account metadata support propagation and author analyses.
- Data collection: 3 053 Tweet IDs were extracted from 7 623 fact-checked articles, yielding 1 565 unique tweets after deduplication.The tweets span 14 January to 10 July 2020.
- Data collection: Tweets were retrieved through fact-checked article links, Twitter identifiers, and the Twitter API.Figure 1 depicts extracting Tweet Links from fact-checked articles and fetching the corresponding tweets.
- Background corpus: The background corpus contains 163 096 English COVID-19 tweets sampled across the same period as the misinformation corpus.It combines TweetsCOV19 for January–April with in-house crawling for May–July; API retrieval produced 92 095 available tweets from the first source.
- Metadata and retweets: Retweets were collected with Twarc to compare propagation of false and partially false tweets.The authors also gathered favourites, friends, followers, account age, profile descriptions, and locations for account analysis.
- Preprocessing: The researchers normalized 18 fact-checker verdict classes into study categories and translated and tokenized multilingual tweets.They removed emojis, mentions, URLs, and the hashtag symbol while retaining hashtag text.
- Filtering: The final study dataset contains 1 500 tweets classified as false or partially false after excluding nonconforming verdict categories.The passage reports 1 274 false and 226 partially false claims.
4. Method
The method uses a two-way exploratory analysis of misinformation-related user accounts and tweet content to investigate how misinformation propagates on social media.
- Analytical approach: The analysis examines both user-account characteristics and tweet content.Together, these analyses target the propagation of false and partially false information on social media.
- Analytical approach: Propagation is measured through retweet-based analyses of misinformation spread.The study compares false and partially false data within its two-way framework.
4.1. Account categorisation
The study categorizes Twitter accounts involved in COVID-19 misinformation by automation, branding, and popularity-related characteristics.
- Account categorisation: Account analysis uses a bot detection API, a brand classifier, and popularity indicators such as follower count.The study also examines whether accounts are verified.
- Bot detection: The account analysis considers whether bots participate in spreading misinformation through Twitter activity such as tweeting or retweeting.The paper describes bots as programs operating accounts through the Twitter API.
- Account categorisation: Brands are defined here as organizations or celebrities with many followers and greater attention or reachability.Their larger follower networks and retweet counts can increase public attention.
- Account categorisation: Verified and popular accounts are treated as potentially more influential because their posts are visible to followers and attract more attention.The paper motivates examining these accounts because misinformation from popular handles may reach more users.
4.2. Information diffusion
The study measures misinformation diffusion through retweet speed, comparing false and partially false tweets across overall, tweet-specific peak, and crisis-peak periods.
- Propagation of misinformation: The analysis uses retweets as a proxy for propagation speed and considers only retweets of the original tweet.A retweet may include an additional comment, but it reposts the tweet to a user’s follower network.
- Propagation of misinformation: Propagation speed is defined as a tweet’s total retweet count divided by the total number of days it receives retweets.Ps denotes propagation speed, rc retweet count per day, and Nd total days.
- Propagation of misinformation: Three metrics capture overall propagation, propagation during a tweet’s first peak, and propagation during 15-03-2020 to 15-04-2020.The tweet-specific peak ends when retweets first reach zero; the crisis period was selected because retweet activity was maximal then.
- Study boundary: Real-time detection is not possible in this design because fact-checking websites may take days to verify claims.The study therefore compares propagation speed between false and partially false tweets rather than detecting them immediately.
4.3. Content analysis
Content analysis compares misinformation tweets with a background corpus to identify distinctive topics, language, hashtags, emojis, and psychologically relevant language patterns.
- Content analysis: Because partially false claims are relatively few, the content analyses combine false and partially false data.The analyses examine hashtags, emojis, distinctive terms, and language differences from COVID-19 content generally.
- Content analysis: The study compares misinformation tweets with English COVID-19 tweets from 14-01-2020 to 10-07-2020 as a background corpus.This comparison is used to identify the most distinctive phrases in the misinformation corpus.
- Hashtags and emojis: Hashtags are analyzed as self-reported topics, while emojis are analyzed as proxies for authors’ self-reported emotions.The study identifies the top 10 hashtags and detects emojis using a dedicated package.
- Distinctive terms: KLIP is applied to unigrams, bigrams, and trigrams to identify terms that distinguish misinformation from the background corpus.Its informativeness component compares probability distributions and identifies terms with the largest information loss.
- Distinctive terms: Duplicate tweets about the same misinformation are removed because they could bias distinctive-term analysis toward particular claims.The removed tweets express slightly different versions of the same misinformation.
- Language analysis: LIWC 2015 measures the relative frequency of words across emotional, social, cognitive, motivational, temporal, personal-concern, and informal-language categories.The categories are based on manually curated word lists.
5. Results
Across 1,500 misinformation tweets and 1,187 unique accounts, false claims propagated faster than partially false claims, especially during the pandemic peak. Compared with general COVID-19 tweets, misinformation differed in topics, language, emotional expression, and apparent author motivations.
- Accounts: 1,500 misinformation tweets came from 1,187 unique accounts, including 24 accounts classified as bots using a CAP threshold above 0.65.The accounts were categorized after filtering the misinformation dataset for unique users.
- Information diffusion: Misinformation activity peaked from mid-March to mid-April 2020, when false and partially false tweets were most numerous and false-category spread was faster.The peak is also described as 16th March to 23rd April 2020 in the timeline analysis.
- Information diffusion: False-category misinformation spread faster than partially false misinformation, with a significant difference in propagation speed, X2 (3, N = 1500) = 10.23, p <.001.Propagation speed was highest during the peak COVID-19 period.
- Hashtag analysis: Common hashtags focused on COVID-19, stopping the virus, and calling out alleged misinformation or hoaxes, including #covid19, #stopcorona, and #fakenews.The analysis also identified location-linked hashtags associated with claims from Madagascar, Mérida, and Daraq.
- Distinctive terms: Compared with general COVID-19 tweets, misinformation more often concerned discrediting information circulating on social media; completely false and partially false misinformation emphasized different topics.Completely false misinformation mentioned health-governing bodies more often, whereas partially false misinformation focused more on transmission, mortality, and updates.
- Psycho-linguistic analysis: Misinformation tweets used less tentative language and fewer emotion-related words than general COVID-19 tweets, while showing greater apparent affiliation-driven motivation and less emphasis on rewards or achievements.Positive and negative emotions were significantly less prevalent, while anxiety did not differ significantly.
6. Discussion
The discussion translates the exploratory findings into recommendations while emphasizing constraints from Twitter data access, selection, interpretation, and ethical profiling risks.
- Lessons learned: Twitter’s API restricts retrospective analysis by limiting access to replies older than seven days and retweets.These restrictions hinder reconstruction of public reactions and misinformation propagation before fact-checking.
- Lessons learned: Around 90% of collected claims were excluded because they lacked a tweet referenced by a fact-checking organisation.Similarity matching could enlarge coverage but would be slower and noisier.
- Recommendations: Authorities should monitor social media, tailor online responses, and use analytics to track developing misinformation.The authors caution that these recommendations are initial aid rather than definite guidelines.
- Recommendations: Approximately 70% of false and partially false misinformation involved brands, organisations, or celebrities creating or circulating content.The authors recommend monitoring these accounts because they may spread misinformation through activities such as liking or retweeting.
- Recommendations: Partially false claims warrant particular study because their slower propagation may make them harder for users to recognise as false.Their truthful elements may complicate recognition despite slower spread.
- Limitations: The study’s interpretation of hashtags and emojis is limited because their meanings are culturally, contextually, and temporally variable.The authors also note that exposure to COVID-19 misinformation was not measured by this study.
7. Conclusion
The paper presents an exploratory analysis of COVID-19 misinformation on Twitter, offering early findings, practitioner recommendations, and research questions. It frames these contributions as preliminary work intended to support crisis mitigation and future quantitative verification.
- Conclusion: The study analyses fact-checked COVID-19 misinformation tweets using social media analytics in response to an unfolding crisis.The exploratory design enabled early insights but also brought severe limitations.
- Conclusion: The paper provides rich results, lessons learned, initial recommendations for practitioners, and open questions about COVID-19 misinformation.The authors describe the research gaps as unsurprising because misinformation was being studied early in the crisis.
- Conclusion: The work aims to contribute to mitigating the crisis and to stimulate research making social media a more reliable data source.The authors also highlight the need for broader action to address the infodemic.
- Future work: The authors plan to verify their early findings quantitatively with much larger datasets and collaboration providing access to historical Twitter data.This is presented as a continuation in a less exploratory fashion.
CRediT authorship contribution statement
The authors divide responsibilities across data collection, methodology, software, visualization, investigation, conceptualization, funding, project administration, and writing.
- Contributions: Gautam Kishore Shahi contributed to data collection, curation, investigation, methodology, software, visualization, and writing.
- Contributions: Anne Dirkson contributed to data curation, investigation, methodology, software, visualization, and writing.
- Contributions: Tim A. Majchrzak contributed conceptualization, funding acquisition, methodology, project administration, and writing.
Author biographies
The authors are researchers spanning web science, data science, health-related social media, natural language processing, information systems, and emergency management.
- Author biographies: Gautam Kishore Shahi is a PhD student whose research interests include web science, data science, and social media analytics.
- Author biographies: Anne Dirkson is a PhD student studying knowledge discovery from health-related social media and researching natural language processing and text mining.
- Author biographies: Tim A. Majchrzak is an information systems professor and member of a centre for integrated emergency management.