Source-linked AI summary
Characterizing COVID-19 Misinformation Communities Using a Novel Twitter Dataset
Shahan Ali Memon, Kathleen M. Carley
TL;DR
COVID-19 misinformation hampers communication and decision-making, making it important to identify and characterize the communities involved in spreading misinformation and correcting it. The paper builds a diverse annotated Twitter dataset and compares competing informed and misinformed communities across network, linguistic, and membership dimensions. Misinformed communities are denser and more organized, may include substantial disinformation activity, and contain many users identified as anti-vaxxers, while informed users use more narrative language.
Problem
COVID-19 misinformation disrupts communication and decision-making, while identifying misinformation and the communities involved remains difficult because data are scarce and themes are diverse.
Method
The paper constructs a diverse annotated COVID-19 Twitter dataset and characterizes informed and misinformed communities through network, sociolinguistic, bot, and cross-community analyses.
Results
Misinformed communities are denser and more organized than informed communities; bots are more prevalent among misinformed users, many may be anti-vaxxers, and informed users use more narratives.
Takeaways & Limitations
COVID-19 misinformation communities are polarized, highly organized, and potentially connected to disinformation campaigns and anti-vaccination communities.
Takeaways & Limitations
Most analyses rely on annotations from one annotator, are correlational rather than causal, and use data collected over a three-week period with timeline augmentation.
Abstract
from arXiv · showhide
From conspiracy theories to fake cures and fake treatments, COVID-19 has become a hot-bed for the spread of misinformation online. It is more important than ever to identify methods to debunk and correct false information online. In this paper, we present a methodology and analyses to characterize the two competing COVID-19 misinformation communities online: (i) misinformed users or users who are actively posting misinformation, and (ii) informed users or users who are actively spreading true information, or calling out misinformation. The goals of this study are two-fold: (i) collecting a diverse set of annotated COVID-19 Twitter dataset that can be used by the research community to conduct meaningful analysis; and (ii) characterizing the two target communities in terms of their network structure, linguistic patterns, and their membership in other communities. Our analyses show that COVID-19 misinformed communities are denser, and more organized than informed communities, with a possibility of a high volume of the misinformation being part of disinformation campaigns. Our analyses also suggest that a large majority of misinformed users may be anti-vaxxers. Finally, our sociolinguistic analyses suggest that COVID-19 informed users tend to use more narratives than misinformed users.
1 Introduction
COVID-19 misinformation has created a global infodemic that disrupts communication and decision-making. Effective debunking requires identifying misinformation and the communities that spread it, but this is difficult because data are scarce and themes are diverse.
- COVID-19 political and medical misinformation has created a global infodemic that hampers communication and affects decision-making.
- Debunking false information is important because undisputed misinformation can exacerbate epidemic spread.
- Effective intervention requires identifying both misinformation and the misinformed communities that spread it.
- This identification is challenging because data are scarce and misinformation themes are diverse.
2 Background
Existing COVID-19 datasets vary in annotation, modality, language, and scope, but many lack manual labels or focus specifically on misinformation. Prior studies also examine misinformation users, bots, credibility, content, and claims from several perspectives.
- Many COVID-19 datasets are generic and lack annotations or labels, including multilingual, longitudinal, location-based, and Arabic Twitter resources.
- COVID-19 misinformation datasets include automatically annotated, distantly supervised, manually annotated, multilingual, multimodal, and large-scale fake-news resources.
- Alam et al.’s dataset offers fine-grained annotation but contains only a few hundred tweets, whereas this study reports broader topic diversity.
- Prior research has categorized or identified misinformed users, studied conspiracy theories propagated by bots, measured low-credibility information, and analyzed COVID-19 content, sources, and claims.
3 Methodology
The study collected COVID-19 tweets through keyword-based Twitter searches on three dates, randomly sampled them for annotation, and developed a 17-category classification scheme. A second annotation phase assigned a subset of tweets to additional annotators.
- Tweets were collected with Twitter’s search API on 29 March, 15 June, and 24 June 2020, with each collection covering its corresponding week.
- The collection used hashtags and keywords in conjunction with “coronavirus” and “covid” to retrieve Twitter data.
- The annotation task classified tweets into 17 categories defined in a publicly available codebook with detailed definitions and examples.
- 4,573 tweets were annotated by one annotator in the first phase, and 651 were randomly assigned to six additional annotators in the second phase.
4 Data Description
The CMU-MisCOV19 dataset was designed to represent diverse misinformation and informed-community categories, including true prevention and correction. It contains 4,573 annotated tweets from 3,629 users and is released with annotations and a codebook for reproducible use.
- The dataset covers diverse information and misinformation categories, including “True Prevention,” “Calling out/correction,” “True Public Health Response,” and “Sarcasm.”
- True-information categories complement false-information annotations, which the paper states is necessary for building models.
- 4,573 annotated tweets comprise 3,629 users, averaging 1.24 tweets per user, across a wide range of categories and topics.
- CMU-MisCOV19 provides tweet IDs, annotations, creation dates, and a public codebook, while omitting full tweet JSONs to comply with Twitter’s terms.
5 Analysis and Discussion
The analysis identifies competing informed and misinformed communities from annotated true- and false-information categories, then augments their data with users’ timelines for behavioral analyses.
- Community identification: The dataset spans true and false information categories used to distinguish informed users from users actively posting misinformation.The supplied passages list categories including True Treatment, Correction/Calling Out, Conspiracy, Fake Cure, and False Public Health Response.
- Data characterization: Figure 1 charts the frequency of each identified topic across all tweets, although some tweets may have multiple topics.Topic frequencies therefore need not correspond to mutually exclusive tweet counts.
- Community identification: Users are assigned community membership from weighted annotation valences, with true-information categories positive and misinformation categories negative.Valence is assigned to annotations rather than tweets, allowing multiple annotators’ labels to be combined before computing each user’s weighted valence.
- Data augmentation: The study augments each community with users’ timelines to reduce survivorship bias in network, bot, and sociolinguistic analyses.Only COVID-19-related tweets are extracted from the collected timelines for these analyses.
5.3 Network Analysis
The network analysis combines retweet, mention, and reply ties to compare the two communities. Both show echo-chamberness, but misinformed sub-communities are denser while reply networks show more intergroup engagement.
- Combined network: Figure 2 displays the combined retweet, mention, and reply network with informed users in green and misinformed users in red.Users with unidentified or ambiguous membership are removed from the graph.
- Combined network: The combined retweet, mention, and reply network shows echo-chamberness in both communities, with misinformed sub-communities much denser than informed sub-communities.Network density is defined as the ratio of actual to potential connections.
- Combined network: Some two-way communication is present between the informed and misinformed sides.This communication is observed despite the broader echo-chamber pattern.
- Network measures: Table 3 reports nodes, links, and network density for the two target sub-communities.These quantities support comparison of the communities’ network structure.
- Separate networks: Reply networks show more intergroup engagement than retweet or mention networks, although the reply network is small.The authors hypothesize that this pattern may reflect corrective or calling-out behavior.
5.4 Bot Detection
Bot detection compares potential bot prevalence between the two competing communities using Bot-Hunter and a two-sample z-test. Misinformed users have the higher bot percentage, with a statistically significant difference.
- Method: Bot-Hunter was used to identify potential bot-like accounts, and a two-sample z-test assessed the difference in bot proportions between groups.Bot-Hunter is reported with precision .957 and recall .704; the test used α = 0.05.
- Results presentation: Table 4 reports the number and percentage of bots within each competing group.The table is the reported location of the bot-group comparison results.
- Bot prevalence: 19% of identified misinformed users were bots, compared with 11% of identified informed users.The difference was statistically significant (p < 0.001; z = −6.23).
- Bot prevalence: 14% (505) of 3629 users were identified as bots overall.Bot-Hunter identified potential bot-like accounts using a probability threshold of at least .75.
- Interpretation: The authors interpret the higher bot prevalence among misinformed users as indicating that more than 1/5th of misinformation-related posts may result from COVID-19 disinformation campaigns.This is presented as a potential contribution rather than a definitive attribution.
5.5 Sociolinguistic Analysis
The study compares COVID-19 informed and misinformed communities using normalized LIWC measures across narrative discourse, tone, and linguistic formality. Informed users use more narrative-associated language, while both groups are highly negative and formality differences are inconclusive.
- Analysis design: LIWC analyzes lexical categories with psychological meaning by calculating their percentage in each text.The analysis uses COVID-19-relevant tweets, excludes identified bots, normalizes by user activity, and averages user-level indices.
- Narrative discourse structure: Informed users use more pronouns, function words, and family-related keywords, while also being less analytical and more authentic than misinformed users.These measures are treated as proxies for narrative discourse, leading the authors to infer greater narrative use among informed users.
- Tone: Both communities are highly negative, with no significant difference in emotional tone.LIWC tone scores above 50 indicate greater positivity, while scores below 50 typically indicate negativity.
- Linguistic formality: Misinformed users tend to be more informal, whereas informed users tend to use more swear words, but these differences are not statistically significant.The authors therefore characterize the formality findings as inconclusive.
5.6 Vaccination Stance
The study examines vaccination stance within the COVID-19 misinformed community using vaccine-related users and hashtag co-occurrence. Anti-vaxxers constitute a larger share than pro-vaxxers, and bot prevalence supports the possibility that some misinformation is intentional and organized.
- Method: The authors identify vaccination stance by analyzing misinformed users who posted vaccine-related tweets and constructing a user-to-hashtag co-occurrence network.Vaccination hashtags receive valence labels through a label-propagation method.
- Vaccination stance: Among 1027 COVID-19 misinformed users, 41% are identified as anti-vaxxers and 22% as pro-vaxxers.The difference between the two proportions is reported as significantly high.
- Bots and disinformation: Bot prevalence is higher among misinformed anti-vaxxers than among misinformed pro-vaxxers, whose bot proportion is 17%.The authors report the difference as significant and use it to suggest that a substantial portion of misinformation may be intentional disinformation.
- Bots and disinformation: Bots in both informed and misinformed communities suggest that some disinformation may be an organized effort to amplify COVID-19 debate and create discord.The paper connects this possibility to earlier observations involving Twitter bots and Russian trolls.
6 Limitations
The paper identifies limitations involving annotation coverage, correlational analysis, temporal scope, prevalence estimation, and bot classification.
- Annotation: Most analyses rely on data annotated by only one annotator, although more than one-seventh of annotations received a second annotation.The authors incorporate the multiply annotated cases when computing user membership.
- Inference: The analyses are correlational and do not establish causation.
- Data scope: Because data were collected over three weeks and sampled for annotation, the dataset cannot assess change over time.The collection strategy supports topic and agent diversity but may limit estimates of the actual prevalence of different story types.
- Bot analysis: Bot analysis relies on second-level inference from a trained model.The authors attempt to mitigate this by using probability-based labels, though the supplied passage ends before specifying the full procedure.
7 Conclusion
The paper characterizes competing COVID-19 misinformation communities across network structure, sociolinguistic variation, and membership in related communities. It finds that misinformed communities are denser and highly organized, while informed users use more narratives and many misinformed users may be anti-vaxxers.
- The methodology compares competing COVID-19 misinformation communities by network structure, sociolinguistic variation, and membership in disinformation and other health-related misinformation communities.
- Misinformed communities are denser than informed communities and appear highly organized.
- Bots occur in both groups, but the significantly higher percentage among misinformed users suggests possible disinformation campaigns.
- Informed users use many more narratives than misinformed users, although both communities show negative emotional tone.
- Many misinformed users may be anti-vaxxers, indicating overlap with another health-related misinformation community.