Source-linked AI summary
Framing COVID-19: How we conceptualize and discuss the pandemic on Twitter
Philipp Wicke, Marianna M. Bolognesi
TL;DR
The paper asks how Covid-19 discourse is framed on Twitter, especially whether non-expert users employ WAR and alternative frames. Using topic modelling and lexical corpus analyses of pandemic-related tweets, it finds that WAR is concentrated in treatment and diagnostic discussions and is the most frequent figurative frame, while FAMILY covers more of the corpus.
Problem
The paper investigates which topics dominate Covid-19 Twitter discourse and whether the conventional WAR frame and alternative figurative frames are used across them.
Method
The authors apply topic modelling and frame-specific lexical analyses to Covid-19 tweets, comparing WAR with MONSTER, STORM, TSUNAMI, and FAMILY.
Results
WAR-related terms occur in 5.32% of tweets and are especially associated with virus treatment and diagnostics, while FAMILY covers a wider corpus portion and WAR is the most frequent figurative frame.
Takeaways & Limitations
Because WAR fits some pandemic aspects but not others, the authors support using a broader metaphor menu to represent varied Covid-19 experiences.
Takeaways & Limitations
Filtering repeated tweets and retweets improves corpus balance but omits Twitter’s retweeting, super-tweeter, and popularity dynamics.
Abstract
from arXiv · showhide
Doctors and nurses in these weeks are busy in the trenches, fighting against a new invisible enemy: Covid-19. Cities are locked down and civilians are besieged in their own homes, to prevent the spreading of the virus. War-related terminology is commonly used to frame the discourse around epidemics and diseases. Arguably the discourse around the current epidemic will make use of war-related metaphors too,not only in public discourse and the media, but also in the tweets written by non-experts of mass communication. We hereby present an analysis of the discourse around #Covid-19, based on a corpus of 200k tweets posted on Twitter during March and April 2020. Using topic modelling we first analyze the topics around which the discourse can be classified. Then, we show that the WAR framing is used to talk about specific topics, such as the virus treatment, but not others, such as the effects of social distancing on the population. We then measure and compare the popularity of the WAR frame to three alternative figurative frames (MONSTER, STORM and TSUNAMI) and a literal frame used as control (FAMILY). The results show that while the FAMILY literal frame covers a wider portion of the corpus, among the figurative framings WAR is the most frequently used, and thus arguably the most conventional one. However, we conclude, this frame is not apt to elaborate the discourse around many aspects involved in the current situation. Therefore, we conclude, in line with previous suggestions, a plethora of framing options, or a metaphor menu, may facilitate the communication of various aspects involved in the Covid-19-related discourse on the social media, and thus support civilians in the expression of their feelings, opinions and ideas during the current pandemic.
Introduction
The paper examines how non-expert Twitter users frame Covid-19 discourse during a rapidly spreading pandemic and widespread lockdowns. It asks which topics dominate this discourse and whether war and other figurative frames characterize it.
- Covid-19 spread rapidly worldwide, while lockdowns, school closures, remote work, and home confinement reshaped daily life.
- Twitter users used social media to express concerns, opinions, beliefs, and feelings during enforced social distancing.
- The study investigates Covid-19 topics on Twitter and the extent to which non-expert communicators use WAR and other figurative frames.
- The paper explicitly addresses which topics are discussed and whether WAR and alternative figurative frames are used to discuss Covid-19.
1. What type of topics are discussed on Twitter, in relation to Covid-19?
The paper situates its analysis in research on social-media health discourse and framing, especially the conventional metaphor that disease treatment is war. It also considers criticisms of war framing and alternative ways to conceptualize the pandemic.
- Related research: Prior Twitter studies used social-media posts to examine epidemic incidence, public concerns, attitudes, reactions, symptoms, transmission, prevention, and treatment.
- Framing and metaphor: Framing selects salient aspects of reality to promote problem definitions, causal interpretations, moral evaluations, or treatment recommendations.
- WAR framing: The conventional DISEASE TREATMENT IS WAR metaphor maps medical professionals to allies, the body to a battlefield, tools to weapons, and treatment to fighting.
- WAR framing: War metaphors are widely used because they provide a familiar framework for abstract topics and convey urgency around negative situations.
- Critiques and alternatives: Critics argue that pandemic war language can narrow perceived solutions and increase xenophobia, fear, and anxiety.
- Critiques and alternatives: The study compares WAR with alternative frames including FOOTBALL, GAMES, STORMS, MONSTER, and TSUNAMI, alongside the literal FAMILY frame.
Study design
The study combines topic modelling with corpus-based lexical analyses to examine Covid-19 discourse, WAR framing, and comparisons with alternative figurative and literal frames.
- The authors first use topic modelling to identify topics in Covid-19 Twitter discourse.
- They then compile war-related lexical units and examine their distribution across the identified topics.
- Finally, they compare the corpus percentages of WAR, three figurative alternatives, and one literal frame.
- The analyses are replicated on a later tweet corpus and the Coronavirus Tweets Dataset.
Constructing the corpus of Covid-19 tweets
The corpus contains English Covid-19-related tweets collected through predefined hashtags and filtered to reduce repeated contributions from highly active users. The resulting dataset contains 203,756 tweets from unique tweeters, with redistribution restricted to tweet IDs.
- Tweets were collected through Twitter’s official API using predefined Covid-19 hashtags, with 25,000 tweets targeted per day and retweets excluded.
- 203,756 tweets from unique tweeters were collected over 14 days from 20.03.2020 to 02.04.2020.
- The corpus mainly represents English-language users residing in the USA because of collection timing and language restrictions.
- The study excludes retweeting, mentions, usernames, hashtags, and URLs from its analysis.
- The dataset is stored and publicly distributed as tweet IDs to comply with Twitter’s privacy and content-redistribution policies.
- The corpus’s word-frequency analysis and word cloud summarize the most common words after excluding stopwords and online tags.
What type of topics are discussed on Twitter, in relation to Covid19?
The study uses topic modelling to identify semantically related categories in Covid-19 tweets, comparing broad and fine-grained topic divisions after preprocessing the corpus.
- Topic modelling: LDA identifies topics as heterogeneous categories based on semantically related word occurrences in documents.The model is used to discover categories within the Covid-19 tweet corpus.
- Topic modelling: The analysis compares four broad topics with sixteen more fine-grained topics.The four-topic model represents a less granular division, while the sixteen-topic model represents a more granular division.
- Preprocessing: Tweets were tokenized, filtered for short tokens and stopwords, stripped of Covid-19 terms, and converted into a bag-of-words.These preprocessing steps produced the representation used for topic modelling.
- Preprocessing: Covid-19 terms were removed because they do not provide information about the topics themselves.The corpus retained inflected word forms rather than lemmatizing or part-of-speech tagging them, since different forms can express different metaphor scenarios.
Topic model analysis
The four-topic LDA model reveals broad, overlapping thematic classes, while the sixteen-topic model provides substantially greater diversity among topic classes.
- Four-topic model: The N = 4 LDA model assigns words and importance weights to four topics, visualized as word clouds.Larger words indicate greater significance, and “pandemic” appears among the most important words in all but topic II.
- Sixteen-topic model: The sixteen-topic results show much greater diversity among the classes.This fine-grained model distinguishes more varied topic groupings than the four-topic analysis.
Discussion
The discussion interprets the LDA topics and examines how WAR-related language is distributed across them, while acknowledging limits in automatically identifying metaphorical usage.
- Topic interpretation: Analysts interpret the LDA topics because the algorithm does not provide topic labels.The four broad topics include Communications and Reporting, Community and Social Compassion, and Reacting to the epidemic.
- Topic interpretation: The sixteen-topic model refines broad domains into distinctions involving world news, lockdown and media, quarantine, treatment, testing, and working or studying from home.Topics #1, #6 and #7 concern treatment and medical needs; #10 concerns testing; and #2, #9 and parts of #12 concern working or studying from home.
- WAR-frame analysis: Because manual metaphor-identification procedures cannot be applied to the large tweet corpus, the study assumes war-related entries are metaphorical and checks this qualitatively on a subsample.The authors acknowledge that some unexamined tweets may use these terms literally.
- WAR-frame analysis: The study identifies WAR-related topics by collecting tweets containing WAR terms and using the LDA model to predict their topics.This reveals which topics contain the most or fewest WAR-related terms.
WAR framing results
WAR language appears in a measurable minority of Covid-19 tweets, but many specialized war terms are virtually absent from the corpus.
- WAR framing: 10,846 tweets, or 5.32% of the corpus, contained at least one WAR-framing term.Among these, 1,253 tweets contained more than one war-related term.
- WAR framing: Combatant, combative, disarmament, gunfight, invader, treaty, bombard, minefield, belligerent, guerilla, insurgency, vanquish, conquest, blitzkrieg, and vanquishment were virtually absent or rarely used.Several terms occurred only once or twice, while conquest and blitzkrieg did not occur.
LDA topic prediction of WAR tweets
WAR-related tweets cluster in selected Covid-19 topics, especially treatment, diagnostics, responses, communications, and politics, while intimate social topics are less associated with the frame. The analysis reports that 5.32% of tweets contain war-related terms, while noting methodological randomness in LDA topic distributions.
- Fine-grained topic distribution: In the sixteen-topic model, WAR terms are particularly represented in topics 2, 7, and 10.Topic 2 concerns online learning and education, while topics 7 and 10 concern virus treatment, diagnostics, and related support language.
- 5.32% of all tweets contain war-related terms and are therefore likely to frame Covid-19 metaphorically as a literal war.
- WAR-frame vocabulary: The WAR frame uses terms such as “fight,” “war,” “combat,” “threat,” and “battle,” which denote negative actions and events.The authors suggest this vocabulary may reflect the emergency stage of the pandemic and could change in later phases.
- Macro-level topic distribution: WAR-related tweets are most likely to belong to macro-topics IV, I, and III rather than topic II.Topic IV concerns epidemic responses; topics I and III concern communications, reports, and politics, whereas topic II covers familiar, communal, and compassionate discourse.
- Fine-grained topic distribution: WAR framing is associated with disease treatment and diagnostics, but not with intimate social relations and personal affective aspects.Tweets about topic 3, characterized by friends, family, sharing, and personal time, do not employ military lexical units.
- Methodological caveat: LDA topic distributions vary because the model uses randomness in training and inference.Training a new model with the same parameters can yield slightly different topic distributions, and topic-model analysis has additional limitations.
The literal frame of FAMILY used as control
The study compares figurative framings with FAMILY as a literal control frame. FAMILY is represented through lexical entries concerning kinship, households, relatives, and domestic relations.
- FAMILY serves as the literal frame used to evaluate the relevance of figurative frames in the tweet corpus.
- The FAMILY lexical list contains 66 entries spanning marriage, household, kinship, relatives, children, and other family relations.Examples include marriage, household, parent, cousin, child, sibling, spouse, and related terms.
Alternative framing results
The literal FAMILY frame covers substantially more tweets than the figurative alternatives, while WAR is the most frequent figurative frame. Frame frequencies differ significantly, and comparable list sizes yield similar corpus coverage.
- Frame frequencies: 12.06% of tweets contain FAMILY terms, compared with 1.49% for STORM, 1.13% for TSUNAMI, and 0.68% for MONSTER.The frame-frequency differences are statistically significant (Cochran’s Q = 47,226.72, df = 4, p < 0.001).
- Lexical distributions: Within each frame, term frequencies follow Zipf distributions, with a few highly frequent words and many rarely used words.The WAR term “fight” occurs more than 3,000 times in the corpus.
- Comparison design: The WAR frame contains more lexical units than the other figurative frames, requiring equal-length lists for frequency comparisons.The study compares lists of 30 and 50 lexical units for each frame.
- Frame frequencies: FAMILY is substantially more frequent than the figurative frames, while WAR covers more tweets than the other figurative alternatives.The comparison uses frame term lists with matched lengths to avoid favoring frames with more lexical units.
Replication studies
Replication analyses across additional and external corpora produce similar frame distributions and preserve the ordering FAMILY > WAR > STORM + TSUNAMI + MONSTER. Differences are associated with corpus keywords and the evolving pandemic discourse.
- Replication findings: The five-frame distribution is very similar across the first corpus and the replication studies.The replications compare a subsequent two-week corpus and an external dataset containing more than 1.2 million tweets.
- Replication findings: >0.22% increase in WAR framing from the two-week to the two-month corpus may reflect new debates emerging as the epidemic developed.The authors present this explanation as partial and tentative.
- External corpus comparison: The relative ordering remains FAMILY > WAR > STORM + TSUNAMI + MONSTER in the comparison with Lamsal’s Coronavirus Tweets Dataset.In that comparison, FAMILY decreases by 3.46% and WAR increases by 1.62%.
- External corpus comparison: Differences between datasets may result from keyword choices, including Lamsal’s use of “Corona” and changing keywords during data mining.The authors retained the same keyword set throughout their own collection.
- Frame interpretation: Qualitative inspection suggests that different frames address different aspects of the pandemic, with MONSTER terms personifying the virus through emotionally negative words.Examples include “devil,” “demon,” “horror,” “monster,” and “killer.”
- Replication findings: The replication analyses are consistent across corpora, although differences between time spans may reflect the pandemic’s changing discourse.The authors distinguish temporal variation from dataset differences associated with keywords.
General discussion and conclusion
The study identifies major Covid-19 Twitter topics and finds that WAR framing is concentrated in treatment and diagnostics rather than applying uniformly across the discourse. Its conclusions are bounded by corpus construction choices and support using multiple frames for different pandemic aspects.
- Findings: The main Twitter topics concern Communications and Reporting, Community and Social Compassion, Politics, and Reacting to the epidemic.Finer-grained topics include disease treatment, healthcare workers, and virus diagnostics.
- Findings: WAR lexical units are particularly concentrated in tweets about virus treatment and diagnostics.Frequently used war terms include “fighting,” “fight,” “battle,” and “combat,” while most war-related words are not used to frame Covid-19 discourse.
- Methodological scope: The study uses topic modelling and automated lexical analysis on a corpus constructed by dropping retweets and retaining one tweet per user.These choices were intended to reduce bias from duplicated content and super-tweeters.
- Methodological scope: Because retweets and super-tweeters were excluded, the findings may not represent Twitter’s distribution as a social network per se.The authors instead characterize the findings as reflecting how a broad selection of American-English speakers conceptualize and discuss Covid-19 on Twitter.
- Conclusion: WAR is relatively pervasive in Covid-19 discourse but is not typically used for aspects such as family closeness during distancing or collaborative efforts to flatten the curve.The authors characterize WAR as apt for treatment and hospital operations but not for every pandemic-related concern.