Source-linked AI summary

A first look at COVID-19 information and misinformation sharing on Twitter

Lisa Singh, Shweta Bansal, Leticia Bode, Ceren Budak, Guangqing Chi, Kornraphop Kawintiranon, Colton Padden, Rebecca Vanarsdall, Emily Vraga, Yanchen Wang

arXiv:2003.13907v1cs.SI

TL;DR

The paper addresses limited early evidence about how COVID-19 information and misinformation circulated on Twitter. It analyzes conversation volume, themes, locations, myths, shared URLs, and relationships with reported cases. Preliminary results show Twitter conversations led COVID-19 cases by 2-5 days in the study's analysis, while myths and misinformation were present but less dominant than other themes.

  • Problem

    Health misinformation is widespread on social media, and the COVID-19 pandemic created an urgent need to understand its rapidly expanding Twitter conversation.

  • Method

    The study analyzes Twitter conversation volume, languages, themes, locations, myths, and shared URLs, using weighted phrase matching to identify five common myths and NewsGuard's curated list to identify low-quality misinformation sources.

  • Results

    Twitter conversations were highly correlated with COVID-19 cases and led them by 2-5 days; misinformation-linked sources were shared less often than credible health sources but retweeted more often.

  • Takeaways & Limitations

    The findings provide preliminary support for using Twitter conversations as a surveillance approach when reliable leading indicators are unavailable.

  • Takeaways & Limitations

    The analysis used simple text methods and its non-representative population creates measurement-error concerns requiring future bias measurement and reweighting.

Abstract

from arXiv · show

Since December 2019, COVID-19 has been spreading rapidly across the world. Not surprisingly, conversation about COVID-19 is also increasing. This article is a first look at the amount of conversation taking place on social media, specifically Twitter, with respect to COVID-19, the themes of discussion, where the discussion is emerging from, myths shared about the virus, and how much of it is connected to other high and low quality information on the Internet through shared URL links. Our preliminary findings suggest that a meaningful spatio-temporal relationship exists between information flow and new cases of COVID-19, and while discussions about myths and links to poor quality information exist, their presence is less dominant than other crisis specific themes. This research is a first step toward understanding social media conversation about COVID-19.

1. Introduction

The paper examines COVID-19 discussion on Twitter because social media is a major channel for health information while also carrying substantial misinformation. It provides an initial analysis of conversation volume, themes, geographic emergence, myths, information quality, and relationships with reported cases.

  • Motivation: Social media is a major conduit for news and health information, making Twitter an important setting for studying COVID-19 communication.The paper notes that Twitter users commonly share news and preventive health information.
  • Motivation: Health misinformation is widespread across social media and includes content about infectious-disease outbreaks such as Ebola and Zika.The paper defines health misinformation as information that counters the best available evidence from medical experts at the time.
  • Motivation: COVID-19 warrants urgent study because more people use social media than during earlier epidemics and institutional trust is eroding alongside increased online misinformation.Twitter users increased from 255 million in February 2014 to more than 330 million in 2019, according to the paper.
  • Scope: The article measures COVID-19 conversation volume, themes, geographic emergence, shared high- and low-quality URLs, and five specific myths on Twitter.The authors characterize these social-media signals as moderate or poor quality indicators that require calibration.
  • Findings: Preliminary findings show growing conversation, a 2-5-day lead of information flow over cases in some countries, dominant health and global-pandemic themes, and lower-volume misinformation and myth discussion.The authors present these findings as preliminary support for using Twitter as a crisis-surveillance approach.

2. Size of COVID-19 Conversation

From January 16 to March 15, 2020, COVID-19 Twitter conversation increased overall and varied across languages. English dominated the collection, while temporal patterns differed by language and reflected the changing geography of outbreaks.

  • Data scope and caveat: The study used Twitter Streaming API data collected with COVID-19 hashtags, but a collection glitch made some hashtags unavailable from March 13 to March 15.The paper warns that the apparent drop in volume during those dates is an artifact.
  • Overall conversation volume: 2,792,513 tweets, 456,878 quotes, and 18,168,161 retweets were collected from January 16 through March 15, 2020.Tweet and retweet volumes increased as the epidemic unfolded, with peaks mapping to a series of crisis events rather than one single event.
  • Overall conversation volume: Late-January conversation ramped up around January 25 after new cases, Hubei quarantine measures, and emergency actions in Hong Kong and the United States.A later increase occurred at the end of February amid rising deaths in Iran and Italy and restrictions in Switzerland.
  • Conversation by language: 57.1% of COVID-19 tweets were in English, followed by Spanish at 11.6%, French at 6.5%, and Italian at 4.8%.The COVID-19 language distribution differed from overall Twitter language use, despite English and Spanish remaining generally dominant.
  • Conversation by language: Chinese conversation was higher in January and early February, while English, French, German, Italian, Spanish, and Turkish generally rose again in March.Spanish had the highest increase among the highlighted languages, whereas Japanese stayed fairly constant after its initial spike and Thai showed large January and March spikes.

3. Location of Conversation and Its Relationship to COVID-19 Cases

The study compares Twitter location signals with reported COVID-19 cases using conversation-location tweets, geotagged tweets, and cross-correlation analysis. Conversation volume generally tracks case counts and appears more closely aligned when it precedes reported cases.

  • 3.1. Description of Data Sets: Conversation-location tweets assign a tweet to locations mentioned alongside COVID-19, while geotagged tweets identify where users sent tweets.The location ontology includes countries, governorates, capitals, and major cities; geotagged tweets were collected in English.
  • 3.2. Locations of Twitter conversation: Conversation-location tweets and confirmed cases were highly correlated across countries, with more cases generally accompanying more location-related conversation.The analysis covered 217 mentioned countries, 184 of which were mentioned at least 10 times.
  • 3.2. Locations of Twitter conversation: The correlation between conversation-location and geotagged tweet volumes was 0.75 when China was excluded.The authors describe this as a reasonable, though imperfect, proxy relationship.
  • 3.3. Comparison between location conversation and COVID-19: Cross-correlation analysis compares location-conversation and case time series across different leads and lags, including separate China analyses because of a testing-procedure spike.A lead means tweets occur before cases; a lag means tweets occur after cases.
  • 3.3. Comparison between location conversation and COVID-19: Conversations were more highly correlated with cases at leads than at lags, suggesting tweets may precede confirmed-case increases.The authors relate this pattern to delays between symptom onset, severe symptoms, and testing, while noting that testing protocols may differ by country.

4. Content of English Conversation

The paper analyzes frequent words and researcher-coded themes in English COVID-19 tweets, then examines how theme prevalence changes over time. Health or virus discussion and the pandemic’s global nature dominate the labeled conversation.

  • 4.1. Words Being Used: The analysis first examines frequent words, then uses a word cloud to identify dominant references to the virus, its spread, responses, and the pandemic’s global nature.Words in the cloud appear in at least 150,000 tweets and are sized by frequency.
  • 4.2.1. Identifying Themes: Researchers grouped 537 frequent words into eight themes: Economy, Emotion, Illness, Global Nature, Information Providers, Social, Government Response, and Individual Response.Three researchers performed the grouping through open coding.
  • 4.2.1. Identifying Themes: Approximately 80% of tweets received one or more theme labels through proportional word-to-theme assignment.Theme proportions were summed by day to estimate overall theme volume.
  • 4.2.2. Findings: Health or virus content accounted for 30% of labeled tweets, while the global-nature theme accounted for 29%.The health theme includes the virus, health consequences, vaccines, testing, and other epidemics; the global theme includes locations and pandemic scale.
  • 4.2.2. Findings: Information providers represented 11%, emotion 9%, and economy 3% of labeled tweets during the study period.The authors note that economy prevalence might change as secondary effects become more prominent, while the emotion share highlights mental-health needs.
  • 4.2.3. Themes Over Time: The top five themes followed trends similar to overall COVID-19 hashtag volume, with global-nature and health/virus themes most common early in the period.Early discussion often referred to China, the virus, and reported cases.

5. Myths About the Virus

The study identifies and measures five prominent COVID-19 myths on Twitter, finding that myth-related tweets increased over time but remained a small share of the conversation.

  • Methodology: The analysis grouped common examples into ten initially identified myths, then focused on five accurately detected categories: virus origin, vaccine development, flu comparison, heat killing disease, and home remedies.Myths were identified from blogs, news media, and medical organizations, then categorized using weighted words and phrases in tweets.
  • Findings: Approximately 16,000 tweets, just under 0.6% of the collection, discussed one or more of the five myths.The distribution of discussion across individual myths is presented in Figure 11.
  • Findings: Myth-associated tweet volume increased since January, although this may reflect growing COVID-19 discussion rather than a larger proportion of attention devoted to myths.Figure 12 addresses the changing share of attention over time.
  • Limitations: The detection approach prioritized precision over recall; manual validation produced 80% precision for myth categories and 100% precision for non-myth classifications.The authors note that some noisier tweets were missed.
  • Findings: The virus-origin myth dominated in January and February, while flu-comparison and home-remedy myths appeared at nearly equal frequencies by late February.These trends describe differences among the sampled myths over time.

6. Sharing of High vs Low Quality Information on Twitter

The study examines URLs and linked sources in COVID-19 tweets, finding that high- and low-quality health sources were both uncommon, while news-linked content more often connected to high-quality sources.

  • URL Sharing: 40.5% of original tweet content included a URL, compared with 5.1% of retweet content and 9.6% of overall content.The authors associate the high URL share with information seeking during uncertainty.
  • URL Sharing: The researchers analyzed more than 60,000 unique domains and examined the most frequently shared domains, including YouTube, news organizations, and retail sites.The top-ten domain analysis required domains to be tweeted by over 100 user accounts.
  • Health Sources: The study identified reputable health domains using public-health agencies, medical journals, hospitals, and related sources, and low-quality sources using NewsGuard’s COVID-19 misinformation list.These source sets were used to classify shared links as high quality or low quality/questionable.
  • Health Sources: Low-quality/questionable sources accounted for 0.4% of original tweets and 0.06% of retweets, while reputable health sources accounted for 0.51% and 0.04%, respectively.Low-quality sources were tweeted less often but retweeted at a higher rate than reputable health sources; both remained a small fraction of conversation.
  • News Sources: Just over 351,000 tweets linked to news organizations; 18% of those news shares connected to high-quality health sources, while less than 0.3% connected to low-quality sources.Among frequently shared news domains with linkable articles, 175 of 178 referred to high-quality sources at least 80% of the time.
  • News Sources: The news-source analysis applies only to articles that link to other sources in the researchers’ Twitter dataset.This scope boundary limits what can be inferred about articles without such links.

7. Discussion and conclusions

The study links Twitter conversation with COVID-19 case patterns and examines information quality, myths, and communication during the pandemic. It finds that misinformation and myths were present but less prominent than broader crisis-related discussion, while noting methodological and measurement limitations.

  • 7. Discussion and conclusions: Twitter conversations led COVID-19 cases by 2-5 days, suggesting potential use for predicting spread when reliable leading indicators are unavailable.The authors caution that measurement error may arise from the non-representative population and that bias adjustment remains future work.
  • 7. Discussion and conclusions: Twitter attention to COVID-19 continued to grow, with discussion concentrated in countries most affected by the pandemic.
  • 7. Discussion and conclusions: 40.5% of original tweets included a URL, but only 0.4% linked directly to highly credible health sources such as the CDC or WHO.Known misinformation sources were also not shared in great numbers, though their links were retweeted more often than links to credible health sources.
  • 7. Discussion and conclusions: Over 350,000 tweets linked to news sources, whose articles more often referenced credible sources than misinformation sources.The cited counts were 63,352 news articles linking to sources such as WHO and about 1,135 linking to misinformation sources.
  • 7. Discussion and conclusions: Five of ten predefined myth themes were identified with high precision, and tweets concerning them accounted for a small fraction of Twitter content.The authors emphasize that additional myths and distinctions between propagators and debunkers require continued monitoring.
  • 7. Discussion and conclusions: The analysis used simple text-analysis methods because of data volume and time constraints, with more robust techniques reserved for future work.Future directions also include refined spatio-temporal analysis and language-model-assisted myth and theme identification.

8. Appendix A

Appendix A provides a table of hashtags and their collection start dates.

  • 8. Appendix A: Appendix A contains Table 6, titled “Hashtags and Start Date of Collection.”
Loading 2003.13907v1…