Source-linked AI summary

Tracking Social Media Discourse About the COVID-19 Pandemic: Development of a Public Coronavirus Twitter Data Set

Emily Chen, Kristina Lerman, Emilio Ferrara

arXiv:2003.07372v2cs.SIq-bio.PE

TL;DR

The paper addresses the need for data on online discourse during the COVID-19 crisis by continuously collecting and sharing a multilingual Twitter dataset. Its initial analysis shows that discourse statistics reflect major pandemic-related events, while the dataset remains constrained by API coverage and English-oriented tracking.

  • Problem

    Social distancing shifted substantial COVID-19 conversation online, creating a need for research data on pandemic-related social media discourse.

  • Method

    The authors continuously collect COVID-19-related tweets using evolving English keywords and accounts, then share the resulting dataset through Tweet IDs.

  • Results

    Twitter discourse statistics reflect major pandemic-related events in the collected dataset.

  • Takeaways & Limitations

    The dataset is intended to support research on online conversation dynamics and social media’s role in the global health crisis.

  • Takeaways & Limitations

    The dataset covers only 1% of Twitter volume and is significantly biased toward English tweets because tracking keywords and accounts are mostly English.

Abstract

from arXiv · show

At the time of this writing, the novel coronavirus (COVID-19) pandemic outbreak has already put tremendous strain on many countries' citizens, resources and economies around the world. Social distancing measures, travel bans, self-quarantines, and business closures are changing the very fabric of societies worldwide. With people forced out of public spaces, much conversation about these phenomena now occurs online, e.g., on social media platforms like Twitter. In this paper, we describe a multilingual coronavirus (COVID-19) Twitter dataset that we have been continuously collecting since January 22, 2020. We are making our dataset available to the research community (https://github.com/echen102/COVID-19-TweetIDs). It is our hope that our contribution will enable the study of online conversation dynamics in the context of a planetary-scale epidemic outbreak of unprecedented proportions and implications. This dataset could also help track scientific coronavirus misinformation and unverified rumors, or enable the understanding of fear and panic -- and undoubtedly more. Ultimately, this dataset may contribute towards enabling informed solutions and prescribing targeted policy interventions to fight this global crisis.

Background:

This paper, by Chen, Lerman, and Ferrara, is titled “Tracking Social Media Discourse About the COVID-19 Pandemic: Development of a Public Coronavirus Twitter Data Set.”

  • The paper is authored by Chen, Lerman, and Ferrara.
  • It is titled “Tracking Social Media Discourse About the COVID-19 Pandemic: Development of a Public Coronavirus Twitter Data Set.”
  • The article was published in JMIR Public Health and Surveillance in 2020.

Introduction

The paper situates COVID-19 as a rapidly spreading global crisis that has shifted social interaction online, then introduces a public Twitter dataset for studying pandemic-related discourse.

  • COVID-19 spread internationally, prompting quarantines, emergency declarations, and government efforts to contain health and economic consequences.
  • Social distancing disrupted everyday activities and moved more social interaction and conversation onto platforms such as Twitter.
  • The authors share a COVID-19 Twitter dataset to support research on online social-network dynamics and social media’s role in the global health crisis.
  • Data collection began in real time in January 2020, using COVID-19-related keywords and accounts while documenting collection methods, statistics, and access procedures.

Methods

The collection combines Twitter’s streaming and search APIs with evolving keyword and account tracking to gather COVID-19-related tweets and make the dataset available for research use.

  • Streaming collection began January 28, 2020, while search over the same terms recovered historical tweets dating back to January 21, 2020.
  • The researchers continuously added keywords and accounts based on evolving Twitter conversations and monitored trending topics and COVID-19-related sources.
  • The tracked keywords were English, creating a heavy bias toward English tweets and events related to English-speaking countries.

Results

The releases provide Tweet IDs and release metadata for a continuously growing COVID-19 collection, while documenting recoverability limits, known gaps, and the scope of the latest release.

  • Releases: The dataset is released as Tweet IDs because Twitter’s terms prevent public release of collected tweet text; researchers can use the IDs to query Twitter’s API.
  • Releases: Known gaps reflect Twitter API restrictions, including limited recovery of tweets older than one week through the free streaming API.
  • Releases: The collection uses hourly Tweet ID files organized by posting year and month, with filenames encoding the tweet date and hour.
  • Releases: Researchers cannot obtain the original tweet when a tweet has been removed from Twitter.

Discussion

The dataset’s initial analyses show that Twitter discourse statistics track major pandemic events across hashtags, languages, and verified-user activity, while coverage remains constrained by API sampling and English-focused tracking. Despite these limitations, the collection captures over 1 million tweets daily and averages 35% non-English tweets.

  • Discussion: Twitter discourse statistics reflect major pandemic events, with a 3/2/2020 dip attributed to internet connectivity failures.The discussion analyzes release v1.2, covering January 21 through March 31, 2020.
  • Limitations: Coverage is limited because Twitter’s free stream API returns only 1% of total Twitter volume, while tracked keywords and accounts are mostly English.These constraints make the dataset significantly biased toward English tweets and dependent on the filter endpoint and network connection.
  • Limitations: The dataset collects over 1 million tweets daily from Twitter’s available 1% and contains an average of 35% non-English tweets.Collection began in late January and was planned to continue as the pandemic developed.
  • Languages: Japanese, Italian, and Spanish tweet activity increased around reported cases, deaths, or related events in their respective countries.Japanese activity rose after the Diamond Princess quarantine, Italian tweets spiked after early cases and a death, and Spanish activity increased after Spain’s first case and subsequent death reports.
  • Verified Users: Verified-user activity peaked during major events, including the emergence of the first COVID-19-related deaths in the United States.Verified accounts include news sources and political figures, which the paper describes as active during breaking news.

Conflicts of Interest

The paper provides an abbreviation for Coordinated Universal Time.

  • UTC means Coordinated Universal Time.
Loading 2003.07372v2…