Source-linked AI summary
Assessing the risks of "infodemics" in response to COVID-19 epidemics
Riccardo Gallotti, Francesco Valle, Nicola Castaldo, Pierluigi Sacco, Manlio De Domenico
TL;DR
Infodemics can accompany COVID-19 emergencies, but evidence on how unreliable information develops and relates to epidemic dynamics remains limited. This paper analyzes worldwide multilingual Twitter messages and finds that unreliable information waves precede epidemic waves, followed by a shift toward reliable sources.
Problem
The paper examines the risks posed by low-quality information during global epidemic emergencies and the relationship between infodemics and biological epidemics.
Method
The study analyzes worldwide COVID-19 tweets and classifies disseminated news by source reliability using a Harm Score framework.
Results
Waves of unreliable and low-quality information precede epidemic waves, while reliable information becomes prominent when epidemics reach the same area.
Takeaways & Limitations
The findings identify early-warning signals in responses to falsehood that may inform adequate communication strategies.
Takeaways & Limitations
The analysis does not track higher-order transmission pathways among exposed users.
Abstract
from arXiv · showhide
Our society is built on a complex web of interdependencies whose effects become manifest during extraordinary events such as the COVID-19 pandemic, with shocks in one system propagating to the others to an exceptional extent. We analyzed more than 100 millions Twitter messages posted worldwide in 64 languages during the epidemic emergency due to SARS-CoV-2 and classified the reliability of news diffused. We found that waves of unreliable and low-quality information anticipate the epidemic ones, exposing entire countries to irrational social behavior and serious threats for public health. When the epidemics hit the same area, reliable information is quickly inoculated, like antibodies, and the system shifts focus towards certified informational sources. Contrary to mainstream beliefs, we show that human response to falsehood exhibits early-warning signals that might be mitigated with adequate communication strategies.
Methods
The study collected COVID-19-related Twitter messages worldwide using predefined hashtags and keywords, then classified users and linked domains with machine-learning and manually checked databases. Twitter’s Stream API provided uninterrupted qualifying messages when their volume remained below 1% of all platform messages.
- Data collection: The study monitored COVID-19 tweets using predefined virus- and disease-related hashtags and keywords through Twitter’s public API.The tracked terms included coronavirus, ncov, #Wuhan, covid19, covid-19, sarscov2, and covid.
- Data collection: From 20 January 2020, the collection covered qualifying tweets continuously and regardless of language, subject to the Stream API’s 1% volume limit.The observation began when China reported more than 6,000 cases; above 1% of total unfiltered platform volume, the API condition was not satisfied.
- User classification: Users were classified as humans or bots with a deep-learning algorithm reporting >90% accuracy and >95% precision for bot identification.The model was described as more stable for broadcaster users than state-of-the-art comparison methods.
- Domain classification: Shared web domains were manually checked using multiple public databases, including scientific and journalistic sources.Because databases used different labeling schemes, the study defined reliable domains as scientifically scrutinized or professionally fact-checked and accountable.
MAINSTREAM MEDIA
The MAINSTREAM MEDIA classification identifies several types of unreliable domains, including satire, clickbait, political content, and fake/hoax material. It also includes unknown categories for unclassifiable content and unresolved shortened URLs.
- SATIRE 3 UNRELIABLE domains intentionally distort events as humor or social critique.
- CLICKBAIT 4 UNRELIABLE domains distort or misrepresent information to capture attention.
- OTHER 5 UNKNOWN covers content that cannot be easily classified, while SHADOW 6 UNKNOWN covers shortened URLs that cannot be successfully unshortened.
- POLITICAL 7 UNRELIABLE domains present partisan interpretations of facts to support one political position over rivals.
- FAKE/HOAX 8 UNRELIABLE domains provide manipulative, fabricated content intended to mislead public opinion and provoke inflammatory responses.
CONSPIRACY/JUNK SCI
The section defines conspiracy/junk science as manipulative, fabricated content that imitates scientific reasoning to legitimize implausible explanations. It classifies sources by intentional knowledge manipulation and data fabrication, with additional filtering producing a final database of 3,892 domains.
- Conspiracy/junk science comprises manipulative and fabricated content that coarsely mimics scientific reasoning to legitimize implausible facts and knowledge.
- 3,892 entries remained after removing hard duplicates, resolving conflicting classifications, and filtering poorly defined or unclassifiable domains.The database initially contained 4,988 domains and was reduced to 4,417 after hard-duplicate removal.
- The Harm Score ranks sources higher when knowledge manipulation and data fabrication are more systematic and intentionally harmful.Scientific content receives the lowest Harm Score because it undergoes rigorous validation through scientific methods.
- Other and Shadow categories include content associated with social manipulation, while anonymized or temporary links in Shadow add unaccountability and increase Harm Score.