Source-linked AI summary
CrisisMMD: Multimodal Twitter Datasets from Natural Disasters
Firoj Alam, Ferda Ofli, Muhammad Imran
TL;DR
Disaster-response research has largely focused on text, despite the usefulness of imagery and the lack of labeled multimodal data. The paper releases CrisisMMD, human-annotated Twitter datasets from seven natural disasters covering informativeness, humanitarian categories, and damage severity. Images contain more damage-related information than corresponding text, while image annotations also have a higher prevalence of not informative content.
Problem
Disaster-response research has focused mainly on text, while labeled imagery and combined textual-visual ground-truth data remain limited.
Method
The paper releases human-annotated multimodal Twitter datasets from seven natural disasters with three humanitarian annotation tasks.
Results
Images tend to contain more damage-related information than corresponding text, while the image informative task has a higher prevalence of not informative content than the text task.
Takeaways & Limitations
CrisisMMD provides multimodal ground-truth data for developing systems that reduce information overload and support humanitarian situational awareness.
Takeaways & Limitations
Images showing banners, logos, and cartoons are not considered informative.
Abstract
from arXiv · showhide
During natural and man-made disasters, people use social media platforms such as Twitter to post textual and multime- dia content to report updates about injured or dead people, infrastructure damage, and missing or found people among other information types. Studies have revealed that this on- line information, if processed timely and effectively, is ex- tremely useful for humanitarian organizations to gain situational awareness and plan relief operations. In addition to the analysis of textual content, recent studies have shown that imagery content on social media can boost disaster response significantly. Despite extensive research that mainly focuses on textual content to extract useful information, limited work has focused on the use of imagery content or the combination of both content types. One of the reasons is the lack of labeled imagery data in this domain. Therefore, in this paper, we aim to tackle this limitation by releasing a large multi-modal dataset collected from Twitter during different natural disasters. We provide three types of annotations, which are useful to address a number of crisis response and management tasks for different humanitarian organizations.
Introduction
Social media offers disaster information for humanitarian response, but extracting useful content requires handling overload and credibility across text and images. CrisisMMD addresses the shortage of labeled multimodal disaster data by releasing human-annotated Twitter datasets for humanitarian tasks.
- Social media posts provide reports about casualties, infrastructure damage, urgent needs, and missing or found people during disasters.
- Processing this information requires classification, clustering, summarization, credibility assessment, and prioritization to support preparedness, response, and recovery.
- Earlier disaster-response research emphasized textual content, although images can also help assess infrastructure damage and support humanitarian aid.
- CrisisMMD releases human-labeled multimodal Twitter datasets from seven natural disasters, addressing the lack of publicly available ground-truth imagery annotations.
- The annotation effort targets informativeness, humanitarian information categories, and image-based damage severity assessment.
Related Work
Existing crisis datasets support mainly textual analysis, while multimodal ground-truth resources combining text and visual annotations remained unavailable. CrisisMMD bridges this gap with multimodal datasets collected during seven natural disasters and annotated for several tasks.
- Crisis informatics has used social media to curate, analyze, and summarize crisis information for decisions and responses.
- CrisisLex provides labeled tweets from six disaster events using categories such as directly related, indirectly related, and not related.
- CrisisLex-related work also annotated tweet formativeness, information type, and source using crowdsourced workers.
- CrisisNLP provides crisis-related tweet resources, including approximately 2,000 Hurricane Sandy tweets and 4,400 Joplin Tornado tweets.
- CrisisMMD releases multimodal datasets from seven 2017 natural disasters with combined textual and visual annotations.
Natural Disaster Events and Data Collection
CrisisMMD collected Twitter data from seven natural disasters using event-specific keywords and hashtags. The dataset details table records each event’s collection keywords and period.
- Twitter data were collected during seven natural disasters using event-specific keywords and hashtags.
- Table 1 lists event names, keywords used for data collection, and data collection periods.
- Hurricane Irma data were collected from September 6 to September 19, 2017, yielding approximately 3.5 million tweets and 176,000 images.
Hurricane Harvey 2017
Hurricane Harvey was a Category 4 storm that struck Texas on August 25, 2017, while Hurricane Maria was a Category 5 hurricane affecting Dominica and Puerto Rico. Twitter collections covered both events during their respective disaster periods.
- Hurricane Harvey 2017: Hurricane Harvey was a Category 4 storm that hit Texas on August 25, 2017, causing nearly $200 billion in damage.
- Hurricane Harvey 2017: Harvey data collection ran from August 25 to September 5, 2017, producing approximately 7 million tweets and 300,000 images.
- Hurricane Maria 2017: Hurricane Maria was a Category 5 hurricane that struck Dominica and Puerto Rico and caused more than 78 deaths.
- Hurricane Maria 2017: Maria data collection ran from September 20 to October 3, 2017, producing approximately 3 million tweets and 52,000 images.
California Wildfires 2017
The California wildfires occurred in October 2017 and generated a dataset of approximately 400,000 tweets and 10,000 images.
- The wildfire series took place in California during October 2017.
- Approximately 400,000 tweets and 10,000 images were collected for this event.
Iraq-Iran Border Earthquake 2017
A magnitude-7.3 earthquake struck the Iran–Iraq border in November 2017, causing substantial casualties, homelessness, and injuries; the study collected approximately 200,000 tweets and 6,000 images.
- A magnitude-7.3 earthquake struck the Iran–Iraq border on November 12, 2017.
- The earthquake caused around 630 casualties, left 70,000 people homeless, and injured 8,000.
- Data collection from November 12 to November 19, 2017 produced approximately 200,000 tweets and 6,000 images.
Sri Lanka Floods 2017
Severe flooding in southwest Sri Lanka during May 2017 worsened after Cyclone Mora, and data collection yielded approximately 41,000 tweets and 2,000 images.
- Heavy monsoon rains caused severe flooding in southwest Sri Lanka in May 2017.
- Cyclone Mora worsened the flooding and caused additional floods and landslides during the last week of May.
- Data collection from May 31 to July 3, 2017 resulted in approximately 41,000 tweets and 2,000 images.
Data Filtering and Sampling
The dataset was prepared by retaining image-containing tweets, removing non-English and duplicate content, and sampling events according to annotation budgets and remaining dataset size.
- Tweets without at least one image URL were discarded, and image URLs were extracted from each tweet’s extended entities.
- Non-English tweets were removed using Twitter-provided language metadata.
- Table 2 reports initial tweets, associated images, and retained tweets for each dataset; image totals can exceed tweet totals because tweets may contain multiple images.
- Tweets with cosine similarity scores greater than 0.7 were treated as duplicates and removed.
- Approximately 4,000 tweets were sampled for Hurricanes Irma, Harvey, and Maria, while all filtered tweets were retained for lower-volume datasets.
Humanitarian Tasks and Manual Annotations
The dataset uses three humanitarian annotation tasks to filter informative content, categorize actionable information, and assess infrastructure damage severity. Manual labels were collected separately for tweets and images, with later tasks applied to progressively narrower subsets.
- Informative Task: Task 1 labels tweets and images as “Informative” or “Not informative” to reduce information overload during disasters.Annotators could also select “Don’t know or can’t judge” for non-English tweets or low-quality images.
- Damage Severity Assessment: Task 3 assesses physical damage severity in images, focusing on infrastructure such as buildings, bridges, and roads rather than non-physical indicators like smoke.Severe damage includes infrastructure that is unsafe, unusable, non-crossable, or non-driveable.
- Annotation Pipeline: The annotation pipeline applies Task 2 only to content labeled informative in Task 1, then applies Task 3 only to images labeled “Infrastructure and utility damage” in Task 2.Tweets with neither informative text nor imagery were dropped.
- Manual Annotations: Manual annotations were collected through Figure Eight, using test questions to screen annotators and agreement among three annotators to determine final labels.Tweets and images were annotated separately for each task.
- Annotation Results: 25%–35% of tweet text was labeled “not informative” across events, rising to around 60% for Sri Lanka floods; images had a higher not-informative proportion than text.These informative-task results describe the filtering stage before humanitarian-category annotation.
Applications and Future Directions
CrisisMMD supports multimodal research and humanitarian applications by linking tweet text and images with structured annotations. Its applications include cross-modal retrieval, image captioning, multimedia event summarization, and prioritization of actionable disaster information.
- Applications: CrisisMMD enables multimodal tasks such as joint text-image embeddings, cross-modal retrieval, image captioning, and multimedia event summarization.These applications use aligned and structured tweet-image data.
- Humanitarian applications: The dataset is intended to reduce the burden of noisy social-media streams by supporting identification of useful information for relief operations.Humanitarian organizations seek useful messages rather than a deluge of irrelevant or personal content.
- Humanitarian applications: Humanitarian organizations can use the annotations to identify actionable reports, including injuries, deaths, infrastructure damage, and rescue demands.Different organizations have different information needs, so fine-grained annotations support varied response priorities.
- Humanitarian applications: Damage-severity annotations can help response organizations focus attention on severely damaged infrastructure.The paper connects this prioritization to reducing suffering among affected people.
- Dataset contribution: CrisisMMD addresses the scarcity of labeled imagery data by providing several thousand manually annotated tweet-image pairs.The authors claim it is the first and largest multimodal dataset of its kind at publication.
Conclusions
CrisisMMD is a multimodal Twitter corpus created to address the lack of labeled imagery data in disaster research. It contains manually annotated tweets and images from seven major natural disasters, with annotations for informativeness, humanitarian categories, and damage severity.
- Dataset: CrisisMMD contains several thousand manually annotated tweets and images collected during seven major natural disasters in 2017.The disasters included earthquakes, hurricanes, wildfires, and floods across different parts of the world.
- Annotations: The dataset provides annotations for informative versus non-informative content, humanitarian categories, and damage severity.These annotations target crisis response and management tasks for humanitarian organizations.
- Applications: The paper presents humanitarian use cases and tasks that can be pursued with more robust and effective computational systems.The stated applications build on the multimodal dataset and its annotations.