Source-linked AI summary

xBD: A Dataset for Assessing Building Damage from Satellite Imagery

Ritwik Gupta, Richard Hosfelt, Sandra Sajeev, Nirav Patel, Bryce Goodman, Jigar Doshi, Eric Heim, Howie Choset, Matthew Gaston

arXiv:1911.09296v1cs.CV

TL;DR

Assessing building damage for disaster response is dangerous and time-consuming, while damage is not adequately represented by binary labels. xBD assembles a large, expert-curated corpus of imagery and building annotations to support remote, automated damage assessment across disaster settings.

  • Problem

    Building-damage assessment requires dangerous, time-consuming ground observation, and binary labels do not represent the continuum of damage needed for disaster response.

  • Method

    xBD combines imagery with building polygons, an expert-created damage annotation scale, and a rigorous, repeatable annotation process with expert quality control.

  • Results

    Over 800,000 building annotations across over 45,000 km2 of imagery form a large corpus spanning visually distinct regions and varied disaster types.

  • Takeaways & Limitations

    The corpus supports development of vision models that automate building-damage assessment remotely for disaster relief operations.

  • Takeaways & Limitations

    The Joint Damage Scale is a best-effort trade-off that cannot match the precision of in-person human assessment because satellite imagery has modality limitations.

Abstract

from arXiv · show

We present xBD, a new, large-scale dataset for the advancement of change detection and building damage assessment for humanitarian assistance and disaster recovery research. Natural disaster response requires an accurate understanding of damaged buildings in an affected region. Current response strategies require in-person damage assessments within 24-48 hours of a disaster. Massive potential exists for using aerial imagery combined with computer vision algorithms to assess damage and reduce the potential danger to human life. In collaboration with multiple disaster response agencies, xBD provides pre- and post-event satellite imagery across a variety of disaster events with building polygons, ordinal labels of damage level, and corresponding satellite metadata. Furthermore, the dataset contains bounding boxes and labels for environmental factors such as fire, water, and smoke. xBD is the largest building damage assessment dataset to date, containing 850,736 building annotations across 45,362 km\textsuperscript{2} of imagery.

1. Introduction

xBD addresses the need for large-scale, diverse satellite imagery and annotations to support remote building-damage assessment. The dataset was developed with disaster-response expertise and supports machine-learning research and the xView 2 challenge.

  • xBD provides a large-scale satellite-imagery dataset covering diverse disasters and geographical locations.
  • Over 800,000 building annotations span over 45,000 km2 of imagery.
  • Disaster-response experts created the annotation scale and supported a rigorous, repeatable, and verifiable quality-control process.
  • The xBD dataset supports the xView 2 challenge, which targets accurate and efficient models for assessing building damage from pre- and post-disaster imagery.
  • The paper reports the finalized xBD dataset alongside its requirements, annotation scale, collection process, statistics, baseline model, and broader use cases.

2. Related Work

Existing HADR datasets and assessment scales provide limited support for comparable, multi-disaster satellite-based damage assessment. The paper identifies gaps involving disaster coverage, damage granularity, imagery quality, and contextual information.

  • Many existing satellite-imagery datasets are limited to single disaster types and lack common damage-assessment criteria.
  • Damage is continuous, but limited personnel and time have constrained analysis centers to binary damaged-versus-undamaged labels.
  • Haze and mild occlusion can complicate electro-optical imagery analysis, while lower-resolution SAR imagery may not be available.
  • Environmental context such as water, fire, and smoke is useful to analysts, yet relatively few datasets provide it at this granularity.
  • Multiple disjoint assessment scales increase the cognitive burden for HADR organizations.
  • HAZUS supports multi-disaster classification, but some criteria, including wall and window condition, are difficult to determine from satellite imagery.

3. Requirements for the xBD Dataset

xBD was designed around operational HADR requirements while remaining useful for broader research. Its requirements emphasize multi-disaster coverage, realistic damage variation, high-fidelity imagery, and negative examples.

  • xBD requirements guided data collection, labeling, quality control, and dissemination for HADR and broader research needs.
  • 3.1. Multiple Levels of Damage: xBD represents damage as a continuum because agencies often lacked capacity for multiple damage levels despite damage not being binary.
  • Imagery was targeted below 0.8 meter ground sample distance to provide visual cues for distinguishing often-minute damage differences.
  • The dataset was intended to represent multiple disaster types so one model could potentially support agencies across events.
  • xBD sampled regions with varying building sizes and densities to represent geographic and structural variation in damage.
  • Negative imagery with no damage or no damaged buildings was considered critical and should comprise a sizable dataset percentage.

4. Joint Damage Scale

The Joint Damage Scale provides a unified, four-level framework for satellite-based building-damage assessment across disaster types, structures, and locations. It is designed as an operational compromise rather than an authoritative in-person rating.

  • The Joint Damage Scale addresses the cognitive burden and limited label transfer caused by agencies using different disaster-specific scales.
  • The scale was developed with insight from NASA, CAL FIRE, FEMA, and the California Air National Guard and grounded in established damage-assessment literature.
  • The scale is a first attempt at unified satellite-based assessment across multiple disaster types, structure categories, and geographical locations, not an authoritative rating.
  • Four levels range from no damage (0) to destroyed (3), balancing annotation utility with ease of labeling.
  • Satellite-imagery limitations such as resolution, azimuth, and smear force a trade-off between operational relevance and technical correctness.

5.1. Imagery Source

xBD imagery was sourced primarily from the Maxar/DigitalGlobe Open Data Program and supplemented with eight additional Tier 3 disaster events. The dataset uses high-resolution imagery spanning disparate global regions and provides three-band RGB formats.

  • Imagery came from the Maxar/DigitalGlobe Open Data Program, selected for high-resolution coverage across disparate world regions.
  • The authors selected 11 disaster events from the Open Data Program catalog and identified eight additional events as Tier 3.
  • Tier 3 areas of interest were activated through partnerships with Maxar and the National Geospatial-Intelligence Agency.
  • All xBD imagery is available in three-band RGB format.

5.2. Annotation Process

xBD pairs pre- and post-disaster imagery, transfers pre-event building footprints to post-event scenes, and assigns damage classifications through reviewed annotation rounds. It also includes polygons for visible environmental factors such as smoke, fire, and flood water.

  • A web-based in-house tool created building polygons and damage classifications through a multi-step annotation process.
  • Annotators reviewed imagery to identify usable damaged areas while retaining buffer regions that supplied negative imagery.
  • Pre- and post-disaster image pairs were aligned before annotators drew visible building footprints on the pre-disaster imagery.
  • Pre-event polygons were overlaid on matching post-event imagery to represent ideal building footprints before damage altered them.
  • Damage labels used the Joint Damage Scale and underwent repeated annotation, consistency review, and expert validation; experts found approximately 2-3% mislabeled annotations.
  • The dataset adds environmental-factor polygons for smoke, fire, flood water, pyroclastic flow, and lava.

5.3. Design Trade-Offs

The annotation pipeline addressed edge cases, image-pair alignment, missing post-disaster polygons, and variable ground sample distance through explicit processing choices and trade-offs.

  • Annotation design incorporated trade-offs for damage-classification edge cases, re-projection, and imagery-quality issues.
  • Post-disaster pixels were shifted to compensate for polygon drift caused by differing acquisition times, sensors, viewing angles, and sun elevation.
  • Average pixelwise shifts were converted to meters using image ground sample distance and applied as UTM shifts to post-disaster imagery.
  • Post-disaster buildings without transferred polygons were ignored when absent from pre-imagery, outside the building definition, or too occluded for accurate polygon drawing.
  • Ground sample distance can vary within the same geographic region, and Maxar’s proprietary pansharpening output is not reported in metadata.

6. Dataset Analysis

xBD combines broad multi-disaster coverage with detailed building annotations, but its imagery, polygons, and damage labels are unevenly distributed across events. The dataset therefore provides substantial scale while retaining important imbalance challenges for damage assessment.

  • 19 disasters, 22,068 images, 850,736 building polygons, and 45,361.79 km2 of imagery comprise the complete xBD dataset.
  • xBD uses train, test, and holdout splits in an 80/10/10% ratio to support the xView 2 challenge’s evaluation structure.
  • Disaster events are unevenly represented by total imagery area, and positive imagery does not directly correlate with each event’s represented area.
  • Certain disasters are polygon-dense; the Mexico City earthquake and Palu tsunami contribute many polygons despite relatively low image areas.
  • Damage labels are highly skewed toward no damage, with 313,033 no-damage polygons versus 36,860 minor, 29,904 major, 31,560 destroyed, and 14,011 unclassified polygons.
  • The dataset’s unbalanced disaster and label distributions, minute visual differences between damage levels, and variable negative imagery create challenges for localization models.

7. Baseline Model

The xBD baselines pair a modified U-Net localization model with a dual-stream ResNet50/CNN classification model and an ordinal loss. Evaluation uses weighted F1, with classification achieving a low overall baseline score amid confusion between adjacent damage classes.

  • Baseline localization and classification models were created to assess xBD difficulty and provide a starting point for the xView 2 challenge.
  • The localization model, based on an altered U-Net, achieved IoU values of 0.97 for background and 0.66 for building.
  • The classification model combines a pre-trained ResNet50 stream with a shallow randomly initialized CNN stream before dense-layer classification.
  • Ordinal cross-entropy penalizes predictions according to their distance from the true ordinal damage class, supporting distinctions among damage levels.
  • Weighted F1 is the primary evaluation metric because xBD is imbalanced, whereas accuracy could reach 75% by always predicting no damage.
  • 0.2654 overall weighted F1 was achieved by the baseline classifier, with major damage often classified as minor damage.

8. Conclusion

Accurate building-damage assessment is important for disaster relief operations but is currently dangerous and time-consuming. xBD supports remote automation by providing imagery, annotated polygons, and an annotation scale across diverse disaster types and regions.

  • Accurate building-damage assessments help allocate aid, personnel, and resources during natural disaster relief operations.
  • Distinct visual damage patterns make building-damage assessment automatable rather than exclusively manual.
  • xBD provides imagery, annotated polygons, and an annotation scale across visually distinct disaster regions, enabling remote vision models for damage assessment.
Loading 1911.09296v1…