Source-linked AI summary
Landslide4Sense: Reference Benchmark Data and Deep Learning Models for Landslide Detection
Omid Ghorbanzadeh, Yonghao Xu, Pedram Ghamisi, Michael Kopp, David Kreil
TL;DR
Landslide detection needs reliable benchmark data spanning diverse regions and triggers, with accurately annotated inventories for supervised deep learning. Landslide4Sense addresses this by combining multi-source imagery, evaluating 11 segmentation models, and finding ResU-Net achieved the best detection performance.
Problem
Accurate landslide inventories are needed to train and test supervised deep learning models, while landslide hazards affect mountainous regions worldwide.
Method
The study constructs a 3,799-patch benchmark by combining Sentinel-2 optical layers with ALOS PALSAR elevation and slope data, using OBIA followed by manual polygon verification, and evaluates 11 deep learning segmentation models.
Results
ResU-Net achieved the best landslide-detection performance, with an F1-score of 71.65%.
Takeaways & Limitations
Landslide4Sense provides a public benchmark for training models, transfer learning to unexplored regions, scaling with additional imagery, and comparing generalization across landslide-segmentation methods.
Takeaways & Limitations
Because Landslide4Sense provides single-shot images, it is limited for change-detection applications requiring multi-temporal imagery to annotate surface changes accurately.
Abstract
from arXiv · showhide
This study introduces \textit{Landslide4Sense}, a reference benchmark for landslide detection from remote sensing. The repository features 3,799 image patches fusing optical layers from Sentinel-2 sensors with the digital elevation model and slope layer derived from ALOS PALSAR. The added topographical information facilitates the accurate detection of landslide borders, which recent researches have shown to be challenging using optical data alone. The extensive data set supports deep learning (DL) studies in landslide detection and the development and validation of methods for the systematic update of landslide inventories. The benchmark data set has been collected at four different times and geographical locations: Iburi (September 2018), Kodagu (August 2018), Gorkha (April 2015), and Taiwan (August 2009). Each image pixel is labelled as belonging to a landslide or not, incorporating various sources and thorough manual annotation. We then evaluate the landslide detection performance of 11 state-of-the-art DL segmentation models: U-Net, ResU-Net, PSPNet, ContextNet, DeepLab-v2, DeepLab-v3+, FCN-8s, LinkNet, FRRN-A, FRRN-B, and SQNet. All models were trained from scratch on patches from one quarter of each study area and tested on independent patches from the other three quarters. Our experiments demonstrate that ResU-Net outperformed the other models for the landslide detection task. We make the multi-source landslide benchmark data (Landslide4Sense) and the tested DL models publicly available at \url{https://www.iarai.ac.at/landslide4sense}, establishing an important resource for remote sensing, computer vision, and machine learning communities in studies of image classification in general and applications to landslide detection in particular.
I. INTRODUCTION
Landslides pose substantial risks, while inventory creation and satellite-based detection remain difficult across diverse environments. The paper addresses these challenges with a multi-source benchmark and evaluation of deep-learning segmentation models.
- Motivation: Landslides are widespread natural hazards whose frequency is increasing with climate change and whose impacts threaten people, infrastructure, and sustainable development.The introduction cites earthquakes, intense rainfall, volcanic activity, and anthropogenic activity as major triggers.
- Challenges: Field surveys and manual visual interpretation are reliable or commonly used for inventory creation but are time-consuming, resource-intensive, dangerous, and dependent on expert judgment.These constraints complicate inventory generation and updating over large or remote areas.
- Challenges: Satellite imagery is broadly available for large-area detection, yet similar spectral signatures, diverse environments, and trigger-dependent landslide characteristics make mapping challenging.Existing semi-automated approaches often use commercial satellite or UAV imagery and supervised feature-extraction pipelines.
- Contributions: The paper introduces a multi-source benchmark containing 3,799 image patches that combine Sentinel-2 optical layers with ALOS PALSAR-derived elevation and slope information.It is designed to support training and testing of machine-learning and deep-learning models for landslide detection.
- Contributions: The study evaluates 11 state-of-the-art deep-learning segmentation models using the benchmark and makes the dataset, models, and findings publicly available.The contributions target advances in machine learning and computer vision for landslide detection.
II. PRIOR WORK
Prior work has established diverse approaches for landslide inventories and deep-learning detection, but evaluations commonly remain geographically local. The paper positions cross-region generalization and extensive annotated data as central challenges.
- Datasets: Annotated remote-sensing datasets support training and validation for applications including semantic segmentation, instance segmentation, and object detection.Their availability is described as a prerequisite for advanced remote-sensing research.
- Inventory mapping: Earlier landslide inventories used polygonal or point-based formats and methods including visual interpretation of very-high-resolution satellite imagery.Reported inventories varied in landslide counts and spatial coverage for comparable affected areas.
- Deep-learning approaches: Landslide detection studies have applied FCN, U-Net variants, GAN-based Siamese networks, CNNs, Mask R-CNN, SAR, and UAV imagery.The reviewed work spans multiple sensors, model families, and spatial resolutions.
- Open challenges: Most deep-learning models have been evaluated within local geographic regions, leaving direct applicability to novel unexplored regions unclear.The paper identifies transferability across different land-cover and morphological conditions as a key challenge.
III. DATASET DESCRIPTION
The dataset is motivated by the need for diverse labeled examples because models can perform poorly on novel or out-of-distribution landslide cases. Its study areas and data sources are selected to broaden geographic and topographic variation.
- Dataset rationale: Small labeled training datasets can produce poor supervised deep-learning classification, while even large datasets may omit contingent situations.The paper therefore treats dataset diversity as important in addition to dataset size.
- Dataset rationale: Models trained on fixed datasets may have limited performance on novel or out-of-distribution inputs, especially when landslides vary in size, shape, and geographic characteristics.The study selects landslide-affected areas from four geographic regions to increase variation.
- Study areas: The four study areas are represented over a global landslide susceptibility map generated from slope, forest loss, geology, roads, and faults.The map provides geographic context for the selected locations.
- Iburi-Tobu: Iburi-Tobu in Hokkaido experienced a magnitude 6.6 earthquake and aftershocks that caused widespread shallow-sliding landslides and related secondary geohazards.The event occurred on September 6, 2018.
- Kodagu District: Exceptional rainfall exceeding 1200 mm over one month triggered severe landslides and flash floods in Kodagu District during the late monsoon season of 2018.The area includes dissected, sloping structural hills and mixed agricultural, agroforestry, and forest land cover.
2) Kodagu District of Karnataka:
The selected case studies span geographically distinct landslide settings and triggering conditions. Kodagu represents a rainfall-triggered environment with agricultural, forested, and anthropogenically disturbed terrain.
- Study-area context: The study locates its case-study areas on a global landslide susceptibility map.The supplied figure caption identifies the map as representing the selected areas’ geographic locations.
- Kodagu District of Karnataka: Anthropogenic land-use disturbances and altered precipitation patterns aggravated rainfall-triggered landslides in Kodagu.Prior work also applied an unsupervised approach there without using an inventory dataset.
- Rasuwa District of Bagmati: Rasuwa District of Bagmati was affected by widespread landslides associated with the 2015 Gorkha and Dolakha earthquakes.The selected area lies in the higher Himalayas and includes Langtang National Park.
- Western Taitung County: Western Taitung County is exposed mainly to typhoon- and earthquake-triggered landslides in a high-precipitation subtropical setting.The area receives substantial monsoonal precipitation and was affected by Typhoon Morakot in 2009.
4) Western Taitung County:
Western Taitung County is a typhoon- and earthquake-exposed study area whose landslide inventory was compiled from earlier studies and Google Earth interpretation. The benchmark uses Sentinel-2 optical imagery together with ALOS PALSAR topographic data for landslide analysis.
- Study area: The area includes frequent small sliding landslides and larger deep-seated landslides triggered by the typhoon.The inventory was compiled from previous studies and visual interpretation of Google Earth archive images from 2011 to 2013.
- Sensor characteristics: Sentinel-2 provides 13 multispectral bands at 10, 20, or 60 m pixel spacing, with a five-day global revisit time under cloud-free conditions.Its spectral coverage extends from visible wavelengths through near-infrared and short-wave infrared.
- Sensor characteristics: ALOS PALSAR supplies DEM data at 12.5 m spatial resolution for topographic applications including hazard mapping.The paper identifies DEM and slope as widely used topographical information.
C. Landslide inventory annotation
The benchmark combines systematic object-based annotation with manual correction and integrates multispectral, topographic, and labelled patch layers. Its study-area diversity and spatial split support challenging landslide detection evaluation, while the annotations encode landslide boundaries rather than additional landslide attributes.
- Annotation workflow: The annotation workflow first uses object-based image analysis and then manually verifies and corrects every landslide polygon.The workflow uses pre- and post-landslide image differences, multiresolution segmentation, and rule-based classification before visual correction.
- Annotation scope: The annotations identify landslide locations and boundaries but exclude forming material, landslide type, and mass-movement volume.Additional imagery and existing inventory data support manual polygon correction.
- Benchmark construction: The benchmark provides 3,799 annotated patches combining Sentinel-2 optical layers with ALOS PALSAR DEM and slope layers resampled to 10 m.The multi-source design spans regions with diverse triggers, topography, and land cover.
- Benchmark construction: Each 128 × 128 patch contains 12 Sentinel-2 bands, slope and DEM layers, and corresponding labels with red landslide polygons.The figure distinguishes multispectral inputs from topographic layers and labels.
- Dataset statistics: Landslide samples generally have higher mean pixel intensities than non-landslide samples, with the largest differences in bands 4 and 5.Band 5 shows a larger category difference than bands 6 and 7 among the red-edge bands.
- Dataset statistics: The dataset varies in triggers, geography, topography, landslide shape, size, distribution, and frequency across study areas.One-quarter of each area is used for training and three-quarters for testing; random or k-fold splits would increase train-test similarity.
IV. METHODOLOGY
The study evaluates representative deep learning methods for semantic segmentation, including FCN-8s, which produces pixel-level predictions using an architecture that combines coarse and fine semantic information. Figure 8 characterizes landslide-to-non-landslide sample ratios in the training and test sets.
- Ten representative deep learning methods are selected for preliminary landslide-detection experiments on the benchmark dataset.
- FCN-8s: FCN-8s accepts arbitrarily sized images and produces pixel-to-pixel predictions for semantic segmentation.
- Figure 8 reports landslide-to-non-landslide sample ratios separately for each study area and for the average across all areas.
- The ratios are shown separately for the training set and the test set.
- FCN-8s: Its skip architecture combines coarse and fine semantic layers to improve the detail of segmentation results.
B. PSPNet
PSPNet uses pyramid pooling to aggregate global context across multiple spatial scales, while ContextNet combines low-resolution contextual processing with high-resolution semantic refinement. DeepLab-v2 captures multi-scale features with atrous convolutions and ASPP, then refines segmentation with a fully connected CRF.
- PSPNet: PSPNet uses pyramid pooling and regional aggregation to learn global context across four pyramid scales.
- ContextNet: ContextNet combines a deep low-resolution network for context with a shallow high-resolution network for semantic detail.
- DeepLab-v2: DeepLab-v2 combines atrous convolution, ASPP, and a fully connected CRF for semantic segmentation.
- DeepLab-v2: Atrous convolution reduces parameters and computation while extending filter field of view.
- DeepLab-v2: ASPP probes convolutional features at different sampling rates to capture multi-scale information, followed by CRF-based refinement.
E. DeepLab-v3+
DeepLab-v3+ uses an encoder-decoder with atrous separable convolutions to capture semantic information and restore spatial detail. LinkNet, FRRN, and SQNet provide alternative strategies for preserving boundaries, combining contextual and pixel-level information, or improving efficiency.
- DeepLab-v3+: DeepLab-v3+ uses an encoder-decoder with atrous separable convolution to capture semantics and restore spatial information.
- DeepLab-v3+: Its encoder reduces feature-map size for high-level semantics, while its decoder fills in spatial information.
- LinkNet: LinkNet bypasses spatial information from encoder layers to corresponding decoder levels, maintaining object boundaries without additional training parameters.
- FRRN: FRRN combines full-resolution boundary information with pooled high-level semantic features through a two-stream architecture.
- SQNet: SQNet combines ELU activations, a SqueezeNet-like encoder, parallel dilated convolutions, and SharpMask-like decoder refinement modules.
I. U-Net
U-Net uses encoder-decoder processing for image segmentation, while ResU-Net incorporates residual learning blocks to enhance learning and help prevent gradient vanishing. The benchmark stacks Sentinel-2 optical bands with ALOS PALSAR topographic layers into image patches for model evaluation.
- U-Net: U-Net uses an encoder path for low-level representations and a decoder path for semantic segmentation.
- ResU-Net: ResU-Net replaces plain convolution layers with residual learning blocks and uses skip connections without adding parameters to the next layer.
- ResU-Net: The ResU-Net design aims to enhance learning capabilities and can help prevent gradient vanishing.
- Benchmark data: The benchmark stacks 12 Sentinel-2 optical bands with slope and DEM layers from ALOS PALSAR to create image patches.
- Model training: The models are trained on combined labeled data from four geographical regions using Adam with a learning rate of 1e-3 and batch size 32.
B. Experimental Results
Across the evaluated deep learning models, ResU-Net achieved the strongest overall landslide-detection performance, while precision–recall trade-offs remained substantial. The benchmark also supports automated detection from freely available Sentinel-2 data, although medium spatial resolution remains a constraint.
- ResU-Net achieved the best overall performance, with an F1-score of 71.65%, followed by SQNet at 70.24%.
- U-Net obtained the highest precision at almost 80%, more than 3 percentage points above the second-highest precision from FRRN-A.
- Most models showed a precision–recall imbalance, indicating reliable detected landslide pixels but missed labeled landslide pixels.
- DeepLab-v3+ detected most landslide regions but produced many false positives, whereas ResU-Net better suppressed false-positive predictions.
- Sentinel-2’s free availability facilitates automation in cloud-free conditions, but its medium spatial resolution has limited landslide-detection research.
VI. DISCUSSION AND CONCLUSIONS
The study presents Landslide4Sense as a reference benchmark combining annotated imagery with deep-learning evaluation and feature analysis. The dataset can support transfer learning, expansion, and model comparison, but its single-shot, vegetation-heavy coverage limits some applications and transferability.
- Landslide4Sense contains 3,799 Sentinel-2 and ALOS PALSAR image patches and benchmarks eleven deep learning segmentation models.
- Shallow PSPNet layers retain detailed landslide-boundary information, while deeper layers become more abstract as feature-map size decreases.
- Combining shallow and deep features is important for preserving boundary information in landslide detection.
- The benchmark addresses limited annotated data availability and can be continuously expanded and updated for landslide research.
- The dataset supports transfer learning, larger-scale extension, and explicit comparison of generalization across new remote-sensing segmentation methods.
- Because Landslide4Sense provides single-shot images, it is limited for change detection requiring multi-temporal imagery.