Source-linked AI summary

Using satellite imagery to understand and promote sustainable development

Marshall Burke, Anne Driscoll, David B. Lobell, Stefano Ermon

arXiv:2010.06988v1cs.CYcs.CVcs.LGstat.ML

TL;DR

Reliable, comprehensive measurements of sustainable-development outcomes are difficult because ground data are scarce, costly, and noisy. The paper synthesizes satellite-imagery research, emphasizing machine-learning approaches, and finds strong, improving performance alongside training-data and adoption constraints. It argues that satellite methods will generally complement rather than replace ground-based measurement.

  • Problem

    Ground data for key sustainable-development outcomes are scarce, costly, infrequent, locally limited, and often unreliable, creating challenges for model training and validation.

  • Method

    The paper reviews and synthesizes satellite-imagery studies, emphasizing machine learning, quantifying ground-data scarcity, imagery growth, model performance, applications, constraints, and research directions.

  • Results

    Satellite-based performance is reasonably strong and improving, with estimates sometimes equaling or exceeding traditional measurement accuracy across key sustainable-development outcomes.

  • Takeaways & Limitations

    Satellite approaches can add substantial information at broad scale and low cost, while generally amplifying rather than replacing ground-based data collection.

  • Takeaways & Limitations

    Quality training labels remain scarce and unreliable, making satellite-model training and validation difficult.

Abstract

from arXiv · show

Accurate and comprehensive measurements of a range of sustainable development outcomes are fundamental inputs into both research and policy. We synthesize the growing literature that uses satellite imagery to understand these outcomes, with a focus on approaches that combine imagery with machine learning. We quantify the paucity of ground data on key human-related outcomes and the growing abundance and resolution (spatial, temporal, and spectral) of satellite imagery. We then review recent machine learning approaches to model-building in the context of scarce and noisy training data, highlighting how this noise often leads to incorrect assessment of models' predictive performance. We quantify recent model performance across multiple sustainable development domains, discuss research and policy applications, explore constraints to future progress, and highlight key research directions for the field.

1 Introduction

The review synthesizes satellite-imagery and machine-learning approaches for measuring human-related sustainable development outcomes, finding strong and improving performance but persistent data and adoption constraints.

  • Scope: The review synthesizes research using satellite imagery to measure human outcomes linked to the Sustainable Development Goals.It focuses on applications where humans or what they produce are predicted from imagery.
  • Main findings: Satellite-based prediction performance is reasonably strong, appears to be improving, and can equal or exceed traditional measurement accuracy for multiple outcomes.Reported performance may understate true performance because evaluation data are noisy.
  • Main findings: Training data, rather than imagery, may now be the largest constraint because quality ground labels remain scarce and unreliable.The review argues that expanding label quantity and especially quality would accelerate progress.
  • Main findings: Satellite approaches will generally amplify rather than replace ground-based data collection efforts.Some outcomes may never be accurately estimated from satellites, while high-quality local training data can improve predictive performance where satellite power exists.
  • Adoption: Few satellite applications have been operationalized in public-sector sustainable-development decision-making, apart from population and agricultural measurement.Limited adoption is associated with the technology's recency, model accuracy and interpretability concerns, and entrenched interests.

2 The availability and reliability of data

Ground data for sustainable-development outcomes are costly, infrequent, locally limited, and noisy, while satellite imagery is becoming more abundant and informative across spatial and temporal scales.

  • Ground-data constraints: Household and field surveys remain the main tools for measuring poverty, agricultural productivity, population, and many health outcomes.These surveys provide detailed information but are expensive and time-consuming to implement.
  • Ground-data constraints: $1.5-2 million USD is the typical cost of conducting a DHS or LSMS survey in one country for one year.The operation takes multiple years and requires enumerators to work in remote and insecure locations.
  • Ground-data constraints: At least 6.5 years separate nationally representative livelihood surveys in half of African nations.Economic household-survey frequency is also substantially lower in less wealthy countries.
  • Ground-data constraints: 25% of countries have gone more than 15 years since their last agricultural census, while 8% have gone more than 15 years since their last population census.Among African nations, 34% have gone more than 15 years since their last agricultural census.
  • Ground-data constraints: Surveys often cannot produce accurate state-, county-, or local-level statistics because they are typically representative only nationally or regionally.Limited public access to household observations and geographic collection information further constrains local research, evaluation, and model training.
  • Ground-data reliability: >25% too low is the expenditure estimate produced by some household-consumption survey designs relative to gold-standard diaries.Ground data also contain sampling variability, administrative inconsistencies, and privacy-related coordinate jitter.
  • Ground-data reliability: r = 0.39 is the highest observed correlation between household-survey and official government maize-yield measures for Ethiopia.The comparison also reveals systematic upward bias in official government estimates relative to household responses.
  • Satellite availability: Multiple times per week rather than multiple times per year is the newer revisit frequency for 1-5m and moderate- to low-resolution imagery.Information on human activity is also visible in 5-30m imagery, including infrastructure, agriculture, and moisture availability.

3 Modeling approaches using satellite imagery to predict sus-

Satellite-imagery models range from hand-crafted feature regressions to spatial deep networks and methods designed for scarce labels. Their evaluation is complicated by noisy training and test data, which can obscure predictive performance.

  • Modeling approaches: Models map satellite-image inputs to sustainable-development outputs using functions such as regressions, random forests, support vector machines, and neural networks.Training selects a model by minimizing a loss that measures differences between predicted and observed outcomes.
  • Shallow models based on hand-crafted features: Hand-crafted features, including vegetation indexes and aggregated pixel statistics, support simple predictions of yields, GDP, and other outcomes.Aggregation can use means, quantiles, or histograms, but discards much of the imagery’s spatial structure.
  • Models that use spatial structure in the imagery: Convolutional neural networks use spatial context and often outperform hand-crafted features and aggregation strategies on image-prediction tasks.These architectures include VGG, DenseNet, and ResNet-style models.
  • Model development with limited training data: Limited labeled data motivates simulation-based augmentation, transfer learning, and unsupervised or semi-supervised pretraining.Transfer learning reuses representations learned from related tasks, while unlabeled imagery helps learn features when sustainability labels are scarce.
  • Model development and evaluation with noisy data: Noisy data can impair feature learning and lead models to learn relationships that do not generalize beyond the training setting.The review highlights noisy training and test data as central concerns for model development and evaluation.
  • Model development and evaluation with noisy data: When trained on increasingly noisy data but evaluated on undegraded test data, model performance remains highly stable across three explored noise types.By contrast, evaluation on noisy data shows declining performance as more noise is added.

4 Applications

Satellite imagery supports applications from agricultural and economic measurement to population estimation and intervention evaluation, but noisy ground data and limited operational adoption constrain its use.

  • Agriculture: Objective harvest data substantially outperformed farmer self-reported data for training and evaluating crop-yield models.Georeferencing and field-area errors are more consequential for smaller fields.
  • Agriculture: 30-50 additional training samples rapidly improved yield-prediction performance, after which performance was largely stable.The result was measured using root mean squared error on held-out test data.
  • Population: Population estimates agree more closely at aggregate scales, with correlations approaching r = 1.0 for 100km pixels.Swedish 100m validation found correlations of r = 0.83 for GHSL, r = 0.82 for WorldPop, and r = 0.7 for LandScan.
  • Population: Bottom-up population models offer an alternative to estimates that inherit inaccuracies from official census data.Standard approaches commonly disaggregate official census estimates, which may be outdated or inaccurate.
  • Economic livelihoods: Satellite-derived information explained more than half, and often more than 75%, of survey-measured asset-wealth variation across the reviewed studies.Performance appeared to trend upward over time, while aggregate-scale and data-fusion models tended to outperform village-level satellite-only models.
  • Economic livelihoods: Consumption-expenditure prediction is typically lower than asset-wealth prediction, partly because consumption data are noisier and georeferenced public data are scarce.
  • Research and policy applications: Satellite information has been used to evaluate agricultural technologies, estimate who benefits, inform humanitarian response, and support population-survey sampling.Documented public-sector use is widest in population applications, while economic-livelihood use remains limited.
  • Constraints to adoption: Decision-makers may hesitate to adopt satellite-based measures because models can lack interpretability, and many promising estimates are not yet operational at global scale.The research community may need partnerships with public- or private-sector organizations to scale and sustain these estimates.

5 Conclusions and directions for future work

The review finds improving satellite-based performance but identifies training data, outcome coverage, operationalization, interpretability, and adoption as central constraints. It proposes better labels, transparent models, data fusion, scaling partnerships, and temporal evaluation as priorities.

  • Conclusions: Satellite-based performance is reasonably strong, improving, and can equal or exceed traditional outcome-measurement accuracy.Reported performance may understate true performance because evaluation data are noisy.
  • Conclusions: Training data, rather than imagery, are now perhaps the largest constraint because quality labels remain scarce and unreliable.Expanding label quantity and especially quality could accelerate progress and improve performance assessment.
  • Conclusions: Satellite approaches are likely to contribute little near term to measuring female empowerment, educational outcomes, or conflict events.Even where satellites are useful, high-quality local training data remain essential and approaches will likely amplify rather than replace ground collection.
  • Conclusions: Documented operational use in sustainable-development decision-making remains limited, with satellite-informed population estimates the main exception.Limited adoption is associated with technology recency, perceived or real accuracy concerns, poor interpretability, and entrenched data regimes.
  • Directions for future work: Future work should prioritize accurate reference datasets, explainable and transparent predictions, creative data fusion, and partnerships that scale estimates.The review also highlights repeated local ground datasets for developing and validating temporal predictions.

a b c e

Figure 6 compares satellite-informed population datasets and satellite-based predictions of wealth and informal settlements across spatial scales. Population estimates show only modest global agreement, with stronger agreement after spatial aggregation.

  • Population rasters are compared at 1km resolution across LandScan, WorldPop, and GHSL using colors to represent population scale.
  • Mean pairwise correlation between population datasets is shown across spatial resolutions from 1km to 100km.
  • Population-dataset agreement improves when estimates are spatially aggregated.
  • Asset-wealth prediction results report test-data coefficient of determination for 16 estimates from 12 studies across developing countries.
  • Informal-settlement prediction results summarize 20 estimates from 17 studies using satellite imagery and machine learning.

Collecting Satellite Revisit Data

The supplementary tables document satellite-based yield, economic-wellbeing, and informal-settlement studies, while revisit rates are computed from cloud-free imagery collected across sampled locations. Sensor groupings and cloud filtering standardize heterogeneous imagery sources for comparison.

  • Figure 3 samples 100 African and 100 US/EU locations and queries satellite imagery from 2010 and 2019.Locations are population-weighted and buffered by approximately 10 meters.
  • Planet imagery comes from SkySat, PlanetScope, and RapidEye products downloaded through the Planet API.
  • LandInfo supplies private-satellite footprints, while Google Earth Engine supplies Landsat, Sentinel, and MODIS footprints.
  • Image footprints are filtered to less than 30% cloud cover, although filtering differs across data providers.
  • Sensor groups use modal resolution values, compressing within-group variation such as DigitalGlobe’s 31cm-to-91cm range.
  • Average revisit rate is calculated as (number of locations*365)/number of images, and values below one represent average intervals rather than daily cloudless coverage.
Loading 2010.06988v1…