Source-linked AI summary

Beyond Point Predictions: Uncertainty-Aware Satellite Poverty Mapping for Public Policy

Markus B. Pettersson, James Bailie, Mohammad Kakooei, Eagon Meng, Adel Daoud

arXiv:2608.23322v1cs.LG

TL;DR

High-resolution poverty mapping is valuable where surveys are scarce, but EO-ML predictions require uncertainty guarantees before supporting policy decisions. The paper combines quantile regression, conformal calibration, and targeted surveys to produce reliable decisions, while finding that intervals remain wide despite state-of-the-art point accuracy. It concludes that EO-ML can complement traditional measurement when uncertainty is explicitly incorporated.

  • Problem

    EO-ML poverty estimates need uncertainty quantification because strong average performance does not ensure reliable decision-making.

  • Method

    The paper combines simultaneous quantile regression, conformal prediction, satellite imagery, and labeled calibration data to construct statistically valid prediction intervals.

  • Results

    EO-ML can combine with surveys to support accountable, uncertainty-aware policy decisions rather than being used naively or discarded.

  • Takeaways & Limitations

    EO-ML predictions and surveys become complements when model outputs are acted on where decisive and follow-up measurement is used where evidence is insufficient.

  • Takeaways & Limitations

    Prediction intervals remain wide because their widths are driven by residual errors in the underlying point predictions.

Abstract

from arXiv · show

Despite their critical importance for policy and research, high-resolution poverty data remain limited across much of Africa. Machine learning (ML) with earth observation (EO) imagery has recently emerged as a way to supplement these data by predicting (i.e., estimating) poverty where it has not been directly measured. Yet to be used reliably, decision-makers and analysts need assurances that they will not be misled by the errors in these predictions. To meet this need, we develop an uncertainty-aware EO-ML method for poverty mapping based on simultaneous quantile regression and a novel form of conformal prediction. Using a spatiotemporal transformer trained on sequences of Landsat and nighttime-light images, we produce prediction intervals for neighborhood-level International Wealth Index estimates across Africa which are statistically guaranteed to achieve their desired coverage rates. While our method's point-prediction performance matches the state of the art, its prediction intervals are wider than might be expected given its high $R^2$ of $0.75$. However, other models of similar accuracy likely suffer from comparable uncertainty, pointing to an inherent limitation: even with its remarkably high explanatory power, EO-ML cannot naively be relied upon for policy-making, such as when designing poverty-targeting programs. To handle this challenge, we develop a procedure to efficiently allocate aid using both ground-truth surveys and model predictions while provably ensuring the risk of excluding eligible neighborhoods remains below a prespecified level. In simulations, this approach delivers substantially more aid per eligible recipient than other strategies, thereby demonstrating that EO-ML can indeed be a reliable supplement to traditional data sources---as long as methods

Significance Statement

Satellite-based machine learning can produce detailed poverty maps where surveys are scarce, but prediction errors can make unadjusted outputs unreliable for policy and research. The paper develops uncertainty-aware methods with explicit reliability guarantees for downstream decisions.

  • Satellite imagery and machine learning produce detailed poverty maps in regions where household surveys are scarce.
  • Prediction errors can make otherwise informative poverty maps lead to incorrect decisions and unreliable findings.
  • Uncertainty-aware methods are needed because these estimates increasingly inform policy and research.

Introduction

The paper addresses the tension between scarce, costly surveys and promising but uncertain EO-ML poverty estimates. It develops calibrated uncertainty and decision procedures that combine model predictions with targeted surveys for safer and more efficient aid allocation.

  • 0.56–0.76 are reported R2 values for EO-ML poverty models, supporting high-resolution mapping at scale but not guaranteeing decision-useful absolute accuracy.R2 measures variation explained relative to predicting the average, while residual errors can still misclassify neighborhoods near a poverty line.
  • Statistically valid uncertainty quantification is obtained by combining simultaneous quantile regression with conformal prediction.The resulting intervals are guaranteed to cover the true asset-wealth proxy at a prespecified coverage rate under mild assumptions.
  • The median prediction interval covers two fifths of observed values despite matching state-of-the-art EO-ML models in R2 and accuracy.The paper attributes wide intervals to residual errors in the underlying point predictions and expects similarly wide intervals from similarly accurate models.
  • Conformalized thresholding controls misclassification rates according to policymakers’ risk tolerances when classifying neighborhoods relative to a poverty line.The method is designed to abstain when evidence is insufficient and to control both error types at preset levels.
  • SAFE uses survey resources where EO-ML predictions are uncertain, supporting reliable aid targeting at substantially lower cost than survey-only approaches.It caps the fraction of eligible neighborhoods not receiving aid below a preset threshold.
  • EO-ML predictions and surveys can function as complements rather than alternatives when each prediction carries a calibrated uncertainty statement.

1 Methods

The methods build neighborhood-level wealth estimates and calibrated decision intervals from satellite imagery, then use conformal error control and targeted follow-up surveys for policy decisions. Evaluation uses cross-validation with designated training, calibration, and test data.

  • 1 Methods: IWI measures household asset wealth on a 0–100 scale using DHS household data aggregated within small geographic clusters.
  • 1 Methods: The EO-SQR model uses time series of Landsat and nighttime-light imagery over a 6.72 × 6.72 km neighborhood footprint.It uses simultaneous quantile regression to predict conditional wealth quantiles.
  • 1 Methods: 90% prediction intervals are constructed from the model’s 5th and 95th predicted percentiles.The 50th percentile supplies the neighborhood point estimate.
  • 1 Methods: Conformalized quantile regression calibrates naive intervals using residuals from a held-out labeled calibration set.The selected residual adjusts deployment intervals to target the desired coverage rate.
  • 1.4 Conformalized thresholding for poverty classification: Conformalized Thresholding classifies neighborhoods only when calibrated intervals lie entirely above or below poverty threshold β, otherwise returning “Indeterminate.”Under exchangeability, the two misclassification probabilities are each bounded by 1 −α.
  • 1 Methods: SAFE screens out confidently ineligible neighborhoods, gives aid to confidently eligible ones, and surveys indeterminate cases.It selects α2 by maximizing estimated aid dollars per truly eligible recipient across available values.
  • 1 Methods: Five-fold cross-validation separates model training, conformal calibration, and exclusive testing, with country-level partitions for SAFE evaluation.

2 Results

The EO-SQR model achieves strong aggregate accuracy, but its conformal prediction intervals remain wide, limiting local precision. Conformal Thresholding and SAFE use these uncertainty estimates to improve reliable classification and aid allocation under explicit error or exclusion constraints.

  • 2.1 Conformalized quantile-regression estimates: R2 = 0.75 on held-out survey locations, placing EO-SQR near the top of reported EO-ML asset-wealth performance.Reported values across eight recent papers range from 0.56 to 0.76.
  • 2.1 Conformalized quantile-regression estimates: 28.7 IWI points is the median 90% prediction-interval width, indicating substantial local uncertainty despite strong point-prediction accuracy.The interval spans 18.1 to 46.8 when centered on the mean IWI score.
  • 2.1 Conformalized quantile-regression estimates: The 90% intervals are too uncertain for many operational tasks requiring confident local decisions, including threshold-based targeting or rank-based allocation.This limitation motivates decision procedures tailored to downstream use.
  • 2.2 Reliable EO-ML classification (CT): 69.3% of samples are classified by Conformal Thresholding while meeting both target error rates exactly: 0.05 for below and 0.05 for above.Traditional CQR achieves only 34.0% decision coverage in this setting.
  • 2.3 Reliable EO-ML for aid allocation (SAFE): SAFE delivers more aid per eligible recipient than other implementable strategies across exclusion rates from 0 to 0.30 and approaches the unattainable Oracle benchmark.The evaluation covers 17 countries and varies the maximum permitted share of eligible neighborhoods left without aid.
  • 2.3 Reliable EO-ML for aid allocation (SAFE): SAFE allocates aid directly to poorer rural areas, withholds it from dense urban centers, and uses follow-up surveys in more ambiguous locations.These spatial patterns vary locally, including survey assignments in relatively wealthier rural cells in South-East Nigeria.

3 Discussion

The discussion argues that accurate EO-ML predictions still require explicit uncertainty treatment for responsible policy use. It presents CT and SAFE as complementary procedures that combine model scale with surveys while preserving safety and efficiency.

  • Conformalized intervals remain wide because interval widths are fundamentally determined by underlying prediction errors.
  • CT and SAFE incorporate uncertainty by acting when evidence is clear and deferring to follow-up measurement when it is not.
  • The procedures are predictor-agnostic and require suitable scores or quantiles plus an exchangeable labeled calibration sample.
  • Conformal guarantees depend on exchangeability and are marginal over the population, so they may not hold for all policy-relevant subgroups.
  • SAFE uses binary eligibility at a fixed poverty threshold, although welfare consequences may differ across classification errors.
  • Strong predictive performance alone is insufficient for responsible policy use of EO-ML poverty estimates.

A Detailed methods and technical setup

The data pipeline combines multispectral Landsat imagery with harmonized nighttime-light data for neighborhood-scale poverty estimation. Temporal sampling and spatial footprints are selected to preserve relevant context while remaining computationally tractable.

  • Landsat inputs use six multispectral channels from 224 × 224 patches covering approximately 6.72 × 6.72 km.
  • Up to 25 low-cloud Landsat frames are selected, plus the least-cloudy frame from the year before the survey.
  • Harmonized DMSP–VIIRS nighttime-light data provide a comparable temporal signal through 7 × 7 patches covering the Landsat area.

A.1.2 Model architecture

The model uses spatial and temporal transformers with quantile conditioning to produce IWI estimates and uncertainty intervals from satellite sequences. Simultaneous quantile regression supplies model-based bounds, while conformal calibration adjusts them toward valid target coverage.

  • Model architecture: Explicit year, month, and hour encodings represent irregular temporal sampling, with year measured relative to the survey year.
  • Model architecture: The architecture encodes Landsat spatially, combines it with nighttime-light images and conditioning tokens temporally, and outputs scalar IWI estimates.
  • Model architecture: MAE pretraining initializes the Landsat encoder using approximately 300,000 additional unlabeled African locations and a 75% masking ratio.
  • Optimization and inference settings: At inference, τ = 0.5 provides the point estimate, while τ = 0.05 and τ = 0.95 define the nominal 90% interval.
  • Simultaneous quantile regression: Simultaneous quantile regression conditions one shared network on τ to estimate conditional quantiles across the full distribution.
  • Conformalized quantile regression: The nominal intervals are informative but model-dependent, so conformal calibration is required because coverage is not guaranteed initially.
  • Conformalized quantile regression: Conformalized quantile regression uses held-out calibration residuals to adjust model intervals under exchangeability and achieve marginal target coverage.

A.3 Reliable EO-ML methods

Conformalized Thresholding (CT) converts prediction intervals into reliable threshold decisions with separately calibrated one-sided errors. SAFE uses CT to screen neighborhoods, assign aid or surveys, and optimize the allocation policy under a fixed budget.

  • Conformalized Thresholding (CT) extends conformalized quantile regression to threshold classification with set error rates.
  • CT calibrates separate conformal corrections for neighborhoods at or below, and above, the poverty threshold.The corresponding quantiles are computed from distinct calibration subsets.
  • CT labels a neighborhood “Above” or “Below” only when its calibrated interval lies entirely on that side of the threshold; otherwise it returns “Indeterminate.”
  • SAFE screens out neighborhoods unlikely to be below the poverty line, assigns aid to likely eligible neighborhoods, and surveys uncertain communities.The framework operates at neighborhood level because EO imagery and DHS coordinates resolve communities rather than individual households.
  • The exclusion-risk parameter α1 controls eligible neighborhoods mistakenly left without aid, while α2 governs the trade-off between direct aid and follow-up surveys.SAFE selects α2 to optimize estimated utility for a fixed budget.

A.3.3 Proof of Theorem 1

Theorem 1 establishes CT’s two threshold-error guarantees under exchangeability and quantile-validity conditions. The proof derives each guarantee from a separately calibrated one-sided conformal correction.

  • CT controls the false negative rate at the α1 level.The proof conditions on the new point being at or below the threshold and applies exchangeability to the lower-tail calibration scores.
  • CT controls the false positive rate at the α2 level.The proof conditions on wealthy points not classified as “Above” after the first decision step and uses the upper-side calibration quantile.
  • The calibration procedure is implemented by constructing a lower bound, finding an upper bound with highest utility, and returning the corresponding threshold rule.The algorithm evaluates candidate α2 values after fixing the FNR target α1.

A.4 Evaluation design and splits

The evaluation design prevents spatial overlap across cross-validation folds by clustering nearby neighborhoods, subdividing large or imbalanced clusters, and balancing country clusters across folds.

  • 6.72 km × 6.72 km image footprints motivate a 9.5 km minimum distance between neighborhoods assigned to different folds.The distance is the diagonal of the input square.
  • DBScan provides initial spatial clusters, while K-means subdivides large clusters and removes points that violate the inter-cluster distance constraint.The procedure was designed to replace visual inspection with a systematic and reproducible process.
  • 67,829 points remained after iterative clustering and country-level Shannon-entropy balancing.
  • For Burundi, successive K-means steps split an initial near-single cluster into more clusters and removed points between them.
  • Five-fold cross-validation uses held-out folds that are further divided into calibration and test subsets for conformal prediction.

B Previous work

Comparisons with prior EO-ML poverty-mapping work are difficult because studies use different datasets, covariates, and assumptions. Training baseline architectures on identical data and splits indicates that the proposed model’s point accuracy is representative and competitive with reported figures.

  • Comparisons across studies remain difficult because datasets, covariates, and assumptions differ.
  • Baseline architectures trained on the same dataset and splits show that SQR point accuracy is not an artifact of model choice.
  • The proposed EO-SQR model appears competitive with reported figures for asset-wealth prediction from DHS data.

C Map creation

The maps combine cross-fold predictions and intervals, but their raster-wide coverage is unestablished, so displayed intervals are descriptive rather than guaranteed.

  • Map construction: Each retained grid cell receives five-fold predictions of the 5th, 50th, and 95th conditional IWI quantiles.Point predictions use the average of the five median predictions.
  • Map construction: The five prediction intervals are aggregated by taking the median lower and upper endpoints across folds when all intervals overlap.
  • Coverage limitation: Raster-wide coverage has not been established because the continental grid is not exchangeable with any fold’s calibration sample.The displayed map and interval widths are therefore treated as descriptive.
  • Spatial resolution: Grid centers are spaced approximately 0.05875 degrees, or 6.5 km north–south, while the Landsat input footprint is 6.72 × 6.72 km.A separate population-enrichment procedure uses 6.27 km squares, and these spatial quantities are not interchangeable.
Loading 2608.23322v1…