Source-linked AI summary
Micro-Estimates of Wealth for all Low- and Middle-Income Countries
Guanghua Chi, Han Fang, Sourav Chatterjee, Joshua E. Blumenstock
TL;DR
Existing wealth and poverty data are often costly, coarse, or unavailable, limiting local policy and research. This paper develops 2.4km machine-learning estimates across all 135 LMICs, explaining 56–70% of household-level wealth variation and supporting finer geographic targeting.
Problem
Reliable, geographically disaggregated socioeconomic data are expensive and scarce, limiting efforts to address global poverty, inequality, and the Sustainable Development Goals.
Method
The authors aggregate heterogeneous satellite, network, map, and connectivity data into 2.4km grid features, training machine-learning models against DHS wealth measures.
Results
56–70% of actual household-level wealth variation is explained across LMICs, while tile-level targeting improves precision and recall relative to fully covered alternatives.
Takeaways & Limitations
The resulting estimates provide granular wealth and poverty information for geographic targeting and downstream research across all 135 LMICs.
Takeaways & Limitations
The estimates reconstruct an asset-based DHS-style relative wealth index, which does not necessarily capture broader human development or well-being.
Abstract
from arXiv · showhide
Many critical policy decisions, from strategic investments to the allocation of humanitarian aid, rely on data about the geographic distribution of wealth and poverty. Yet many poverty maps are out of date or exist only at very coarse levels of granularity. Here we develop the first micro-estimates of wealth and poverty that cover the populated surface of all 135 low and middle-income countries (LMICs) at 2.4km resolution. The estimates are built by applying machine learning algorithms to vast and heterogeneous data from satellites, mobile phone networks, topographic maps, as well as aggregated and de-identified connectivity data from Facebook. We train and calibrate the estimates using nationally-representative household survey data from 56 LMICs, then validate their accuracy using four independent sources of household survey data from 18 countries. We also provide confidence intervals for each micro-estimate to facilitate responsible downstream use. These estimates are provided free for public use in the hope that they enable targeted policy response to the COVID-19 pandemic, provide the foundation for new insights into the causes and consequences of economic development and growth, and promote responsible policymaking in support of the Sustainable Development Goals.
Affiliations:
The paper addresses severe gaps in local wealth and poverty data by producing fine-grained estimates across all 135 LMICs. It validates these estimates, quantifies uncertainty, and demonstrates their potential for more precise geographic targeting and broader policy and research use.
- Motivation: Reliable, locally disaggregated poverty data are scarce, costly to collect, and sometimes vulnerable to political capture or censorship.Only half of countries have adequate poverty data, and existing data rarely support disaggregation beyond the largest administrative level.
- Contribution: The study creates the first complete micro-estimate set covering all 135 LMICs at 2.4km resolution.It estimates absolute and within-country relative wealth for roughly 19.1 million populated micro-regions.
- Method: The model uses household wealth surveys from 56 LMICs together with heterogeneous spatial and connectivity data to generate the estimates.The training ground truth comprises 1,457,315 households in 66,819 villages, while input data are aggregated to 2.4km grid cells partly to preserve privacy.
- Validation: 56-70% of actual household-level wealth variation is explained by the estimates across LMICs.Independent validation includes census and geocoded household survey data, including a Togo test where predictions explain 76% of grid-cell variation and 84% of canton variation.
- Uses: The estimates and confidence intervals are freely available through an interactive interface for research, policy, intervention monitoring, and poverty tracking.The authors also report that mobile connectivity is highly predictive and that models generalize best to neighboring or observably similar countries.
- Applications: Geographic targeting with the micro-estimates directs a higher share of benefits to poor households than targeting based on recent nationally representative surveys.The estimates cover 100% of Nigerian wards and provide 9,770 tiles in Togo, compared with survey coverage of only 13.8% of Nigerian wards and five representative regions in Togo.
units
The supplied passages identify supplementary materials for the paper and list its title and authors.
- Document: The document contains supplementary materials for the paper.
- Title: The paper is titled “Micro-Estimates of Wealth for all Low- and Middle-Income Countries.”
- Authorship: The listed authors are Guanghua Chi, Han Fang, Sourav Chatterjee, and Joshua E. Blumenstock.
SM1. Ground truth wealth measurements
The paper uses DHS household surveys to construct standardized, asset-based village wealth labels and trains models to estimate them at 2.4km resolution. The resulting target is a DHS-style relative wealth index, not a complete measure of human development.
- DHS surveys provide nationally representative, internationally standardized household wealth data with sub-regional geographic markers.
- The model reconstructs a DHS-style relative wealth index at finer spatial resolution, while its asset-based definition does not necessarily capture broader human development.
- The Relative Wealth Index is computed as the first principal component of 15 standardized asset and housing questions.The indicators include electricity, vehicles, appliances, utilities, and housing characteristics.
- Village-level wealth labels are the mean Relative Wealth Index across surveyed households within each DHS cluster.Clusters roughly correspond to villages in rural areas and neighborhoods in urban areas.
- All prediction inputs are aggregated to 2.4km grid cells, partly because that is the highest common resolution and partly to help protect household privacy.The features include terrain, climate, connectivity, and compressed satellite information.
SM3. Spatial join
The spatial join matches jittered DHS village locations to nearby 2.4km grid cells and aggregates their features using population weights.
- DHS village coordinates are approximate because the survey program masks centroids with up to 2km of urban jitter and 5km of rural jitter.
- The procedure uses 2x2 urban or 4x4 rural grids around each village centroid to cover the village’s possible true location.
- Population-weighted averages of 112-dimensional feature vectors across these surrounding cells produce the training inputs for 66,819 villages.
SM4. Supervised machine learning
The paper predicts village wealth with gradient boosted regression trees and evaluates performance using increasingly geographically realistic cross-validation schemes. Spatially stratified validation is adopted for the main analyses because spatial autocorrelation can otherwise inflate performance.
- A gradient boosted regression tree maps 112 village-level features to average Relative Wealth Index values without ex ante feature selection.Hyperparameters are tuned through three cross-validation approaches.
- Cross-validation: Basic 5-fold cross-validation randomly partitions each country’s labeled data, but can substantially over-estimate performance because nearby observations are spatially correlated.
- Cross-validation: Leave-country-out cross-validation trains on 55 countries and tests on the held-out country, measuring geographic generalization across national borders.
- Cross-validation: Spatially stratified cross-validation separates training and test samples geographically to reduce bias from spatial autocorrelation.
- Cross-validation: 56 countries contribute separate model evaluations to the distributions of R2 values compared across validation methods.
- Cross-validation: The main analyses use spatially stratified cross-validation because the authors regard it as the most conservative and appropriate approach for geographic data.This choice lowers the reported R2 values.
SM5. Feature importance
Feature-importance analyses distinguish individual-feature associations from each feature’s contribution to the fitted model. Connectivity variables are most predictive individually, while satellite features contribute collectively.
- Unconditional feature importance is measured by each feature’s univariate regression R2 with the true wealth label across 56 countries.
- Model gain measures each feature’s average contribution across random-forest splits that use that feature.
- Connectivity features, including cell towers and mobile devices, are generally the most predictive, followed by nightlight radiance and population density.
- Individual satellite-derived features are not especially predictive alone, but their large number collectively contributes to model accuracy.The paper compares predictive performance with and without satellite imagery.
- Spatial autocorrelation can let a flexible model recognize adjacent parts of the same town, illustrating why feature performance requires geographically careful validation.
SM6. Out-of-sample estimates
The final model pools data from 56 countries, maps 112-dimensional features to relative wealth estimates, and applies spatially stratified cross-validation. Privacy protection suppresses estimates for sparsely populated tiles by aggregating neighboring regions until at least 50 people are represented.
- SM6. Out-of-sample estimates: The final model pools data from 56 countries and uses spatially stratified cross-validation to tune parameters before producing relative wealth estimates.It maps 112-dimensional feature vectors for each 2.4km LMIC grid cell to an RWI estimate.
- SM6. Out-of-sample estimates: 112-dimensional feature vectors are passed through the trained model to estimate each LMIC grid cell’s relative wealth.
- SM6. Out-of-sample estimates: 2.4km regions with estimated populations of 50 or fewer are not displayed publicly to help preserve privacy.Neighboring tiles are aggregated using population-weighted average RWI until the combined estimated population reaches at least 50.
SM7. Cross-sectional estimation
The estimates target current cross-sectional wealth, using recently available inputs despite survey data spanning multiple years. This timing mismatch introduces error and makes the estimates more suitable for permanent-income applications than for studying poverty dynamics.
- SM7. Cross-sectional estimation: The model targets accurate estimates of the current cross-sectional distribution of wealth and poverty within LMICs.
- SM7. Cross-sectional estimation: Input data are primarily from 2018, whereas the ground-truth wealth measurements span a wide range of years.Historical versions of most input data are unavailable, creating a timing gap between predictors and survey measurements.
- SM7. Cross-sectional estimation: The timing mismatch likely introduces error into the model’s estimates.
- SM7. Cross-sectional estimation: The estimates are better suited to applications requiring permanent income than to applications requiring an understanding of poverty dynamics.
- SM7. Cross-sectional estimation: The model is presented as a benchmark that can improve as more input and survey data become available.
SM8. Independent validation with census data
The study validates its estimates against independently collected census and household-survey data with geographic detail. Across administrative units and fine-scale grid cells, the model shows substantial agreement with observed wealth, though the Kenya test is more stringent and uses a different poverty measure.
- SM8. Independent validation with census data: Independent census validation covers 15 countries on 3 continents and 27 million individuals whose data were not used to train the model.
- SM8. Independent validation with census data: The pooled census comparison across 979 administrative units yields R2 = 0.72, while the average country-level population-weighted R2 is 0.86.
- SM8. Independent validation with census data: Togo’s grid-cell comparison explains 76% of observed wealth variation across 922 2.4km cells.The analogous canton-level analysis covers 260 cantons.
- SM8. Independent validation with census data: Nigeria’s model explains 50% of wealth variation across 2,446 grid cells and 71% across 774 Local Government Areas.
- SM8. Independent validation with census data: Across 44 Kenyan grid cells, predicted RWI explains 21% of PPI variation, with within-region correlations ranging from 0.41–0.78.This test is more stringent because it compares spatially proximate units within three relatively homogeneous villages and uses PPI rather than a strict wealth index.
SM10. Model accuracy in high-income nations
Because high-income countries generally lack the asset-based wealth indices used for training, the paper evaluates predictions there using regional GDP per capita as a comparison measure. It constructs population-weighted regional absolute-wealth estimates for this imperfect assessment.
- SM10. Model accuracy in high-income nations: Performance in high-income nations is assessed for completeness rather than as the model’s primary target.The model is designed to estimate wealth in LMICs and is trained using LMIC ground-truth data.
- SM10. Model accuracy in high-income nations: The high-income comparison is imperfect because these nations typically do not collect the asset-based wealth indices used to train the model.
- SM10. Model accuracy in high-income nations: The analysis compares model Absolute Wealth Estimates with OECD estimates of average GDP per capita for each small TL3 region.
- SM10. Model accuracy in high-income nations: For each region, the model’s Absolute Wealth Estimate is calculated as the population-weighted average across its 2.4km grid cells.
SM11. Confidence intervals and model error
The paper quantifies uncertainty in its 2.4km wealth estimates, finding that error depends on geographic proximity to training data and providing location- and country-level error summaries for users.
- Error patterns: The model does not show evidence of performing worse in poorer regions, unlike the reported pattern for nightlights data.This supports using the uncertainty estimates across wealth levels while retaining attention to location-specific error.
- Error patterns: Model error is lower when target countries are near countries with ground-truth data and when nearby training observations are more numerous.The authors estimate each location’s error by regressing model residuals on observable location and country characteristics.
- Robustness: Error estimates remain qualitatively stable when different subsets of predictors are used in the regression.Although coefficient estimates vary somewhat across specifications, actual location-level error estimates are not very sensitive to included variables.
- Cross-country transfer: Models trained on one country perform best when applied to countries with similar characteristics.The analysis evaluates test error as the dissimilarity between training and target countries changes.
- Uncertainty outputs: Expected model error is available as a granular map and as country-level mean, median, and standard deviation summaries.These outputs are intended to show policymakers where the model is accurate and where it is not.
SM12. Absolute wealth estimates
The paper converts within-country relative wealth rankings into approximate cross-country absolute wealth estimates. This conversion uses national GDP and inequality information but requires assumptions that make AWE less reliable than the validated RWI estimates.
- Purpose and construction: Absolute Wealth Estimates (AWE) provide a rough per-capita wealth measure for each grid cell that can be compared across countries.The method converts a country’s relative wealth distribution into a per-capita GDP distribution.
- Purpose and construction: AWE combines each grid cell’s within-country RWI rank with national mean GDP per capita and the country’s modeled wealth distribution.The distribution is parameterized using the inverse cumulative distribution of wealth and country-level Gini and GDP data.
- Limitations: The conversion relies on national wealth-distribution assumptions that may be unjustified where Gini estimates are unreliable or the ICDF approximation fits poorly.These assumptions constrain interpretation of AWE estimates across countries.
- Limitations: AWE estimates should be treated with more caution than RWI estimates, which were validated against several independent survey sources.The distinction reflects the additional assumptions required to obtain absolute rather than relative wealth.
- Distributional comparison: The predicted average wealth distribution is uniformly higher than the independently estimated 2013 global income distribution.The paper identifies average wealth as a measure of per-capita GDP and income as actual family incomes.
SM13. Targeting simulations
The simulations compare high-resolution ML-based geographic targeting with approaches based on recent DHS surveys, including targeting at tile and administrative-unit levels. They show that finer spatial resolution and complete coverage can improve targeting performance, while practical delivery constraints and simplified assumptions remain relevant.
- Simulation design: High-resolution ML estimates enable targeting at the 2.4km tile level, a spatial resolution unavailable from traditional survey data.Panel A targets households in the poorest 2.4km tiles or administrative units using population-weighted tile-level RWI estimates.
- Simulation design: Alternative DHS-based approaches use recent surveys and vary the geographic aggregation level, with imputation for units containing no surveyed households.The simulations use Nigeria’s 2018 DHS and Togo’s 2013-14 DHS; Panel C imputes unsurveyed units using nearby surveyed households.
- Simulation results: Targeting at the tile level increases both precision and recall, reducing errors of exclusion and inclusion relative to fully covered alternatives.The authors note that delivering benefits to such small geographic units may be logistically challenging.
- Simulation results: Admin-region targeting with ML estimates performs at least as well as, and often better than, targeting based on recent nationally representative surveys.ML estimates cover 100% of administrative units, whereas DHS data covered 47.8% of Togo’s cantons and 13.8% of Nigeria’s wards.
- Simulation results: Only 24 of 135 LMICs had conducted a DHS since 2015, so the micro-estimates provide geographic-targeting options that might otherwise be unavailable.The comparison therefore illustrates a potential advantage where recent nationally representative survey data are absent.
- Scope and caveats: The simulations compare universal geographic transfers and omit additional eligibility criteria, which would be expected to improve all listed methods.Real-world programs may also use proxy means tests and participatory wealth rankings.