Source-linked AI summary
A global mobile network coverage raster product at 1km resolution, 1999--2030
Till Koebe, Theophilus Aidoo, Ali El Chami, Ali Kanso, Akansh Maurya, Purushottam Sharma, Ingmar Weber, Ridhi Kashyap
TL;DR
Sub-national mobile coverage is inconsistently observed even though usable signal availability underpins research on mobile connectivity and its societal impacts. The paper combines three independent coverage models to produce annual 1 km 2G, 3G and 4G probability layers. Track A exceeds the population-density baseline, whose AUC is 0.673 / 0.679 / 0.710 for 2G/3G/4G on the same reliable labels.
Problem
Sub-national mobile coverage is inconsistently observed even though usable signal availability underpins research on mobile connectivity and its societal impacts.
Method
The paper combines three independent coverage models to produce annual 1 km 2G, 3G and 4G probability layers.
Results
Track A exceeds the population-density baseline, whose AUC is 0.673 / 0.679 / 0.710 for 2G/3G/4G on the same reliable labels.
Takeaways & Limitations
The released maps support studying the global digital divide, linking connectivity with survey outcomes, and planning humanitarian and infrastructure interventions.
Takeaways & Limitations
The product cannot represent interrupted service or acute infrastructure loss, and its simulator relies on strong structural assumptions.
Abstract
from arXiv · showhide
Where a mobile signal is available shapes who can work, learn, bank, seek health care and respond to crises in the digital age, yet no globally consistent, sub-national record of mobile network coverage exists. We present such a record: annual 1km maps of the probability of 2G, 3G and 4G coverage for 214 countries and territories for the years 1999 to 2030. The maps are produced by three independent models: a calibrated machine-learning model, a techno-economic simulator of network build-out, and a spatial deep-learning model. The three estimates are then combined, per country and technology and in proportion to their measured accuracy, into a single best estimate with per-pixel 90% uncertainty bands; all four layers are released as part of the dataset. Because mobile roll-out closely follows a country's socio-economic conditions (population distribution, electrification, physical infrastructure), the models are grounded in existing geospatial data and tuned on 2,409 quality-screened operator-reported coverage maps, which are available up to 2020. For 2021--2024 the maps are predicted from recent geospatial data alone; for 2025--2030 they are extrapolated from demographic and infrastructure projections. On countries held out during training, the machine-learning model attains AUC 0.89--0.92. Baseline comparisons and the combined product's external validation are reported in Technical Validation. The dataset supports mapping the global digital divide, linking connectivity to household-survey outcomes, and humanitarian and infrastructure planning.
Background & Summary
The paper introduces an open global product of annual 1 km mobile-coverage probability maps, addressing inconsistent sub-national observations with filtered labels and three complementary modelling tracks.
- Background & Summary: The dataset addresses sparse and inconsistent sub-national coverage observations by screening operator maps and filling remaining country-technology-year combinations with model predictions.The reliable set contains 2,409 distinct country-technology-year maps, while 5G is excluded because too few reliable triples support supervised filtering.
- Background & Summary: Three independent tracks combine tabular learning, techno-economic rollout simulation and spatial image translation to estimate per-pixel coverage probabilities.The tracks trade off covariate exploitation, interpretable deployment assumptions and spatial-context learning.
- Background & Summary: Annual 1 km rasters provide 2G, 3G and 4G coverage estimates for 214 countries from 1999–2030, with a combined layer and three per-track estimates.The product uses a common grid and releases four estimation layers.
- Background & Summary: Coverage estimates are combined using held-out accuracy, constrained to be temporally monotone, and accompanied by conformal 90% uncertainty bands across label-anchored, modelled and projected periods.The released regimes are 1999–2020, 2021–2024 and 2025–2030.
- Background & Summary: The open product complements country-level, paywalled, quality-focused or access-gated alternatives by providing persistent cross-national coverage-extent maps.The paper states that no comparable open product exists.
Methods
The methods construct reliable modern and legacy labels through operator-recency and consistency filters, harmonise them across a common grid, and document their uneven composition.
- Methods: The reference grid retains every valid 1 km country-land cell, including zero-population cells, so labels, covariates and predictions remain co-registered.The grid is derived from WorldPop rasters in EPSG:4326.
- Reliability filter: The modern filter accepts 1,058 of 1,203 candidate triples, or 87.9%, spanning 2008–2020 across 168 countries.Acceptance increases from 80.9% for 2G to 98.1% for 4G.
- The pre-2010 legacy archive and its substitute filter: Legacy admission requires all applicable checks to pass, including ITU agreement, temporal monotonicity and 2010-boundary consistency, while unverifiable triples are rejected.Checks with unavailable inputs are skipped, but triples with no applicable check are rejected.
- The pre-2010 legacy archive and its substitute filter: The legacy substitute filter accepts 1,430 of 1,947 candidate triples, or 73.4%, covering 2000–2009 for 2G and 3G.The accepted set includes ITU-checked and consistency-only triples; no legacy 4G exists.
- The pre-2010 legacy archive and its substitute filter: The reliable label set is tilted toward legacy 2G, with 1,325 legacy 2G triples versus 448 modern 2G, 511 3G and 204 4G triples.Per-country weighting, pixel caps and temporal hold-out validation are used to address this imbalance.
Covariate layers
The global covariate stack represents demand, economic value, and deployment feasibility while preserving annual coverage across observed and projected years.
- Covariate layers: Covariates capture network-building forces through population and built-up demand, settlement structure, nighttime lights, terrain, deprivation, and infrastructure-related layers.Population and built-up surface quantify demand, while settlement classes distinguish urban, rural, and unpopulated land; terrain and deprivation reflect deployment conditions and economic value.
- Covariate layers: Every global input is computable for all 214 countries, enabling consistent spatial prediction across the product.
- Covariate layers: Annual covariate coverage is completed by interpolation between observations and carrying the nearest observed values beyond each series’ span.Nighttime lights and World Bank telecom indicators are held at 2024 through 2030, while GHSL layers use epoch-based interpolation or nearest-epoch values.
- Covariate layers: The most informative network predictors are unavailable as open, global, time-resolved data, limiting the model to less direct proxies.Confidential operator data, incomplete crowdsourced tower data, and the absence of a reliable annual global electricity-grid series constrain feature selection.
Buffer-aggregated economic features
The economic features and structural simulator translate local population, revenue potential, deployment costs, and technology constraints into network build-out predictions.
- Buffer-aggregated economic features: Revenue-potential features aggregate population, population weighted by ARPU, and built-up surface within 10 km to approximate commercially attractive deployment areas.These buffer aggregates are computed from WorldPop and align with the simulator’s economic logic.
- Buffer-aggregated economic features: Track B evaluates each pixel as a potential antenna site by comparing expected annual revenue with upgrade or greenfield construction costs.Revenue depends on population within the antenna footprint, country-year ARPU, and penetration; costs include backhaul and distance to existing towers.
- Buffer-aggregated economic features: The simulator uses fixed technology-specific antenna radii of 12, 7, and 5 km for 2G, 3G, and 4G, respectively, without density-dependent cell sizes.
- Buffer-aggregated economic features: Track B assumes operators greedily maximize predicted profit, deploy omnidirectional sites until budgets are exhausted, and models coverage extent rather than capacity or congestion.
Track C: pix2pix spatial deep learning
Track C uses a supervised U-Net image-translation model to predict per-pixel coverage from spatial covariates, with technology availability gates preventing historically impossible deployment.
- Track C: pix2pix spatial deep learning: Track C outputs per-pixel coverage probabilities from a nine-channel, year-encoded covariate tensor using a U-Net generator trained with reconstruction losses.The input is a 256×256 km patch containing normalized population, roads, built-up surface, terrain, deprivation, nighttime lights, and sine/cosine year encoding.
- Track C: pix2pix spatial deep learning: The model is trained separately for each technology in a pix2pix-style conditional image-translation setup, but the released models disable the adversarial loss.
- Track C: pix2pix spatial deep learning: Every modelling track is gated by each country’s first commercial availability year so spatial similarity cannot generate coverage before a technology existed.
- Track C: pix2pix spatial deep learning: Commercial launch years are sourced preferentially from operator, regulator, industry, and trade records, while uncertified cases fall back to ITU-derived or global technology-floor years.The certified set covers 427 of 642 country-technology cells; fallback years use the first reported 1% population coverage or global floors of 1995, 2002, and 2010.
Track combination and uncertainty
The product combines three calibrated track probabilities using accuracy-based weights and supplies empirically calibrated 90% uncertainty intervals alongside reproducible released layers.
- Track combination and uncertainty: The best-estimate layer is a per-pixel weighted mean of Tracks A, B, and C, with weights based on country-technology Brier scores.A tempered softmax of negative loss favors locally accurate tracks while retaining a small uniform floor for every available track.
- Track combination and uncertainty: A shared 46 km Sudan 2G example shows smooth probabilities, hard-edged antenna footprints, and learned segmentation boundaries aligned over identical observed pixels.The best-estimate layer combines these complementary spatial structures rather than treating the tracks separately.
- Track combination and uncertainty: The released 90% intervals use stratum-specific empirical residual quantiles around the combined mean, clipped to the [0, 1] probability range.Calibration strata combine technology, era, World Bank region, and income group, pooling to coarser strata when cell counts are small.
- Track combination and uncertainty: The intervals are empirical rather than strict split-conformal guarantees because the same reliable cells train tracks and fit combination weights.Leave-country-out validation provides the out-of-sample evidence for the released bands, while projected years and label-free countries remain extrapolations.
- Track combination and uncertainty: The full acquisition, modelling, validation, and packaging pipeline is released as a BSD-3-Clause Git repository for reproduction and audit.
Data Records
The released dataset provides global, multi-technology coverage rasters with combined estimates, uncertainty bands, model-specific layers, and supporting metadata. It spans 214 countries through 2030, but reliable observational labels are uneven across countries and technologies.
- Released layers: The release includes combined Cloud-Optimised GeoTIFF rasters, separate layers for Tracks A–C, and metadata recording held-out Brier scores, BMA weights, and recommended tracks.The combined product provides calibrated uncertainty, whereas per-track layers are point estimates.
- Released layers: The primary combined product stores a weighted mean coverage probability alongside split-conformal 90% lower and upper uncertainty bounds.Track weights are based on per-country, per-technology Brier performance.
- Illustrative record: Figure 3 illustrates the dataset’s temporally consistent 1 km record by showing Nigeria’s 2G, 3G, and 4G rollout from 1999 through 2030.Pixels are classified by the highest technology whose combined probability exceeds 0.5.
- Coverage scope: The dataset covers all 214 reference-grid countries, with reliable-triple labels for 207 and model-only coverage for seven countries.Reliable-label availability also varies by technology within countries.
- Audit metadata: The release provides audit metadata for 2,488 reliable triples, including source distinctions between modern MCE and legacy-archive labels.The metadata also includes reliability flags, publication years, and configuration snapshots.
Technical Validation
Track A achieves strong country-held-out discrimination, while temporal hold-out supports annual prediction through 2024. External validation, uncertainty coverage, and sensitivity analyses assess aggregation, calibration, gating, and training-data choices.
- Spatial 5-fold cross-validation: Population-density baselines score AUC 0.673 / 0.679 / 0.710 for 2G/3G/4G, below Track A’s multi-covariate booster.The margin is widest for 2G and narrowest for 4G, while the booster captures structure beyond population location.
- Spatial 5-fold cross-validation: AUC ranges 0.89–0.92 across technologies, with 2G highest at 0.92, followed by 3G at 0.90 and 4G at 0.89.The country-block design uses out-of-distribution calibration and reveals per-fold AUC standard deviations of 0.01–0.03.
- Temporal hold-out validation: Forward-in-time AUC is 0.891–0.944 when training on years ≤2015 and evaluating on years ≥2016, supporting annual predictions through 2024.This test concerns year extrapolation for countries seen during training; the projected 2025–2030 tail extends beyond any label.
- External validation: External validation compares population-weighted country-year shares from all four released layers with ITU’s at-least-2G/3G/4G indicators using identical aggregation and samples.Differences partly reflect distinct coverage definitions: the model uses MCE operator service-area declarations, whereas ITU relies on largest-operator reporting and imputation.
- Effect of the technology availability gate: The technology-availability gate improves agreement with ITU on every technology and both reported metrics, indicating that removed pre-launch coverage was spurious.Launch years came from deployment records and were never fitted to ITU, providing an independent check on the gate.
- Uncertainty interval coverage: Blocked uncertainty coverage is 0.899 overall and 0.897–0.904 across every technology×era stratum, while mean per-pixel band widths are 0.63, 0.64, and 0.56 for 2G/3G/4G.Bands are wider for newer technologies and sparsely labelled regions; country- and region-aggregated estimates have substantially smaller uncertainty.
Usage Notes
The dataset supports digital-divide measurement, micro-data linkage, and operational planning, but users must distinguish observed, modelled, and projected periods and account for explicit scope limits.
- Usage Notes: Coverage probabilities support sub-national digital-divide mapping, linkage with georeferenced surveys or mobile-money records, and humanitarian, targeting, and universal-service planning.Per-technology arrival timing can support event-study designs; causal analyses are recommended only on the observed label subset.
- Usage Notes: Coverage is constrained to be non-decreasing by year, so late-year 2G and 3G values represent historical-maximum footprints rather than operating networks after sunsets.Technology-retirement analyses require an external sunset source.
- Usage Notes: The model excludes interrupted service, direct-to-cell satellite connectivity, and abrupt conflict- or disaster-driven infrastructure loss.In acute supply shocks, estimates and intervals should be read as upper bounds on actual coverage.
- Usage Notes: Area-based aggregates should use the pixel area km2 column rather than raw pixel counts because the EPSG:4326 grid does not preserve area at high latitudes.The same weighting applies when producing global aggregates.
Appendix A Full technical specifications
Appendix A collects the complete technical settings referenced from the Methods section.
- Appendix A Full technical specifications: Appendix A serves as the repository for the paper’s complete technical settings referenced from Methods.It provides the detailed specifications needed to interpret the methodological configuration.
- Appendix A Full technical specifications: The appendix consolidates technical settings that are referenced elsewhere rather than introducing a separate methodological contribution.Its role is specification and documentation.
- Appendix A Full technical specifications: Readers using the Methods section can consult Appendix A for the corresponding full technical specifications.The passage identifies the appendix as the collection point for those settings.
Operator-name matching.
Operator identities are standardized by matching OpenCellID MNC records to MCE names with scoped substring matching and curated brand normalization, supplemented by documented overrides.
- Operator-name matching: Major operators are matched from the OpenCellID MNC table to MCE NAME entries using lower-cased, ISO2-scoped substring matching.A 218-entry brand-normalisation list handles known cross-database spelling variants.
- Operator-name matching: The normalization configuration embeds research notes and source citations alongside each operator entry.This makes individual matching decisions inspectable in the released configuration.
- Operator-name matching: Table A1 documents 33 operators across 16 countries that are treated as major despite falling below the OpenCellID tower-count threshold.These additional-major overrides are sourced from national regulators and industry trackers.
Track B priors and constants.
Track B combines literature-anchored priors, fixed global constants, and CMA-ES optimization, while the three-track configuration is summarized in Table A2.
- Track B priors and constants: Track B assigns log-normal priors to revenue-capture coefficients, greenfield tower cost, and fibre cost, plus a Beta(2, 38) prior to the new-tower CapEx share.The stated medians are $80k–$150k for greenfield towers and $15k–$40k/km for fibre.
- Track B priors and constants: Fixed Track B constants include upgrade cost, microwave-link cost, footprint radii, maintenance, and technology-specific maturity ramps.The maturity-ramp time constants are τ_t ∈ {3, 3, 5} years for 2G, 3G, and 4G.
- Track B priors and constants: CMA-ES fits Track B’s per-country parameters using 200 evaluations, population size 12, and σ0=0.3 in log-parameter space.Table A2 also records the three tracks’ hyperparameter, calibration, and uncertainty settings.
Data availability
The dataset is openly available from Zenodo under CC-BY-4.0, with the full 1999–2030 raster series organized into regional archives. Supporting reliability, provenance, metadata, and integrity information is distributed alongside the archives.
- Data availability: The dataset is available from Zenodo under a CC-BY-4.0 license.The release is hosted at DOI 10.5281/zenodo.21594337.
- Data availability: Twenty-two ZIP archives organize the full 1999–2030 series by UN M49 subregion, covering all countries, layers, and technologies.Each archive contains the four layers and three technologies for every country in its subregion.
- Data availability: A country-to-archive mapping and manifest report archive contents, file counts, member countries, and SHA-256 checksums.These files support locating releases and verifying their integrity.
- Data availability: Separate companion files provide reliability records, per-cell launch-year certification, source-level provenance, and machine-readable metadata.These materials identify which country–technology–year maps were used and document their provenance.