Source-linked AI summary
AlphaEarth Foundations: An embedding field model for accurate and efficient global mapping from sparse label data
Christopher F. Brown, Michal R. Kazmierski, Valerie J. Pasquarella, William J. Rucklidge, Masha Samsikova, Chenhui Zhang, Evan Shelhamer, Estefania Lahera, Olivia Wiles, Simon Ilyushchenko, Noel Gorelick, Lihui Lydia Zhang, Sophia Alj, Emily Schechter, Sean Askay, Oliver Guinan, Rebecca Moore, Alexis Boukouvalas, Pushmeet Kohli
TL;DR
High-quality labels remain scarce for global Earth-observation mapping, motivating methods that translate sparse measurements into maps. AlphaEarth Foundations generates a general geospatial embedding representation from multiple spatial, temporal, and measurement contexts. Across diverse mapping evaluations, it consistently outperforms the tested featurization approaches without retraining, while annual global embedding layers have been released for 2017–2024.
Problem
High-quality labeled data are scarce for global Earth-observation mapping because physical measurements and observations require substantial effort.
Method
AlphaEarth Foundations generates 64-byte embeddings that summarize Earth-surface and climatic activity over conditioning time periods from multiple encoded data sources.
Results
AEF consistently outperforms designed and learned featurization approaches across diverse sparse-data mapping evaluations, reducing error magnitudes by approximately 23.9% on average versus the next-best approach.
Takeaways & Limitations
AEF embeddings are broadly applicable across biodiversity, ecology, agriculture, and other mapping domains where spatial and temporal changes must be modeled from sparse annotations.
Takeaways & Limitations
Adequate general 1-shot or 10-shot performance remains an unsolved research frontier because extreme low-shot results are highly variable.
Abstract
from arXiv · showhide
Unprecedented volumes of Earth observation data are continually collected around the world, but high-quality labels remain scarce given the effort required to make physical measurements and observations. This has led to considerable investment in bespoke modeling efforts translating sparse labels into maps. Here we introduce AlphaEarth Foundations, an embedding field model yielding a highly general, geospatial representation that assimilates spatial, temporal, and measurement contexts across multiple sources, enabling accurate and efficient production of maps and monitoring systems from local to global scales. The embeddings generated by AlphaEarth Foundations are the only to consistently outperform a suite of other well-known/widely accepted featurization approaches tested on a diverse set of mapping evaluations without re-training. We have released a dataset of global, annual, analysis-ready embedding field layers from 2017 through 2024.
Introduction
AlphaEarth Foundations addresses the challenge of converting sparse, heterogeneous Earth-observation measurements into accurate maps by learning a universal geospatial feature space. Across tested domains, it consistently outperforms designed and learned alternatives while supporting temporally precise, planet-scale mapping.
- Introduction: AEF generates a universal feature space whose features consistently achieve top performance across tested application domains without a single dominant prior approach.The comparison spans general and domain-specific approaches.
- Introduction: Sparse, high-quality labels and localized measurements make accurate, efficient scaling of detailed maps an open challenge.Global mapping must balance measurement precision against spatial coverage.
- Introduction: Designed Earth-observation features can extrapolate labels efficiently, but are often noisy, sensor-dependent, and region- or application-specific.Learned approaches also do not always outperform designed features in scarce-data regimes.
- AlphaEarth Foundations: AEF reduces error magnitudes by ∼23.9% on average while providing 10-meter resolution with 64-byte representations.The representation requires 16x less information than the next-most compact learned method.
- AlphaEarth Foundations: The work releases annualized planet-scale embedding-field layers and an evaluation suite under an open license.The embedding-field layers cover global mapping outputs, while the suite is designed to replicate realistic mapping scenarios.
- AlphaEarth Foundations: AEF summarizes multisource observations over specified valid periods, supporting interpolation and extrapolation even without observations inside the requested interval.The embeddings are 64 bytes and can be applied to time-dependent problems without fine-tuning.
- Evaluation in realistic data-scarce scenarios: Evaluations found AEF consistently outperformed designed and learned featurization methods, with average error reductions of ∼10.4% in ten-shot and ∼4.18% in one-shot trials.The next-best approach varied across evaluation datasets and transfer methods.
Thematic mapping
AEF supports thematic mapping across 11 classification evaluations spanning land use, land cover, crop, and species-distribution applications. It generally improves spatial coherence while retaining precision and shows broad performance generality across diverse evaluations.
- Thematic mapping: 11 classification evaluations cover land use, land cover, crop detection, crop type, and species distribution mapping.The datasets differ in class count, semantic complexity, and summary period.
- Thematic mapping: The evaluation results are reported using Balanced Accuracy for classification and R2 for biophysical-variable assessments.Table 1 identifies the evaluation domains, geographic coverage, temporal cadence, and maximum trial sizes.
- Thematic mapping: AEF demonstrates improved spatial coherence without loss of spatial precision in qualitative comparisons.
- Thematic mapping: AEF’s consistent performance across diverse thematic mapping evaluations suggests generality previously unavailable even with higher-dimensional learned embeddings.
Change detection
AEF outperforms other models in directly supervised change detection for the reported land-cover and land-use evaluations. Under unsupervised thresholding, it leads for land-cover change but not land-use change, indicating a role for supervision in this use case.
- Change detection: Direct classification treats change between two summary periods as a binary label, whereas unsupervised detection thresholds continuous deviation from an expected value.The evaluations use labels combining observations from different years.
- Change detection: 78.4% ± 1.11 BA (linear) and 79.3% ± 1.67 BA (kNN, k=3) are AEF’s direct-supervision results for land-cover and land-use change detection.The corresponding next-best baselines achieved 72.0% ± 1.28 BA and 71.5% ± 2.33 BA.
- Change detection: 71.3% ± 1.14 BA for land-cover change exceeded the next-best 67.0% ± 1.28 BA under unsupervised thresholding.For land-use change, AEF reached 71.4% ± 2.08 BA versus ViT at 72.9% ± 1.97 BA.
- Change detection: The results suggest that supervision has value for this change-detection use case.
Scaling source data quantity and type
AEF performance generally improves as training observations and source groups increase, although gains diminish and saturation varies by evaluation. Its performance usually exceeds approaches trained with equivalent observation counts, with a notable US trees exception.
- Scaling source data quantity and type: The AEF training dataset contains over 3 billion observations across nine gridded sources and one unstructured text source, covering approximately 1.1% of Earth’s land surface.
- Scaling source data quantity and type: AEF performance generally exceeds approaches trained with an equivalent number of observations and always outperforms other methods with the full training set.
- Scaling source data quantity and type: Performance saturated between 100 million and 1 billion observations for LUCAS land use and Africa crop mask, but saturation was not obvious for US trees.
- Scaling source data quantity and type: AEF requires approximately 100x more observations than SatCLIP for US trees, possibly because AEF receives no coordinate information and therefore needs more examples to learn climate gradients.
- Scaling source data quantity and type: AEF is most performant when trained on the full set of optical, radar, LiDAR, environmental, and annotated source groups, with diminishing returns as groups are added.
Global embeddings dataset
The authors released annual AEF embedding summaries as an image dataset on Google Earth Engine to facilitate use by Earth-observation practitioners. The dataset is intended to reduce the compute and storage burden of mapping workflows.
- Global embeddings dataset: Annual embedding summaries generated by AEF are hosted as an image dataset on Google Earth Engine.
- Global embeddings dataset: The released embedding fields are intended to reduce compute and storage overhead for Earth-observation mapping workflows.The authors describe these workflows as often requiring large training datasets, compute-intensive models, and custom inference systems.
Conclusions
AlphaEarth Foundations combines diverse geospatial observations into a time-continuous embedding space that models temporal dynamics and cross-source relationships. It is presented as a robust representation for compactly describing Earth’s surface despite noise and sparse observations.
- AEF combines diverse geospatial observation records into a time-continuous embedding space by modeling temporal dynamics and cross-source relationships.
- Separating measurement-specific information from mutual information across sources yields compact Earth-surface descriptions robust to noisy and sparse Earth-imaging data.
- AEF embeddings are broadly applicable across biodiversity, ecology, and agriculture, including settings where large annotation corpora are unavailable.
Data and Materials Availability
The project releases annualized embedding field layers from 2017–2024 and related resources under an open license. These materials include evaluation datasets and training-site locations for further exploration and applied use.
- Annualized embedding field layers covering 2017–2024 are released for further exploration and applied use.
- The release includes evaluation datasets and the locations of training sample sites under an open license.
- Training used publicly available data from Copernicus, USGS, NASA, JAXA, and the Copernicus Climate Change Service of the European Commission and ECMWF.
S1. Data sources and preprocessing
The training data combine diverse image and text sources with standardized spatial preprocessing and sensor-specific processing. The sources span optical, radar, climate, and elevation-related observations at varied resolutions and temporal refresh rates.
- AEF training uses image and text sources spanning diverse imaging modes and measurement spaces, with raster data sampled from the Earth Engine Data Catalog.
- Raster inputs are reprojected to UTM coordinates, resampled to 10 m resolution, standardized using per-band statistics, and clipped beyond 6 standard deviations.
- Sentinel-2 supplies moderate-resolution multispectral optical imagery, while Landsat 8 and 9 add optical and thermal observations with 15 m, 30 m, and 100 m bands.
- Sentinel-1 contributes calibrated C-band SAR observations, including available polarization and angle bands, with preprocessing for decibel scaling and value masking.
- PALSAR-2 contributes ortho-rectified, radiometrically terrain-corrected L-band SAR imagery with HH, HV, and LIN bands when available.
- ERA5-Land monthly aggregates contribute precipitation, temperature, dewpoint temperature, and surface-pressure variables, while GEDI provides LiDAR observations.
S2. Modeling
AEF is trained from globally sampled, time-bounded multimodal observations to produce embeddings that represent spatial and temporal surface variation. Its temporal conditioning and reconstruction objectives support continuous predictions from irregular observations, while consistency training reduces—but does not eliminate—tile artifacts.
- Training data: AEF training uses 8,412,511 video sequences containing 3,047,520,515 time-stamped frames from 5,145,244 sites.Each frame covers a 1.28 km x 1.28 km area, and sites are divided into approximately year-length periods.
- Training data: The global training dataset samples terrestrial and near-shore ecosystems across space, time, and data-source availability.Sampling combines geocoded text, ecoregions, coral reefs, and intertidal ecosystems.
- Training data: 10,203,798 unique (x, y, t_start, t_end) rows proceed to data-source collection after assigning two non-overlapping temporal support periods per site.Some sequences are later dropped because of insufficient imagery availability.
- Temporal modeling: AEF conditions temporal summaries on a valid period [t_s, t_e), including interpolation and extrapolation when observations are absent within that interval.Time-axial attention pooling uses a learned query derived from sinusoidal timecodes.
- Reconstruction: Conditional source decoders reconstruct each source from an embedding, timecode, and measurement metadata, producing spatially continuous predictions at arbitrary timestamps.The decoders operate across the output grid and can generate dense, superresolved predictions such as LiDAR profiles.
- Training objectives: Consistency training balances reconstruction quality with teacher–student agreement, but irregular inputs still leave visible tile artifacts in embedding-field layers.The overall contrastive objective weight is c=0.02.
S3. Evaluation datasets
The evaluation suite is designed to test low-shot geospatial representation performance on realistic classification, regression, and change-detection tasks. It uses precisely georeferenced, time-valid public reference data and produces 15 derivative evaluations from openly available datasets.
- Benchmark principles: The benchmark avoids requiring identical source imagery for all methods so representations using time and additional sources are not artificially penalized.Annotations are treated as ground truth rather than measurements tied to a specific observation.
- Processing: Reference datasets are standardized into 15 derivative datasets, with spatial filtering used to reduce autocorrelation between training and test sites.Some source datasets yield multiple evaluations through alternative label hierarchies or combinations.
- LCMAP: LCMAP contributes separate land-cover, land-use, land-cover-change, and land-use-change evaluations representative of CONUS operational mapping.Change labels are derived from sequential annual labels, while classification labels cover 2017–2021.
- LUCAS: LUCAS provides a challenging detailed assessment because its ground-based survey distinguishes fine-grained land-use and land-cover concepts.The processed data retain observations from 2017 onward with location precision of at least 10 meters.
- Canada crops: Canada crops (fine) contains 5,565 training points and 9,001 test points after preprocessing and spatial proximity filtering.The evaluation uses proportional class allocation and requires more than 100 samples per retained category.
S4. Details on evaluation setup
The evaluation tests representations across label-sparsity settings using simple transfer predictors, change-detection procedures, and uncertainty estimates tailored to trial size. Metrics are evaluated through cross-validation, bootstrap resampling, or nested resampling when datasets are small.
- Each evaluation varies label sparsity and transfer method to test downstream model performance.
- Predictors: Linear probes and kNN are used because they support low-shot domains with minimal parameterization.Linear classification uses one-vs-rest ordinary least squares with λ = 0, while regression uses a simple least-squares fit.
- Change detection: Change classification concatenates before-and-after embeddings and applies the same downstream predictors used elsewhere.Unsupervised change detection instead compares normalized embeddings using a threshold selected to maximize balanced accuracy.
- Validation design: Unequal class counts determine k-fold construction from the least represented class, while equal class counts use one fold.For Canada crops coarse, the least represented class has 75 examples and produces 273 folds.
- Uncertainty estimation: Low-shot trials estimate metric distributions by k-fold cross-validation over randomly sampled, class-balanced training sets.Each sampled training set receives an independent predictor, and the resulting metrics form a normal distribution.
- Uncertainty estimation: Full-trial uncertainty uses bootstrap statistics by resampling the validation split with replacement 100 times.Smaller datasets additionally combine multiple training sets with nested bootstrap resampling of validation data.
S5. Baseline comparisons
The study compares AlphaEarth Foundations with designed features, learned geospatial representations, geographic controls, and a generic vision model. Baseline compatibility requires additional extraction, resampling, and temporal aggregation choices that differ across methods.
- Compared representations: The baseline suite spans EO feature engineering, deep geospatial models, geographic controls, and a generic ImageNet-trained vision model.These controls assess the predictive value of location and of a model trained on ordinary camera imagery.
- Comparison constraints: Baseline comparisons are not always straightforward because methods handle spatial, temporal, and channel dimensions differently and may require specific data sources.For each study window or extent, baseline features must be regenerated from the appropriate Earth-observation data and model pipeline.
- Spatial extraction: Coarser baseline outputs are bilinearly resampled to 10 m, with embeddings extracted at the precise evaluation location and sub-pixel aliasing retained.
- Spatial extraction: For coarse-token ViT-based models, evaluation fits a per-pixel linear decoder after resampling rather than a full-patch decoder.All evaluations use the same center pixel, matching one of the linear combinations available under full-patch decoding.
- Geographic controls: XY and XYZ geographic controls are not time-varying, preventing change-detection evaluation and making poor performance generally expected for dynamic landscapes.XYZ extends XY with normalized elevation, producing five-dimensional positional embeddings.
- Vision baseline: The ViT control accepts only Sentinel-2 RGB inputs, embeds images independently across time, and averages outputs into a 1024-dimensional embedding.Its hyperparameters were tuned on the evaluation set, and this version was more performant than some non-controls in several instances.
- Composite baseline: Composite baselines combine observations across time, using optical medians and radar means after preprocessing, masking, and valid-period filtering.The selected composite configuration was chosen after tuning alternatives involving date ranges, aggregation, source combinations, and masking.
S6. Additional results & discussion
AEF performs strongly across classification and regression evaluations, but extreme low-shot settings remain highly variable. Its comparative advantage differs by application, while change-detection methods are generally less differentiated.
- Low-shot comparisons: AEF error reductions exceeded 1.0x in only 8/15 evaluations for 10-shot trials and 5/15 for 1-shot trials, with substantial variation.The authors describe adequate general 1-shot or 10-shot performance as an unsolved research frontier.
- Classification: AEF showed consistently strong classification performance, while the next-best approach varied and sometimes matched null expectation.This pattern held across binary, detailed land-cover, and genus-level tree-mapping problems.
- Application-specific comparisons: AEF often improved over other feature spaces, while different applications favored phenological, localized, or spatial-context representations.CCDC was strongest for several crop tasks, SatCLIP for Ethiopia crops and US trees, and MOSAIKS or Clay for land-cover tasks.
- Classification: AEF typically achieved the highest balanced accuracies in max-trial linear-classifier experiments, except Ethiopia crops, where kNN with k=1 was preferred.The Ethiopia crops evaluation was especially sparse and fine-scale, with generally poor performance across methods.
- Regression: AEF achieved the best overall regression performance in both R2 gains and MAE reductions, while several baselines performed worse than null on selected datasets.On OpenET, AEF was the only approach with viable results across all considered kNN and linear predictors.
- Change detection: Change-detection performance differed less across methods, with Prithvi and linear MOSAIKS near the random baseline; location-only baselines were omitted.SatCLIP, XY, and XYZ lack time handling needed to distinguish observations across dates.
S7. Ablations
Ablations show that AEF generally benefits from additional observations and source groups, although performance depends on the evaluation. Bottleneck dimension and noise also interact with task complexity.
- Observation scaling: AEF improved monotonically as additional observations were added for 9-of-15 evaluation datasets, with nonmonotonic behavior in some others.No obvious grouping or regional bias was identified.
- Source-group ablations: The source groups comprised Optical, Radar, LiDAR, Environmental, and Annotated measurements.These groups were evaluated as ablations of the AEF input sources.
- Source-group ablations: All source groups produced the best performance in 11-of-15 evaluations, although individual tasks preferred different measurement combinations.Descals oil palm favored Optical + Radar + LiDAR, consistent with information about sub-canopy structure.
- Bottleneck characteristics: The production AEF setting used embedding dimension D=64 and VMF κ=8e3 across trial sizes and transfer methods.Figure S22 evaluates nearest neighbors with k=1 or k=3 and linear probes.
- Bottleneck characteristics: Performance as a function of embedding dimension and VMF κ varied considerably across evaluation datasets.Larger-legends tasks generally favored higher embedding dimensions and more concentrated noise in max-trial settings.
S8. Inference
Inference and postquantization procedures support compact, globally tiled embedding fields. The released data use 8-bit quantization selected for its storage–performance tradeoff, with tiled inference designed to avoid seams.
- Quantization: Quantization converts 32-bit floating-point values to signed 8-bit or 16-bit integers, with clipping and restoration through dequantization.The exponentiation preserves information in the least significant digits of dequantized values.
- Quantization: The released embedding fields use 8-bit quantization with power=2 because it offered the best storage–performance tradeoff.The authors observed little performance variability relative to non-quantized embeddings.
- Global inference: Global annual embedding fields are generated by tiling UTM zones into 960m x 960m tiles buffered to 1.28km x 1.28km for inference.The outer 80m is trimmed before rendering to reduce boundary artifacts.
- Global inference: The inference system runs entirely on Earth Engine and scales to hundreds of billions of observations.The authors connect annual embedding fields with reducing the need for costly field campaigns.
- Model version: The reported results reflect AEF v2.0, followed by fixes and improvements informed by feedback on annual embedding fields.The revised training dataset increased video sequences from 8,412,511 to 10,182,450.