Source-linked AI summary
Earth Observation Foundation Models for Terrestrial Ecohydrology: From Representation Learning to Process Inference
Yi Yu, Jian Peng, Yucheng Lin, Trevor F. Keenan, Thomas F. A. Bishop
TL;DR
Ecohydrology lacks a synthesis connecting EOFM representations to observability, process timescales, and reference uncertainty. This paper develops an observation-to-inference framework and finds strongest evidence for spatial context, label-efficient adaptation, and hybrid workflows, while deeper process inference remains weakly validated.
Problem
Existing syntheses do not jointly evaluate EOFM representations against ecohydrological observability, process timescales, and reference independence.
Method
The paper combines an observation-to-inference framework with meta-analysis, application synthesis, and benchmark audit to assess EOFMs for ecohydrological process inference.
Results
Current EOFM evidence is strongest for spatial context, label-efficient adaptation, and hybrid workflows, while deeper process inference and physical consistency remain weakly validated.
Takeaways & Limitations
Future EOFM development should align inferred variables, observation pathways, process timescales, uncertainty budgets, and physical diagnostics through process-aware benchmarks.
Takeaways & Limitations
Benchmarking remains constrained by scarce public, diverse, openly reusable field- and sub-field-scale reference data.
Abstract
from arXiv · showhide
Earth observation foundation models (EOFMs) are emerging as reusable representation frameworks for data-driven retrieval, prediction and process modelling within ecohydrology, which integrate EO, meteorological forcing and process models to characterise coupled water, energy and carbon dynamics in vegetation and soil across scales. However, there is yet to be an ecohydrology-specific synthesis assessing the EOFM relevance, application evidence or evaluation requirements under uncertain reference data, scale mismatch and temporal dependence. Here, we develop a framework for determining when EOFMs support interpretable inference and identify a mismatch between EOFMs and ecohydrological requirements. Firstly, an observation-to-inference hierarchy shows that relevance depends on target-specific sensing pathways, spatial-temporal support and traceable uncertainty. Secondly, a meta-analysis shows that pretraining is dominated by reflected optical and active-microwave data, with sparse thermal coverage and no passive-microwave-emission sources. Thirdly, our synthesis of ecohydrological applications finds strongest support for spatial context, label-efficient adaptation and hybrid workflows. Evidence declines with inference depth; independent validation of fluxes, coupled dynamics, event trajectories, calibrated uncertainty and decision benefits remains sparse. Fourthly, our benchmark audit finds stronger coverage of fair adaptation and reproducibility in general EOFM suites, and of process targets, direct reference evidence and distribution shifts in ecohydrological evaluations; physical consistency and uncertainty remain weakly assessed. These findings motivate a process-aware framework aligning EOFM design and evaluation with the target variable, observation pathway and process timescale, supporting trustworthy monitoring and interpretation of coupled water, energy and carbon dynamics.
I. INTRODUCTION · II. INFERENCE OF ECOHYDROLOGY · A. A Hierarchical Framework
Ecohydrology inference links EO measurements with physical relationships, models and contextual information to estimate coupled water, energy and carbon dynamics across scales. The hierarchical framework shows that inference moves from calibrated signals and retrievals to estimated states, fluxes, mechanisms, events and decisions, with increasing dependence on reference evidence, spatial-temporal support and process constraints.
- I. INTRODUCTION: Ecohydrology examines how water availability and movement interact with soils, vegetation and the atmosphere to shape ecological processes across space and time.These interactions partition precipitation and stored soil water among runoff, evaporation and transpiration, divide available energy between sensible and latent heat, and couple stomatal water loss with photosynthetic carbon uptake.
- I. INTRODUCTION: Data-driven EO methods learn nonlinear spatiotemporal relationships for retrieval, prediction, gap filling, data fusion and flux upscaling, but performance is sensitive to temporal memory, physical constraints and distribution shifts.Shifts can occur across sensors, regions, ecosystems, scales, hydroclimates and event magnitudes; process attribution requires independent process evidence, although broader geographic coverage and transfer learning can reduce labelling requirements.
- I. INTRODUCTION: EOFMs learn reusable representations from extensive heterogeneous EO data and adapt them through linear probing, fine-tuning, parameter-efficient adapters or prompting.Their multitemporal, multispectral, radar–optical and generative multimodal capabilities are relevant because complementary sensors provide partial information about coupled water, energy and carbon dynamics.
- I. INTRODUCTION: The paper asks when EOFMs can provide reliable and scientifically interpretable information about coupled water, energy and carbon dynamics across space and time.It frames this through observation pathways, EOFM capabilities, demonstrated applications and evaluation requirements.
- II. INFERENCE OF ECOHYDROLOGY: Ecohydrology inference combines measurements with physical relationships, models and contextual information to estimate states, fluxes, mechanisms or decision-relevant conditions at specified spatial and temporal support.Inference also transfers information across observational supports because point measurements and flux-tower footprints sample different portions of heterogeneous landscapes than EO pixels.
- A. A Hierarchical Framework: The framework links biophysical organisation from leaves to the biosphere with inference progressing from measured signals through retrievals and estimated states or fluxes to mechanisms, events and decisions.The tiers are calibrated signals (T0), retrievals and indices (T1), estimated states and fluxes (T2), and mechanisms, events and decisions (T3).
- A. A Hierarchical Framework: Tier 1 retrievals and indices transform sensor signals into quantities such as vegetation indices, LAI, LST, surface SM and SIF, while Tier 2 combines retrievals with forcing, spatial transfer and model structure.Tier 2 estimates plant or canopy water status, biomass, root-zone SM, ET, GPP and NEE, with sensing pathways differing in depth and vegetation sensitivity.
- A. A Hierarchical Framework: Deeper inference adds uncertainty from calibration, radiative transfer, retrieval inversion, meteorological forcing, modelling and mechanistic interpretation, while temporal support must resolve antecedent conditions, timing, lags and nonlinear interactions.Agreement with a retrieved or model-derived product establishes fidelity to that product; recovery of an underlying state, flux or mechanism requires independently matched reference evidence and biophysical and temporal support.
B. EOFM for Ecohydrology · III. META-ANALYSIS OF EOFMS
EOFMs connect ecohydrological inference to reusable representations learned from multisource EO, with relevance determined by whether they preserve target-specific spatial, temporal and cross-modal information. The meta-analysis examines the field’s growth, multimodal composition and spatiotemporal design through publications available by 31 July 2026.
- B. EOFM for Ecohydrology: Representation learning links EO archives to transferable statistical structure in coupled water, energy, vegetation and carbon dynamics.Masked reconstruction, cross-modal prediction and contrastive alignment derive supervisory signals from the data themselves.
- B. EOFM for Ecohydrology: AlphaEarth Foundations learns a time-conditioned 64-dimensional embedding field from spatially and temporally irregular multisource data for downstream tasks with sparse labels.The example reconciles irregular multisource inputs or prediction targets before ecohydrological targets are introduced.
- B. EOFM for Ecohydrology: Early applications indicate that EOFM embeddings can support ecohydrological prediction when combined with dynamic observations and forcing.The supplied example reports improved station-held-out LST recovery after fusion with GOES-18 observations and additional inputs.
- B. EOFM for Ecohydrology: Existing EOFMs learn reusable representations through radar–optical alignment, multimodal targets, wavelength-conditioned transfer, generative translation and heterogeneous-measurement assimilation.These design strategies expose multiple EO streams for downstream ecohydrological inference.
- B. EOFM for Ecohydrology: Representation relevance depends on encoding spatial organisation and temporal evolution in forms aligned with the target process.Capability assessment must examine reconciliation of asynchronous, scale-mismatched sources and preservation of spatial context and temporal memory.
- B. EOFM for Ecohydrology: Capability assessment must determine whether relevant information is available at inference and whether representations support adaptation through frozen embeddings, linear probes and parameter-efficient methods.Input catalogues identify model exposure, but do not by themselves establish downstream capability.
- III. META-ANALYSIS OF EOFMS: EOFМ research has expanded rapidly since 2021 alongside more reviews, evaluations and model releases using multiple EO data pathways.The literature includes conference proceedings, preprints and model documentation, motivating broad search coverage.
- III. META-ANALYSIS OF EOFMS: The meta-analysis characterises EOFM growth, multimodal composition and spatiotemporal design using publications available by 31 July 2026.Annual publication counts use each work’s earliest public scholarly release, while the 2026 record is partial through 31 July.
A. Overall Trends in EOFM Development … 1) Input Pathways:
EOFMs are expanding rapidly across publication volume, multimodal inputs, and spatiotemporal designs. Across 60 eligible releases, reflected optical and microwave pathways dominate pretraining, while passive-microwave-emission and SIF pathways are absent.
- A. Overall Trends in EOFM Development: EOFMs are developing across three related characteristics: publication growth, multimodal development, and spatiotemporal diversity.These characteristics structure the synthesis of overall EOFM development.
- 1) Publication Growth:: Model releases increased from 1 in 2021 to 21 in 2025, while annual publications rose from 2 to 33.The partial 2026 record already contains 10 model releases and 17 total publications through 31 July.
- 2) Multimodal Development:: The share of releases using at least two EO data pathways increased from 0% in 2021 to 66.7% in 2025 and 70.0% in partial 2026.Optical–microwave and broader multisensor or ancillary configurations became more prominent after 2023; early percentages require caution because release counts were small.
- 3) Spatiotemporal Diversity:: Across 49 releases, EOFM designs span fine-resolution landscape information to kilometre-scale geostationary data and range from snapshots to explicit time series.The mapped dimensions include pre-training sampling distance, temporal support, EO pathway group, and reported parameter count.
- B. Quantitative Analysis of EOFMs: The quantitative analysis covers 60 eligible EOFM releases and evaluates input pathways, scale, adaptation and outputs, and ecohydrological relevance.Applications comprise downstream evaluations reported alongside the original model releases.
- 1) Input Pathways:: Reflected optical data occur in 57 releases (95.0%), microwave data in 34 (56.7%), ancillary inputs in 15 (25.0%), and thermal data in 5 (8.3%).The corpus contains 0 releases with an explicit passive-microwave-emission or SIF pre-training pathway.
2) Spatiotemporal and Model Scale: … C. Synthesis of Model-Level Capabilities
The 60 EOFM releases span heterogeneous spatial, temporal, parameter, adaptation and output scales, but direct ecohydrological evaluation remains limited and uneven. The clearest relevance comes from a small set of models evaluating ecohydrological variables with field, flux, product or proxy references, while process support and inference quality require application-level assessment.
- 2) Spatiotemporal and Model Scale:: Parameter counts have a median of 300 million, an interquartile range of 88–600 million, and a range from 0.402 million to 14.7 billion.Sampling distance and parameter count describe input and model scale; process support and ecohydrological skill require empirical evaluation.
- 3) Adaptation and Outputs:: Fine-tuning is reported for 55 releases (91.7%), while linear probing, frozen-encoder or frozen-embedding use, parameter-efficient fine-tuning and zero-shot evaluation are less common.Downstream evaluations concentrate on segmentation and classification, followed by regression and change detection.
- 4) Ecohydrological Relevance:: Direct ecohydrological evaluation is reported for 6 releases (10.0%), including four releases (6.7%) using independent field, in situ, or flux-reference data.Examples include surface soil moisture, soil organic carbon, water chlorophyll, live fuel moisture content, gross primary production, in situ soil moisture and observed streamflow.
- 4) Ecohydrological Relevance:: Two releases (3.3%) use product or proxy references, while one release provides flux-tower-referenced evaluation of a partitioned carbon flux.The eligible corpus contains no independent flux-tower validation of evapotranspiration.
- 4) Ecohydrological Relevance:: The other 54 releases (90.0%) establish indirect relevance through classification, segmentation, change detection, retrieval or related land-surface tasks.Reference variables derived from retrievals, reanalysis or models carry source assumptions, smoothing, spatial and temporal support, and uncertainty into downstream evaluation; no release evaluates NEE directly.
- C. Synthesis of Model-Level Capabilities: The 60 releases form a heterogeneous design space across backbone architecture, parameter scale, observation pathways, spatial-temporal support, adaptation and downstream outputs, with uneven coverage of ecohydrologically important observations and outputs.AlphaEarth Foundations, HySens and Mini-JEPAs provide the clearest direct relevance because each evaluates an ecohydrological variable; application-level assessment is needed to connect model capabilities to observability, reference independence and validation design.
IV. CAPABILITIES OF EOFMS ACROSS ECOHYDROLOGICAL TASKS … 1) LST:
The application-level evidence spans 71 heterogeneous EOFM records, progressing from retrievals and indices toward mechanisms, events and decision-relevant outcomes. T1 evidence is concentrated in LST and SM, where representations provide mixed benefits and require independent, shift-aware validation.
- A. Overview of EOFM Applications: 71 eligible primary-evidence records form a broad but uneven corpus spanning agriculture, hazards, ecosystems, hydrology and other application families.The audit extracted model, application, representation, reference, validation, baseline and limitation information.
- A. Overview of EOFM Applications: Application evidence is assessed across T1 retrievals and indices, T2 estimated states and fluxes, and T3 mechanisms, events and decision-relevant outcomes.The synthesis compares EOFM representations with observations, covariates and baselines while examining reference support, validation and distribution shifts.
- B. T1: Retrievals and Indices: T1 studies focus mainly on LST and SM, testing EOFM contributions alongside contemporaneous observations, handcrafted predictors and task-specific models.These applications remain comparatively close to EO observables and include satellite recovery, field spectroscopy and gridded-product comparisons.
- 1) LST:: LST recovery combined AlphaEarth embeddings with GOES-18 predictors, reducing leave-one-station-out RMSE from 3.11 to 2.87 K, with the largest improvement under cloud.The study recovered 5-min LST at 53 Hawai‘i Mesonet stations, indicating complementary spatial context from embeddings.
- 1) LST:: Independent LST retrieval requires separate reference observations because gridded products contain their own retrieval and model histories.Operational tests should compare embeddings with thermal harmonisation, conventional gap filling and local ancillary variables across station, seasonal, heat-island and extreme-heat holdouts.
- 1) LST:: HySens provides rare direct evidence that wavelength-aware pretraining supports small field-spectroscopy datasets for bare-soil SM, but evaluation remains limited to snapshot measurements.The study explicitly compared scratch, frozen and fine-tuned regimes and also tested co-located SOC and water-chlorophyll targets.
- 1) LST:: SM results range from negligible to small incremental gains: Prithvi features changed R2 from 0.514 to 0.515, while AlphaEarth added ∆R2 = 0.031 before regional holdout degradation.The broader evidence also includes task-specific SAR–optical–weather models and dynamic global in situ evaluation, though architecture, inputs and pretraining effects can be difficult to separate.
C. T2: Estimated States and Fluxes … 4) GPP:
T2 evidence shows EOFMs most consistently provide transferable spatial context for estimated states and fluxes, while stronger reference data, temporal forcing, local calibration and uncertainty modelling remain necessary. Support is strongest for hybrid, label-efficient applications but weakens as inference requires more assumptions and direct process validation.
- C. T2: Estimated States and Fluxes: T2 applications cover latent states, ecosystem properties and fluxes, with EOFM representations most consistently providing transferable spatial context.These applications require stronger reference and modelling assumptions than T1 retrievals, while temporal forcing, local observations and adaptation remain important.
- 1) Biomass:: Biomass has the largest evidence base, with transferable representations supporting LiDAR-referenced canopy height, Finnish BioMassters AGB and agroforestry stocking indices.TESSERA embeddings matched strong task-specific baselines with high label efficiency, while attentive neural processes improved cross-biome GEDI prediction calibration.
- 1) Biomass:: R2 = 0.79 and 0.82 were achieved by a northeastern United States LiDAR–inventory–embedding workflow after growth-adjusted temporal augmentation.The embedding-specific contribution and uncertainty in the adjusted reference require separate evaluation.
- 1) Biomass:: Biomass gains are conditional: spectral indices outperformed AlphaEarth embeddings in one Andean forest, and pretraining advantages could disappear with abundant multimodal supervision.Performance depends strongly on canopy-structure information, LiDAR support, local calibration, adaptation and uncertainty modelling.
- 2) Canopy Chlorophyll, Nitrogen and Phosphorus:: AlphaEarth channel subsets retrieved canopy chlorophyll, nitrogen and phosphorus across 24 approximately one-hectare maize plots, demonstrating label efficiency and vegetation information.Robust nutrient monitoring across regions, seasons and canopy structures requires broader evidence than this study provides.
- 2) Canopy Chlorophyll, Nitrogen and Phosphorus:: R2 = 0.92 and RMSE= 0.37 mm d−1 were achieved for temporally held-out FLUXNET urban ET, versus R2 = 0.56 over Shenzhen observations.A temporal Transformer combined annual AlphaEarth embeddings, Sentinel-2 NDVI and daily meteorology; meteorology supplied daily variability and embeddings represented spatial heterogeneity.
- 4) GPP:: R2 = 0.81 was achieved by fused Prithvi–MERRA-2 GPP prediction across 37 sites, compared with 0.75 for a same-input ResNet.The leave-one-year-out result supports hybrid prediction from EO representations and meteorological forcing, subject to uncertainty from NEE partitioning and tower-footprint mismatch.
5) Groundwater: … 1) Crop Type, Stress, Tillage and Yield:
Across groundwater, vegetation, soil, streamflow, water-quality and agricultural applications, EOFMs provide useful contextual representations and label-efficient prediction, but transfer, causal inference, uncertainty and decision relevance remain constrained by hydrogeology, phenology, scale mismatch and retrospective evaluation.
- 5) Groundwater:: Groundwater depth prediction reached R2 = 0.793 from AlphaEarth embeddings and TabPFN across 87 wells, but regional transfer remained unresolved.In Danish peatlands, embedding-only prediction was 3% less accurate than an expert-covariate model, while topography recovered part of the deficit and improved responses under synthetic wet and drained conditions.
- 6) LFMC:: For LFMC, fine-tuned time-series representations improved over random initialisation, while a task-specific random forest remained stronger across 1,578 geographically separated field samples.A multimodal Galileo encoder subsequently produced spatially complete 10-m LFMC maps with more than 20% lower RMSE than random initialisation.
- 7) SOC, Soil N and pH:: Soil-property studies favoured embeddings combined with environmental covariates or local reference data rather than embedding-only models.The Danish peatland embedding-only SOC model was 6% less accurate than the expert-covariate model, whereas AlphaEarth outperformed bare-soil and vegetation-index approaches across more than 1,800 cropland topsoil samples.
- 8) Streamflow:: Streamflow studies used EOFMs mainly as learned catchment descriptors alongside meteorological forcing and hydrological memory, with gains declining for environmentally dissimilar donor basins.Water-balance and regime diagnostics remain necessary to assess causal relationships among storage, runoff generation and groundwater connectivity.
- 9) Water-Quality Indicators:: Across eight South Korean river-network water-quality indicators, gain-ranked AlphaEarth channels achieved leave-one-year-out correlations of 0.62–0.89, but spatial transfer was not evaluated.For seasonal farmland-groundwater nitrate, measured hydrochemistry exceeded embedding-based predictors by approximately 10–20% in R2, while coordinate effects and shoreline artefacts raised concerns.
- D. T3: Mechanisms, Events, and Decision-Relevant Outcomes: T3 applications span crop, ecological, hazard and event outcomes, but evidence mainly concerns retrospective mapping or classification under geographic, temporal, event or sensor holdouts.Prospective decision benefit and causal process evidence remain limited in the available evaluations.
- 1) Crop Type, Stress, Tillage and Yield:: County-aggregated AlphaEarth embeddings predicted corn and soybean yield with mean R2 values of 0.825 and 0.814, respectively, while purpose-built EO predictors transferred more effectively across ecoregions and scales.Strong residual spatial autocorrelation and retrospective designs limit prospective interpretation, and latent-space drift did not reliably predict yield error under held-out drought.
- 1) Crop Type, Stress, Tillage and Yield:: Retrospective AgriFM stress alerts had median lead times of 8, 13 and 24 days for satellite-only, late-fusion and AgriFM models, respectively.The historical drought and heat cases demonstrate earlier detection under multimodal fusion, but aggregated or commercially supplied reference data do not establish field-scale operational skill.
2) Ecological Traits and Composition: … 5) Landslide Susceptibility:
Across ecological, flood, glacial-lake, and landslide applications, EOFMs support label-efficient mapping, spatial transfer, and hybrid workflows that combine foundation-model context with local detail. Deeper process inference, including event evolution, timing, runout, and impacts, remains less directly supported by the reported evidence.
- 2) Ecological Traits and Composition:: BotaCLIP improved plant, butterfly, and soil-trophic-group prediction by aligning DOFA with botanical relevés, indicating benefits from explicit domain information.The result links ecological semantic performance to domain-specific information rather than representation alone.
- 2) Ecological Traits and Composition:: Under cross-year transfer, weighted F1 declined by 15% for AlphaEarth and 9% for TESSERA, with the largest effects on rare species.Both embeddings supported label-efficient tree-species mapping before transfer degradation was assessed.
- 3) Flood Inundation:: TerraMind global-event fine-tuning improved flood-inundation accuracy and precision, whereas U-Net retained higher recall.The comparison shows that model advantages depended on the evaluation metric.
- 3) Flood Inundation:: Prithvi-CAFE improved held-out-site IoU by combining Prithvi adapters with a local CNN attention branch, demonstrating the contribution of complementary local detail.The evaluation specifically tested held-out-site performance.
- 4) Glacial-Lake Extent:: TerraMind attained out-of-domain IoU 0.931 for glacial-lake mapping under spatially blocked evaluation.Cryo-Bench additionally found that the strongest model varied with sensor, cryosphere task, and tuning strategy.
- 4) Glacial-Lake Extent:: Glacial-lake-outburst-flood early warning requires lake evolution, dam stability, triggering processes, and downstream exposure information beyond lake-extent delineation.This requirement distinguishes spatial mapping from early-warning inference.
- 5) Landslide Susceptibility:: Adding Clay context to a CNN increased Landslide4Sense F1 from 59.9% to 64.5 ± 1.8%, while Clay alone underperformed U-Net.The evidence supports EOFM representations as contextual inputs within susceptibility workflows.
- 5) Landslide Susceptibility:: A separate AlphaEarth landslide study remained dependent on conventional conditioning factors, while event timing, runout, and impact assessment require additional evidence.The reported findings support hybrid susceptibility workflows but do not establish complete event-process inference.
6) Wildfire Burned Area and Severity: … A. Evaluation Scope of Existing EOFM Benchmarks
Existing EOFM evidence is strongest for spatially contextual, label-efficient applications, while support declines for deeper ecohydrological inference and rigorous validation. Benchmarking therefore requires target-matched evaluation of reference independence, temporal dynamics, environmental shifts, physical consistency, uncertainty and reproducibility.
- 6) Wildfire Burned Area and Severity:: Across 3,820 burned-area events, LoRA generalized more effectively than full or decoder-only fine-tuning while updating less than 1% of parameters; Prithvi-v2 was strongest overall.TESSERA approached a strong spectral burned-area baseline in Portugal, but event-date spectral features remained preferable when suitable imagery existed.
- E. Synthesis of Application-Level Evidence: Among 71 application records, fine-tuning or PEFT was most frequent (22 records), followed by feature extraction (19) and mixed or benchmark workflows (16).Fusion or similarity accounted for 8 records, probing or reconstruction 4, and retrieval 2.
- E. Synthesis of Application-Level Evidence: Reported comparisons were positive in all five hydrological-modelling records, 86% of ecosystems, land-cover and cryosphere records, and 80% of disturbance and hazards records.Water, energy and climate states had the lowest positive share (50%), with 40% neutral and 10% null or negative.
- E. Synthesis of Application-Level Evidence: Application-family proportions cannot support pooled effect estimates or general model rankings because family sizes, targets, metrics, baselines, reference data and validation designs were heterogeneous.Vegetation, carbon and fuels was the only other family containing a null or negative result (8%).
- E. Synthesis of Application-Level Evidence: Evidentiary support generally decreased with inference depth: map reproduction and sparse-label spatial regression were broadest, whereas flux, geographic-transfer and event-trajectory evidence remained limited.Field-referenced estimation was emerging for VWC, LFMC, SM, groundwater, water quality and SOC; flux evidence included a partitioned GPP tower example, ET product reproduction and one hybrid daily urban ET application.
- V. BENCHMARKING EOFMS FOR ECOHYDROLOGY: A credible ecohydrological benchmark aligns targets with spatial-temporal support, traces reference uncertainty, evaluates environmental and sensor shifts, and tests water, energy and carbon consistency.The inference hierarchy distinguishes T1 retrievals and indices, T2 estimated states and fluxes, and T3 mechanisms, events and decision-relevant outcomes.
- A. Evaluation Scope of Existing EOFM Benchmarks: The audit covered 20 comparative studies, split evenly between 10 general EOFM benchmark suites and 10 ecohydrological evaluations, using eight dimensions scored from 0 to 2 and rescaled to 0–100%.The dimensions cover process target, reference directness, temporal dynamics, transfer or shift, physical consistency, fair adaptation, uncertainty or reliability, and reproducibility.
1) General EOFM Benchmark Suites: … 1) Established Practices for Standardisation:
General EOFM suites emphasise fair adaptation and reproducibility, whereas ecohydrological evaluations prioritise process targets and distribution shifts. Future benchmarks should integrate these complementary strengths while addressing shared gaps in physical consistency and uncertainty.
- 1) General EOFM Benchmark Suites:: General suites reach 90% coverage for fair adaptation and 85% for reproducibility, establishing strong comparison protocols.These suites provide standardised comparison practices across EOFM evaluations.
- 1) General EOFM Benchmark Suites:: MMEarth-Bench combines explicit process targets, controlled adaptation, reproducible implementation and geographic transfer tests.It has the broadest profile among the general suites, with partial coverage of reference directness.
- 2) Ecohydrological Evaluations:: Ecohydrological evaluations reach 100% coverage for transfer or shift, 90% for process targets, 85% for fair adaptation and 70% for temporal dynamics.Reference directness reaches 60%, while reproducibility, physical consistency and uncertainty or reliability remain lower at 40%, 25% and 15%, respectively.
- 2) Ecohydrological Evaluations:: Benchmark strengths are distributed unevenly across eight dimensions, with complementary coverage and shared diagnostic gaps motivating future development priorities.The two evaluation groups contribute different strengths rather than a uniformly complete assessment.
- B. Priorities for Future Ecohydrological Benchmarks: Future EOFM benchmarks should combine established standardisation practices, complementary strengths and diagnostic-gap development across the eight evaluation dimensions.This three-part agenda is mapped from current coverage patterns and evidence-linked priorities.
- 1) Established Practices for Standardisation:: Fair adaptation is the most consistently established dimension, reaching 90% coverage in general suites and 85% in ecohydrological evaluations.General suites also reach 85% coverage for reproducibility through shared protocols, fixed splits and open benchmark assets.
2) Complementary Strengths for Integration: … 2) Emerging Data Infrastructure for EOFM Evaluation:
The synthesis identifies complementary benchmark strengths, diagnostic gaps, and data-infrastructure priorities for evaluating EOFMs in ecohydrology. Future benchmarks should combine reproducibility and controlled comparison with process-centred targets, direct references, shift-aware testing, physical checks, uncertainty, and harmonised provenance.
- 2) Complementary Strengths for Integration:: Process targets, direct references, temporal dynamics, transfer or shift, and reproducibility show complementary coverage across general suites and ecohydrological evaluations.Coverage is 20% versus 90% for process targets, 5% versus 60% for reference directness, 20% versus 70% for temporal dynamics, 45% versus 100% for transfer or shift, and 85% versus 40% for reproducibility.
- 2) Complementary Strengths for Integration:: Target-matched benchmarks should define the physical quantity, units, inference tier, spatial-temporal support, reference layers, temporal tests, and relevant distribution shifts.Temporal tests should resolve seasonal phase, event onset, antecedent memory, lag, and recovery; shift tests should hold out sites, basins, biomes, years, sensors, scales, regimes, and extremes.
- 3) Diagnostic Gaps for Priority Development:: Physical checks and uncertainty remain limited, with coverage of 0% versus 25% and 5% versus 15% in general suites and ecohydrological evaluations.Future evaluations should connect predictive accuracy with physical bounds, water–energy–carbon closure, calibrated uncertainty, reproducible implementation, and multiple model baselines.
- 3) Diagnostic Gaps for Priority Development:: Implementing these priorities requires harmonised, diverse infrastructure preserving observation provenance, spatial-temporal support, versioned splits, and access to evaluated models or embeddings.These requirements convert the coverage assessment into a practical design for future benchmark development.
- C. Multiscale Data Infrastructure for Ecohydrological Benchmarking: Ecohydrological benchmarking can integrate established process-evaluation resources with emerging EOFM datasets linking harmonised multimodal EO products to ecosystem observations.This integration connects EOFM assessment with process evidence across spatial and temporal scales.
- 1) Established Data Infrastructure for Model Evaluation:: Established infrastructure includes half-hourly-to-decadal eddy-covariance observations and PLUMBER2, which harmonises 1,040 site-years from 170 flux towers for land-model evaluation.STEMMUS-SCOPE and related resources extend site-scale evaluation with hydrological, photosynthetic, radiative, flux, and soil-moisture variables.
- 1) Established Data Infrastructure for Model Evaluation:: Agricultural benchmarking is constrained by scarce public field- and sub-field-scale references, leaving recent yield benchmarks heavily dependent on aggregated statistics.AlphaEarth evaluation used county-level USDA yield records, while CropClimateX provides county-level yield and farm-management variables.
- 2) Emerging Data Infrastructure for EOFM Evaluation:: WorldTensor supplies harmonised planetary context, whereas FluxCubes30 provides 30-m, tower-centred EO data for process-focused evaluation of landscape-scale carbon dynamics.WorldTensor aligns hundreds of environmental and socioeconomic variables on a common 0.25° annual grid; FluxCubes30 contains GPP and EVI cubes with pixel-level uncertainty for 404 landscapes from 1999 to 2025.
VI. DISCUSSION AND PROSPECTS … D. Target-First Benchmarking and Future Development
The discussion revisits the four research questions to define observability limits, assess EOFM representation and application evidence, and translate remaining gaps into target-first benchmarking requirements. It proposes process-aware development grounded in target-specific pathways, uncertainty, validation, and evolving evaluation domains.
- A. Observability, Inference Contracts and Uncertainty: EOFM claims require target-specific observation pathways constrained by observation physics, reference uncertainty, spatial scale, and process constraints.These factors define observability limits from calibrated measurements through properties, states, fluxes, mechanisms, and decisions.
- A. Observability, Inference Contracts and Uncertainty: Inference contracts should document each target’s physical quantity, observation pathway, processing level, spatial-temporal support, reference basis, scale matching, assumptions, and intended interpretation.A quantitative uncertainty budget should accompany the contract and preserve identifiable sources of uncertainty.
- B. Representation Capacity and Trustworthiness: Current EOFMs provide diverse modalities, spatial-temporal supports, representations, and adaptation routes, but event-resolving support, complementary non-optical pathways, lineage transparency, and retained process information remain limited.Reflected optical observations dominate the landscape, while SAR provides most microwave coverage.
- B. Representation Capacity and Trustworthiness: Trustworthiness should be a primary EOFM design requirement, encompassing robustness, calibrated uncertainty, interpretability, reproducibility, efficient adaptation, and human- and environment-centred use.The discussion identifies incomplete access to model weights and minimal releases as an operational trustworthiness gap.
- C. Ecohydrological Applications and Process-Targeted Representations: Application evidence most consistently supports reusable contextual representations, label-efficient adaptation, and multimodal fusion, while null and negative results constrain claims of process-specific benefit.Examples include negligible soil-moisture improvement over engineered predictors and spectral indices outperforming AlphaEarth embeddings for field-referenced biomass in one forest.
- C. Ecohydrological Applications and Process-Targeted Representations: Future models should expose process-targeted inputs, outputs, and validation, with readiness requiring independent support for spatial detail, event timing and recovery, and coherent coupled-process behaviour under intended shifts.BERTH illustrates a shared temporal architecture producing daily 30-m estimates of ET, precipitation, soil moisture, and runoff from optical, forcing, and terrain inputs.
- D. Target-First Benchmarking and Future Development: Target-first benchmarking should connect model outputs to ecohydrological claims through explicit observation and reference pathways while preserving target identity, observational identity, and domain validity.Evaluation must retain the target’s quantity, units, inference tier, sensor, processing level, reference directness, and spatial-temporal support.
- D. Target-First Benchmarking and Future Development: Benchmark development should proceed in stages: near-term comparable evidence, medium-term process and distribution-shift diagnostics, and long-term versioned evaluation as data, models, and use contexts evolve.The staged agenda treats benchmarking as an evolving specification of scientific validity.
1) Near Term: Comparable Protocols: … VIII. CODE AND DATA AVAILABILITY
The paper proposes a staged agenda for reproducible, process-aware EOFM evaluation, progressing from comparable protocols and broader process and shift coverage to persistent benchmarking. Its conclusions identify current evidence gaps and provide resources for continued evaluation and reuse.
- 1) Near Term: Comparable Protocols:: Near-term evaluations should use shared protocols with reusable weights, variance reporting, common harnesses, and controls for data, architecture, and algorithm effects.PANGAEA and REOBench additionally contribute stratified subsets, controlled runs, versioned metadata, and maintained releases.
- 1) Near Term: Comparable Protocols:: Ecohydrological records should document targets and units, inference tier, observation and reference pathways, spatial and temporal support, and provenance.
- 2) Medium Term: Process and Shift Coverage:: Medium-term benchmarks should expand sensor, task, regional, uncertainty, and real-world-shift coverage across soil moisture, biomass, streamflow, crop stress, and flood inundation.Flux networks and land-model evaluation datasets provide references for coupled water, energy, and carbon exchange.
- 3) Long Term: Persistent Evaluation:: Long-term benchmarking should become a maintained evaluation programme that tracks evidence as models, data, and environmental conditions evolve.EarthShift frames benchmarking as a living testbed, while REOBench and PANGAEA provide versioning, preservation, and extensible protocols.
- VII. CONCLUSION: EOFMs support credible ecohydrological process inference only when reusable representations are linked to sensing pathways, spatial and temporal support, process timescales, and reference uncertainty.The review applies this observation-to-inference framework to EOFM designs, ecohydrological applications, and benchmark studies.
- VII. CONCLUSION: Current evidence is strongest for spatial context, label-efficient adaptation, and hybrid workflows, but weakens with inference depth and remains sparse for independent process validation, calibrated uncertainty, and decision benefit.Pretraining is dominated by reflected optical and active-microwave data, with sparse thermal coverage and no explicit passive-microwave-emission or SIF sensing data sources.
- VII. CONCLUSION: Future progress requires process-aware inference contracts and uncertainty budgets aligned with inferred variables, observation pathways, process timescales, provenance, failure conditions, and calibrated uncertainty.Process-targeted hybrids should integrate EO, meteorological forcing, in situ data, and process constraints while preserving event dynamics and water–energy–carbon relations.
- VIII. CODE AND DATA AVAILABILITY: The collated resources will be distributed through the EOFM4EcoHydrol GitHub repository.Repository URL: https://github.com/yuyi13/EOFM4EcoHydrol.