Source-linked AI summary
Bridging short- and medium-range weather forecasting with machine learning
Timothy A. Smith, Mariah Pope, Sergey Frolov, Brett Basarab, Daniel Abdi, Paul Madden, Isidora Jankov
TL;DR
NOAA’s short- and medium-range forecasts are produced by separate systems, motivating a unified model that combines regional detail with global coverage. Nested-EAGLE nests a 6 km CONUS HRRR state in a 0.25° global GFS state and improves near-surface CONUS forecasts, while precipitation amounts remain less skillful than HRRR despite accurate longer-lead storm locations.
Problem
Separate forecast systems provide different short- and medium-range information, motivating a single system that synthesizes regional impacts and global weather.
Method
Nested-EAGLE combines HRRR data over CONUS with GFS data elsewhere in a 6 km regional, 0.25° global ML weather model.
Results
Nested-EAGLE significantly lowers near-surface CONUS RMSE versus NOAA’s operational systems and a GFS-only ML baseline, while remaining competitive aloft and outside CONUS.
Takeaways & Limitations
Nested-EAGLE demonstrates that one model can bridge short- and medium-range forecasting, with regional training data improving CONUS near-surface skill and storm-location forecasts at longer leads.
Takeaways & Limitations
The deterministic 6 hour, 0.25°/6 km model is too coarse to resolve convective scales, and its precipitation amounts are less skillful than HRRR.
Abstract
from arXiv · showhide
The National Oceanic and Atmospheric Administration (NOAA) employs independent prediction systems for distinct forecast products. While some separation is practical, we argue that combining short- and medium-range weather into a single prediction system would provide the public with a useful distillation of global weather and its impacts. To this end, we present Nested-EAGLE (Experimental Artificial intelligence Global and Limited-area Ensemble): a 0.25° global weather model with a 6 km refinement over the Contiguous United States (CONUS). The model achieves significantly lower mean-squared error in near-surface and low-level quantities over CONUS compared to NOAA's Global Forecast System and High-Resolution Rapid Refresh (HRRR), while remaining competitive throughout the rest of the global atmosphere. We show that the skill gains for near-surface fields stem from incorporating high-resolution regional analysis data into training through the nesting process. Forecasts of precipitation amounts are less skillful than those from HRRR, owing to deterministic training. However, we show that Nested-EAGLE provides the most accurate forecasts of storm locations at longer leads, despite blurred extrema. Our results motivate future work to extend the skill gains beyond CONUS and improve precipitation representation.
1. Introduction
Nested-EAGLE combines HRRR-informed 6 km forecasting over CONUS with 0.25° GFS-based global forecasting to bridge short- and medium-range weather prediction. The design targets a single system while preserving global medium-range coverage and improving regional skill through nested training data.
- Motivation: NOAA uses distinct systems because forecast products span different applications and timescales, although some separation remains practical.HRRR supports local impacts at one- to two-day leads, while GFS provides synoptic context and two-week outlooks.
- Related work: Machine-learning weather prediction provides a route toward joining global medium-range models with high-resolution regional emulators.Existing regional systems inherit high spatial resolution but commonly remain separate from global forecasting pipelines.
- Related work: Bris demonstrates that an all-in-one ML weather model can provide high-quality forecasts for a region of interest while representing the rest of the globe more coarsely.Nested-EAGLE extends this combined-system idea to CONUS and longer leads.
- Model design: Nested-EAGLE uses a 6 km mesh over CONUS nested inside a 0.25° global grid.Its state space combines HRRR data over CONUS with GFS data elsewhere.
- Model design: The model is trained on HRRR data over CONUS and GFS data elsewhere to combine regional and global forecasting information.It is deterministic, uses a 6 hour time step, and is evaluated to 15 days inside and outside CONUS.
- Limitations: Nested-EAGLE does not yet resolve convective scales and has limitations in representing precipitation because deterministic 6 hour MSE training blurs extremes.The model generally predicts storm locations successfully despite struggling with extreme precipitation amounts.
2. Prognostic Forecast Skill
Nested-EAGLE is evaluated against NOAA’s GFS and HRRR and a matched GFS-only ML baseline. It improves near-surface CONUS skill through HRRR-informed training, remains competitive aloft and outside CONUS, and rapidly corrects less accurate initial conditions.
- Evaluation: Nested-EAGLE is compared with GFS, HRRR, and ML-GFS-Base using 15 day forecasts and RMSE against in situ observations.The test protocol contains 293 forecast initializations per lead time.
- Forecast Skill Over CONUS: Nested-EAGLE generally has the lowest CONUS RMSE, especially for near-surface quantities, while remaining competitive through the atmosphere to 15 days.The comparison includes HRRR to 2 days and the other models to 15 days.
- Forecast Skill Over CONUS: 78 hours: Nested-EAGLE’s one-day skill gap over GFS for 10m wind speed and 2m temperature is about 78 hours.The corresponding gap for 2m specific humidity is 60 hours.
- Forecast Skill Over CONUS: 54 and 66 hours: Nested-EAGLE leads ML-GFS-Base at one day for 10m wind speed and 2m temperature, respectively.The statistically significant lead persists until about 6.5 days for these variables.
- Forecast Skill Over CONUS: Near-surface gains reflect HRRR analysis advantages, including higher resolution, hourly assimilation, a more advanced land model, and better U.S.-centric observations.HRRR analysis improves near-surface fields over GFS by roughly 30–48%, and the authors attribute Nested-EAGLE’s gains to propagating these improvements through training.
- Forecast Skill Over Europe: Beyond CONUS, Nested-EAGLE and ML-GFS-Base show modest gains over GFS for selected European fields but are otherwise statistically indistinguishable.The model does not carry HRRR-derived skill globally, while CONUS gains do not degrade skill elsewhere.
- Disentangling Training and Initial Conditions: Swapped-initial-condition experiments show that the improvements are learned during training and rapidly reappear despite less accurate inference-time initial conditions.For 10m wind speed, convergence occurs within the first 6 hour time step.
3. Precipitation Evaluation Over CONUS
Nested-EAGLE’s precipitation skill is evaluated over CONUS with FSS, separating amplitude representation from storm-location accuracy. Deterministic training blurs precipitation extremes, but Nested-EAGLE becomes strongest at locating events beyond the shortest lead times.
- Evaluation setup: FSS compares 1,426 six-hour precipitation forecasts against AORC over CONUS, with higher scores indicating better spatial and amplitude agreement.The evaluation examines forecast hour and precipitation amplitude, while confidence intervals summarize the forecast ensemble.
- Amplitude skill: HRRR is most skillful across precipitation thresholds, except at 1–2 mm/6h, where Nested-EAGLE has a slight advantage.At 6 hours, Nested-EAGLE is only marginally more skillful than GFS despite its higher resolution.
- Amplitude skill: Deterministic MSE training blurs small-scale precipitation features, diminishing local maxima and reducing skill at higher thresholds.The resulting double-penalty effect contributes to lower precipitation skill relative to HRRR and GFS.
- Location skill: Percentile-based FSS controls for amplitude bias and isolates each model’s ability to capture precipitation-event locations.This framing evaluates the spatial position of events using amplitudes each model can represent.
- Location skill: At 6 hours, Nested-EAGLE and HRRR significantly outperform ML-GFS-Base and GFS, consistent with the benefit of incorporating HRRR data through nesting.The ML models are indistinguishable from their physical counterparts at this lead because 0–6 hour forecasts are training targets.
- Location skill: Beyond 6 hours, Nested-EAGLE has the highest percentile-based FSS, while ML-GFS-Base gradually surpasses GFS and HRRR falls below GFS.All models lose skill with lead time, and ML-GFS-Base approaches Nested-EAGLE’s skill at longer leads.
4. Discussion
The discussion frames Nested-EAGLE as one system bridging global medium-range and regional short-range forecasting. Its main gains occur for near-surface CONUS fields, while deterministic training and coarse resolution constrain precipitation and convective-scale skill.
- Contribution and scope: Nested-EAGLE is a global 0.25° MLWP model with a 6 km CONUS refinement trained on HRRR data nested inside GFS data.The model is intended to combine information otherwise delivered through separate short- and medium-range forecast systems.
- Results: Nested-EAGLE has significantly lower CONUS near-surface RMSE than operational physics-based systems and the GFS-only ML baseline.Its 2m temperature forecasts maintain a significant lead over GFS for 7.5 days and ML-GFS-Base for 6.5 days.
- Results: Nesting produces the largest gains for near-surface CONUS fields, while skill in the upper atmosphere and outside CONUS is similar to the GFS-only ML model.In many cases, both ML models are comparable to GFS in those regions.
- Limitations: Nested-EAGLE’s 6 hour step, 0.25°/6 km resolution, deterministic formulation, and limited convective-scale resolution constrain short-range storm prediction.These constraints manifest as precipitation forecasts that under-represent high-to-extreme amounts and are less skillful than HRRR, though comparable to GFS.
- Future directions: Nested-EAGLE predicts precipitation locations accurately despite blurred amplitudes, motivating direct use of observations in future MLWP development.The discussion links this behavior to training on short-lead HRRR forecasts that remain constrained by data assimilation.
- Future directions: The current deterministic six-hour formulation cannot surpass HRRR’s six-hour skill because HRRR forecasts are its training target.The authors suggest incorporating observations directly into the learning objective.
- Future directions: Observation-constrained losses may reduce the need to nest or combine analyses, but retraining would still be necessary.Future models could operate directly on observational inputs to respond to improved observational constraints.
5. Methods
Nested-EAGLE combines global GFS data with high-resolution HRRR data over CONUS in a nested, autoregressive ML weather model. The methods define its architecture, training, baseline comparisons, evaluation metrics, and uncertainty procedures.
- Model design: Nested-EAGLE maps weather states at t and t−6 hours to a prediction at t+6 hours using an autoregressive deep neural network.Its network contains encoder, processor, and decoder modules.
- Model design: The data space joins GFS fields on a 0.25° latitude–longitude grid with HRRR fields on a CONUS Lambert conformal conic grid.The latent mesh combines a global O96 grid with HRRR nodes coarsened to 24 km resolution.
- Model design: A custom latent mesh with sliding-window attention removed boundary artifacts and reduced computational costs by ∼33% relative to a multimesh graph-transformer processor.Early prototypes using an icosahedral multimesh and graph-transformer processor exhibited artifacts at global–regional boundaries.
- Training: Training used all available HRRR and GFS data since February 2015, without ERA5 pretraining.The dataset was divided into 8 years for training, 1 year for validation, and 1 year for testing.
- Baselines: ML-GFS-Base is a matched single-resolution baseline trained solely on GFS data to isolate the effect of incorporating HRRR data through nesting.Nested-EAGLE additionally reweights the HRRR subdomain to account for 10% of the loss fraction.
6. Data Availability
The study uses openly available GFS and HRRR archives, together with conventional observations accessed through NOAA- and NASA-related data resources.
- Data sources: GFS archives were obtained from NCAR’s Research Data Archive and, for forecasts initialized from January 2021 onward, NOAA’s Open Data Dissemination program via AWS.HRRR archives were accessed through the AWS Registry of Open Data.
- Data sources: Conventional observations used for evaluation came from NNJA and its AI-Ready version accessed through the Brightband Python API.
7. Code Availability
The project releases evaluation scripts, configuration files, related resources, and Nested-EAGLE neural-network weights through public repositories and checkpoints.
- Code and resources: Evaluation scripts and configuration files are available in the NOAA-PSL/nested-eagle GitHub repository.The repository specifies data ingestion, model design, and training configuration.
- Code and resources: Related resources include neural-network weights for Nested-EAGLE and the earlier wxvx evaluation package.
Supporting Information for “Bridging short- and medium-range
The supporting information accompanies the paper on bridging short- and medium-range weather forecasting with machine learning and identifies its authors.
- Document: The supporting information is associated with the paper “Bridging short- and medium-range weather forecasting with machine learning.”
- Authors: The listed authors include Timothy A. Smith, Mariah Pope, Sergey Frolov, Brett Basarab, Daniel Abdi, Paul Madden, and Isidora Jankov.
S1. Extended Prognostic Evaluation
The extended evaluation compares Nested-EAGLE, ML-GFS-Base, and GFS across global and regional domains using RMSE against observations. Lead-time gaps are reported for near-surface and atmospheric fields, with evaluations based on 293 forecasts.
- 293 forecasts underpin RMSE-based lead-time-gap comparisons between Nested-EAGLE and baseline models.Tables S1 and S2 define positive gaps as additional lead time with lower Nested-EAGLE RMSE; statistically insignificant differences are marked gray.
- Global RMSE is evaluated against in situ observations over the globe.
- CONUS RMSE is evaluated over 20°N–55°N and 135°W–50°W at 0.25° resolution, unlike the 6 km evaluation in Figure 2.
- Europe RMSE is evaluated over 35°N–75°N and 25°W–65°E, with more variables shown than in Figure 3.
- Additional regional evaluations cover the Northern Hemisphere, tropics, Southern Hemisphere, and both polar regions.The Northern Hemisphere spans 20°N–80°N, the tropics 20°S–20°N, southern regions 80°S–20°S, polar North 60°N–90°N, and polar South 90°S–60°S.
S2. Model Development Details and Ablation Experiments
Model-development experiments examine data resolution, training assumptions, loss weighting, architecture, optimization, pressure scaling, and the number of initial states. Several choices were selected using coarse-resolution validation experiments, with results assumed to generalize to the final higher-resolution model.
- Development setup: Development experiments used HRRR data regridded to 15 km and GFS data regridded to 1°, with results assumed to carry over to the higher-resolution model.Most experiments used 30,000 optimization iterations and a sliding-window transformer processor; the development model required about 4.5 hours across 8 GPU nodes.
- Development setup: 158 forecasts from February 2023–January 2024 support the sensitivity results, which sometimes use HRRR analysis rather than in situ observations.Only the highlighted hyperparameter changes in each experiment, but the authors assume observed trends generalize despite other model changes.
- S2.2. Attention Window Size: A final attention window of 8,168 latent mesh nodes was chosen from a scaling rule after higher window sizes showed largely statistically insignificant differences.The development rule was N_window := N_mesh/N_proc/2; a window size of 1,080 was too small, while 3,564 was used in the smaller development model.
- S2.3. Training Iterations: 60,000 training iterations performed best, but gains over 30,000 were insignificant, whereas 15,000 appeared too short.The model requires about 10–20% of the training iterations used by some comparable models.
- S2.4. Latent Channels: 512 and 1,024 latent channels showed no statistical difference, despite computational cost scaling linearly with channel count.
- S2.5. Pressure Scaling: Treating all pressure levels identically in the loss function improved performance for variables in the middle and upper troposphere.
- S2.6. Number of Initial States: Two initial states appeared to improve MAE across variables, but differences were usually not statistically significant.Two states were retained because of gains in 500 hPa geopotential-height skill, while one state may be acceptable for many applications during early development.
S3. Attribution of Perturbed 2m Temperature Bias to Orography Differences
The analysis attributes slower 2m-temperature error convergence with GFS initial conditions and static fields to orography differences between HRRR and GFS configurations. Regression correlations rise rapidly with lead time and imply a lapse rate near the standard tropospheric value.
- Orography differences are a leading cause of systematic 2m-temperature errors between the Nested-EAGLE configurations.The comparison contrasts standard Nested-EAGLE with a version using only GFS initial conditions and GFS static variables such as orography.
- By 3 days, Spearman’s ρ² approaches 0.6, while Pearson’s r² saturates near 0.9 after rising rapidly during the first day.Both coefficients relate location-averaged temperature bias to signed HRRR–GFS orography differences.
- Regression-implied lapse rates level off near −6 K km−1, close to the standard tropospheric lapse rate of −6.5 K km−1.
S4. Extended Precipitation Evaluation
Monthly mean precipitation forecasts provide qualitative support for the evaluation of precipitation placement over CONUS. Nested-EAGLE and HRRR place events similarly overall, with muted Nested-EAGLE extrema and some cases of improved feature placement.
- Monthly mean 6 hour precipitation accumulations were forecast at a 24 hour lead for March and August 2023.These visualizations use validation-period forecasts.
- Nested-EAGLE and HRRR show qualitatively similar skill in placing precipitation events over CONUS, except for muted Nested-EAGLE extrema.
- During the illustrated months, Nested-EAGLE places several precipitation features noticeably better than HRRR against the AORC reference.
S5. Sample Forecasts
Sample forecasts examine a strong atmospheric river striking the U.S. West Coast around 10 March 2023. The case supports qualitative assessment of feature resolution across the GFS–HRRR boundary.
- The sample forecasts focus on a particularly strong atmospheric river event that struck the U.S. West Coast around 10 March 2023.The event is presented as societally relevant.
- The case enables qualitative evaluation of how forecast features are resolved across the GFS–HRRR boundary.
- The displayed quantities include 10m wind speed, 2m temperature, 2m specific humidity, and accumulated precipitation.