Source-linked AI summary
The rise of data-driven weather forecasting
Zied Ben-Bouallegue, Mariana C A Clare, Linus Magnusson, Estibaliz Gascon, Michael Maier-Gerber, Martin Janousek, Mark Rodwell, Florian Pinault, Jesper S Dramsch, Simon T K Lang, Baudouin Raoult, Florence Rabier, Matthieu Chevallier, Irina Sandu, Peter Dueben, Matthew Chantry, Florian Pappenberger
TL;DR
Standard NWP forecasting continues to improve but its computational cost constrains resolution and ensemble growth, motivating faster ML-based alternatives. This study compares PanguWeather and ECMWF’s IFS from the same initial conditions using operational forecast verification, finding comparable skill in several global and extreme-event evaluations while identifying smoothness and bias drift as drawbacks.
Problem
NWP improvements are constrained by computational cost, while evidence remains needed on whether ML forecasts can match operational NWP quality and predict extremes.
Method
The study evaluates deterministic PanguWeather forecasts against ECMWF’s IFS in an operational-like setting using shared initial conditions and standard verification tools.
Results
PGW shows good performance for upper-air and surface variables, comparable performance for some extreme events, and similar error-growth sensitivities to NWP forecasts.
Takeaways & Limitations
Data-driven forecasts can provide a promising alternative forecasting paradigm based on ML inference and state-of-the-art analysis and reanalysis datasets.
Takeaways & Limitations
ML forecasts are currently smoother and exhibit an almost linear bias drift with forecast lead time.
Abstract
from arXiv · showhide
Data-driven modeling based on machine learning (ML) is showing enormous potential for weather forecasting. Rapid progress has been made with impressive results for some applications. The uptake of ML methods could be a game-changer for the incremental progress in traditional numerical weather prediction (NWP) known as the 'quiet revolution' of weather forecasting. The computational cost of running a forecast with standard NWP systems greatly hinders the improvements that can be made from increasing model resolution and ensemble sizes. An emerging new generation of ML models, developed using high-quality reanalysis datasets like ERA5 for training, allow forecasts that require much lower computational costs and that are highly-competitive in terms of accuracy. Here, we compare for the first time ML-generated forecasts with standard NWP-based forecasts in an operational-like context, initialized from the same initial conditions. Focusing on deterministic forecasts, we apply common forecast verification tools to assess to what extent a data-driven forecast produced with one of the recently developed ML models (PanguWeather) matches the quality and attributes of a forecast from one of the leading global NWP systems (the ECMWF IFS). The results are very promising, with comparable skill for both global metrics and extreme events, when verified against both the operational analysis and synoptic observations. Increasing forecast smoothness and bias drift with forecast lead time are identified as current drawbacks of ML-based forecasts. A new NWP paradigm is emerging relying on inference from ML models and state-of-the-art analysis and reanalysis datasets for forecast initialization and model training.
A FIRST STATISTICAL ASSESSMENT OF MACHINE LEARNING-BASED WEATHER FORECASTS IN AN OPERATIONAL-LIKE CONTEXT
The passage identifies the authors associated with this work.
- Zied Ben Bouallègue and Mariana C A Clare are listed among the authors.
- Linus Magnusson, Estibaliz Gascón, and Michael Maier-Gerber are also listed as authors.
- Matthieu Chevallier, Irina Sandu, Peter Dueben, Matthew Chantry, and Florian Pappenberger complete the displayed author list.
1 Introduction
Traditional NWP forecasting has steadily improved but faces major computational constraints, while ML-based forecasting offers a faster alternative enabled by reanalysis data. The study introduces an operational-like comparison of PanguWeather with ECMWF’s IFS using common verification methods and shared initial conditions.
- 1 Introduction: Computational and timeliness constraints force operational NWP systems to balance model resolution against ensemble size, limiting further skill improvements.Both higher resolution and larger ensembles are known to improve ensemble forecast skill.
- 1 Introduction: ML-based weather models promise forecasts at much lower computational cost, potentially with increased timeliness and accuracy.
- 1 Introduction: ERA5 provides a continuous, high-quality reanalysis dataset used to train ML models, although most methods discussed train from 1979 because earlier data have lower accuracy.ERA5 spans 1940 to the present and is produced by blending observations with short-range forecasts through data assimilation.
- 1 Introduction: ML forecasts use inductive inference from learned data patterns, raising concerns about unseen extremes and model interpretability.
- 1 Introduction: The study compares PanguWeather with an operational NWP forecast in an operational-like setting, using the same framework and initial conditions.Standard verification techniques are used to assess which forecast attributes can match ECMWF’s IFS.
2 Methodology and experiments
The experiments compare PanguWeather with IFS and ERA5-related forecasts at differing resolutions and initializations. Results show variable- and lead-time-dependent performance, including early advantages from operational analysis initialization and a rapid Arctic error in some PGW fields.
- 2 Methodology and experiments: PGW uses a vision transformer that maps 3D weather fields to forecasts while minimizing latitude-weighted RMSE.
- 2 Methodology and experiments: IFS is evaluated at 9km resolution, ERA5 at 30km, and the study also includes a lower-resolution IFS forecast and a 50-member ensemble.The lower-resolution IFS configuration is close to PGW’s 28km grid spacing.
- 2 Methodology and experiments: PGW is initialized from the operational IFS analysis for a fair comparison, despite being trained on lower-resolution ERA5 data.A complementary PGW_E5 experiment starts PGW from ERA5 reanalysis.
- 2 Methodology and experiments: PGW initialized from operational analysis generally outperforms PGW initialized from ERA5 during roughly the first four forecast days.The exact duration varies by variable and domain.
- 2 Methodology and experiments: IFS_LR falls between ERA5 and IFS in performance, while horizontal-resolution effects become smaller at longer lead times.For T850 at two days, reduced resolution clearly degrades accuracy; at six days for Z500, differences are small.
- 2 Methodology and experiments: PGW performs better than IFS at day 2 for T850 but worse for Z500, especially during summer.PGW also develops a rapidly growing central-Arctic error in Z500 and mean sea-level pressure.
3 Data and a case-study
The study verifies upper-air, surface, and tropical-cyclone forecasts using operational analyses and SYNOP observations across seasonal and case-study settings. Figure 2 compares RMSE behavior for T850 and Z500, while Figure 3 examines a 2m-temperature event at Sodankylä.
- T850 and Z500 forecasts are verified against operational IFS analyses interpolated to 1.5° and aggregated over the Northern Hemisphere.
- Figure 2 presents three-month-averaged RMSE at forecast days 2 and 6 across one year for T850 and Z500.
- The evaluation also assesses 2m temperature against European SYNOP observations and includes tropical-cyclone forecast verification.
- The main verification periods are Summer 2022 and Winter 2022/2023, with winter upper-air results and tropical-cyclone verification spanning 2018.
- In the Sodankylä case, -29°C was observed; PGW indicated event severity earlier than IFS, but both forecasts substantially overestimated the temperature.
- Figure 3 compares ensemble, PGW, and IFS 2m-temperature forecasts with the Sodankylä SYNOP observation and the verifying European analysis.
4 Comparing forecast performance
PGW forecasts match or approach IFS performance across several verification settings, including RMSE, observations, temperature extremes, and tropical-cyclone position. However, PGW exhibits small-scale smoothing, forecast-lead-time bias drift, and weaker tropical-cyclone intensity forecasts.
- Global forecast performance: For lead times beyond 3 days, PGW is better than ERA5 and as good as operational IFS in RMSE.The ensemble mean minimizes RMSE but is not necessarily a plausible atmospheric scenario.
- Forecast attributes: PGW and IFS have similar forecast activity for T850 and Z500, although PGW shows clear small-scale smoothing.The activity metric is dominated by larger scales, so it can understate small-scale dampening.
- Forecast attributes: PGW bias grows faster than IFS or ERA5 bias, with particularly strong Z500 drift that remains present at longer forecast horizons.IFS and ERA5 biases stabilize at longer lead times, unlike PGW bias.
- Surface temperature verification: Against SYNOP observations, PGW retains good RMSE performance and summer bias drift, while the ensemble mean outperforms other forecasts from day 4 onward in both seasons.Observation-based verification gives similar results to analysis-based verification for these key metrics.
- Statistical consistency: PGW captures summer temperature extremes with an offset, but neither PGW nor IFS fully captures extremely low winter temperatures.Rank histograms also show systematic PGW bias in summer and undercoverage of the lowest winter temperatures.
- Extreme-event forecasting: PGW outperforms IFS for summer temperature extremes at statistically significant percentiles from 75% to 90%, while winter performance is similar.After accounting for PGW’s less-consistent summer climatology, its summer extreme forecasts are more accurate than IFS.
- Tropical-cyclone forecasting: Tropical-cyclone position errors differ only slightly and nonsignificantly, but PGW clearly underestimates cyclone intensity and has larger central-pressure errors than IFS.PGW’s positive central-pressure bias is associated with too-weak gradients and maximum wind speeds.
5 Summary and outlook
The assessment finds that PanguWeather forecasts are skillful across several variables and some extreme events, while identifying forecast smoothness, bias drift, training-data limitations, and deterministic-only evaluation as areas for improvement.
- PanguWeather is skillful for upper-air variables and 2m temperature verified against analyses and observations.The study also reports consistent findings for 10m wind speed, although those results are not shown.
- The model is trained on ERA5 rather than operational IFS analysis, and ERA5 resolution may limit small-scale structure forecasts.The experiments do not include fine-tuning on operational IFS analysis.
- Operational IFS analysis initialization improves PanguWeather accuracy into the medium range.Similar error growth between data-driven and standard NWP forecasts indicates comparable sensitivities to chaos.
- PanguWeather forecasts are smoother than IFS forecasts and show an almost linear bias drift with forecast lead time.Statistical post-processing is identified as one possible way to correct systematic errors.
- PanguWeather shows good performance for some extreme events, supported by case studies.For tropical cyclones, preliminary investigations indicate similar track quality but weaker intensity and structure capture than IFS.
- The study evaluates deterministic forecasts, while ensemble forecasts remain important for decision-making uncertainty.An initial-condition perturbation approach has produced promising ensemble results, with model uncertainty left for future work.
- ML forecasts could complement physical models because they run much faster and at much lower computational cost.The paper recommends that operational centres explore their strengths and weaknesses as additional forecasting-system components.
Contextualising the forecast skill
The paper contextualizes forecast skill using summary statistics and accuracy scores that quantify forecast–observation relationships and errors.
- Forecast verification measures the relationship between a forecast and its corresponding observation.The study uses metrics and diagnostics beyond generic scores to investigate forecast and observation distributions.
- Bias is the average forecast–observation difference, while forecast and observation activity are anomaly standard deviations.
- RMSE measures forecast error, while anomaly correlation measures forecast accuracy.
- The observation is used broadly as an assumed truth and may be either an analysis or an observation.
- The forecast, observation, climatology, and latitude-weighted averaging operator define the notation used in the verification metrics.
Checking for statistical consistency
Statistical consistency is evaluated regionally with Q-Q plots and locally with observation rank histograms, while high-altitude stations are excluded to reduce representativeness effects.
- Q-Q plots compare empirical forecast and observation quantiles across the verification sample.Quantiles are estimated separately for forecasts and observations over Europe in summer or winter.
- Stations above 1000m altitude are excluded to avoid emphasizing representativeness issues rather than model characteristics.
- Observation rank histograms assess whether forecasts cover the observed empirical range at each station.The method ranks observations over the verification period and records each forecast’s position within the station-specific distribution.
- A flat observation rank histogram indicates similar forecast and observed distributions.
Forecasting weather events
The event-forecasting analysis tests discrimination for locally defined temperature extremes, using bias-adjusted climatologies and block bootstrapping for significance.
- Forecasts and observations are converted to binary event or non-event values and evaluated with contingency tables.Each contingency table is a 2 × 2 table of event outcomes.
- Separate observation and forecast climatologies remove forecast bias when assessing discrimination.This eigen-climatology approach distinguishes discrimination from calibration attributes that could be improved by post-processing.
- Statistical significance is estimated at the 5% level using a 1000-member block bootstrap with 5-day blocks.
Forecasting tropical cyclones
The study evaluates tropical-cyclone forecasts from PGW, IFS, and ERA5 against IBTrACS using the ECMWF operational tracker. Verification covers forecast lead times up to 5 days.
- Forecast verification: Tropical cyclones are tracked in PGW, IFS, and ERA5 forecasts using the ECMWF operational TC tracker.The observations come from the IBTrACS database.
- Forecast verification: Forecasts are verified through 5 days because longer lead times are unlikely to be statistically significant.
- Forecast verification: If only IFS were validated, the maximum number of cases would be 988 initially and 592 at the 5-day forecast.