Source-linked AI summary

WeatherBench 2: A benchmark for the next generation of data-driven global weather models

Stephan Rasp, Stephan Hoyer, Alexander Merose, Ian Langmore, Peter Battaglia, Tyler Russel, Alvaro Sanchez-Gonzalez, Vivian Yang, Rob Carver, Shreya Agrawal, Matthew Chantry, Zied Ben Bouallegue, Peter Dueben, Carla Bromberg, Jared Sisk, Luke Barrington, Aaron Bell, Fei Sha

arXiv:2308.15560v2physics.ao-phcs.AI

TL;DR

Data-driven global weather forecasting has advanced rapidly, creating a need for an updated benchmark that supports fair comparison with operational systems. WeatherBench 2 provides an open-source framework, shared datasets, operational-style metrics, and headline scores for physical and data-driven models. The reported results show competitive deterministic models at shorter lead times, ensemble-mean advantages later, and similar upper-level probabilistic skill for IFS ENS and NeuralGCM ENS.

  • Problem

    Rapid advances in data-driven weather modeling created a need for an updated benchmark that enables easy comparison between different approaches.

  • Method

    WeatherBench 2 combines operational-style forecast verification, open-source data and code, headline scores, and evaluations of physical and data-driven models.

  • Results

    Deterministic data-driven models have lower errors than IFS ENS (mean) up to 3–6 days, after which the ensemble mean has the lowest error; IFS ENS and NeuralGCM ENS have very similar upper-level probabilistic scores.

  • Takeaways & Limitations

    WB2 provides a reproducible framework for evaluating new data-driven weather methods against operational baselines while making headline metrics and model comparisons publicly accessible.

  • Takeaways & Limitations

    No single metric or metric set fully describes forecast quality, and ERA5 does not represent all variables equally well, especially precipitation.

Abstract

from arXiv · show

WeatherBench 2 is an update to the global, medium-range (1-14 day) weather forecasting benchmark proposed by Rasp et al. (2020), designed with the aim to accelerate progress in data-driven weather modeling. WeatherBench 2 consists of an open-source evaluation framework, publicly available training, ground truth and baseline data as well as a continuously updated website with the latest metrics and state-of-the-art models: https://sites.research.google/weatherbench. This paper describes the design principles of the evaluation framework and presents results for current state-of-the-art physical and data-driven weather models. The metrics are based on established practices for evaluating weather forecasts at leading operational weather centers. We define a set of headline scores to provide an overview of model performance. In addition, we also discuss caveats in the current evaluation setup and challenges for the future of data-driven weather forecasting.

1 Introduction

WeatherBench 2 updates the original benchmark in response to rapid advances in data-driven global weather forecasting. It supports higher-resolution evaluation, additional metrics, and comparison with operational and AI-based models.

  • Motivation: Global medium-range NWP supports forecasting significant weather events and supplies boundary conditions, reanalyses, and research tools.The relevant forecast horizon is 1–14 days.
  • Motivation: Traditional NWP still relies on parameterizations for unresolved processes and data assimilation to estimate the atmosphere’s initial state.Cloud physics and radiation are examples of unresolved processes.
  • Prior benchmark: WeatherBench provided a common benchmark for emerging data-driven medium-range weather models.It followed initial machine-learning attempts at global weather forecasting.
  • Recent advances: Recent models span graph neural networks, vision transformers, spherical Fourier neural operators, recurrent-diffusion systems, and hybrid ML-physics approaches.These models operate across resolutions and report competitive or state-of-the-art forecast skill.
  • WeatherBench 2: WeatherBench 2 responds to rapid progress by adding higher-resolution data and evaluation, additional metrics, headline scores, and discussion of benchmark limitations.The paper presents design principles, datasets, evaluation metrics, and results for current models.

2 Design decisions for WeatherBench 2

WeatherBench 2 follows operational forecast verification while keeping modeling choices open and emphasizing probabilistic prediction. Its open-source framework is intended to evolve with the machine-learning weather community.

  • Evaluation principles: No single metric can fully describe forecast quality because weather is high-dimensional, multi-faceted, and use-case dependent.WB2 therefore provides headline scores rather than claiming a complete definition of a good forecast.
  • Evaluation principles: WB2 makes standard WMO and operational-center verification accessible through open-source ground truth data, evaluation code, and centralized model comparisons.The protocol follows established verification practices rather than reinventing them.
  • Benchmark scope: WB2 evaluates entire forecast systems, leaving inputs, training setup, and model resolution open instead of prescribing a model architecture.This includes the interaction of input data, training, and architecture.
  • Probabilistic prediction: Probabilistic prediction is emphasized because chaotic error growth creates a range of possible future outcomes even with highly accurate models and initial conditions.Ensembles support decision-making by representing different potential weather realizations.
  • Open framework: All evaluation data and code are publicly available and designed for extension by the authors or community contributors.The framework can accommodate more detailed evaluation as community needs change.

3 Data, baselines and data-driven models

WB2 combines ERA5 data, operational IFS forecasts, climatological baselines, and data-driven model baselines in a shared evaluation framework. The included models differ in architecture, resolution, inputs, and initialization.

  • Datasets: ERA5 is WB2’s ground-truth and common training dataset, providing hourly data from 1940 to the present through reanalysis and data assimilation.ERA5 is based on a 0.25° version of ECMWF’s HRES model and 4D-Var assimilation.
  • Datasets: ERA5 is a model simulation whose fidelity varies by variable, with especially large discrepancies possible for precipitation relative to rain-gauge measurements.The paper includes precipitation evaluation but advises caution in interpreting it.
  • Data and model inventory: Table 1 summarizes datasets and models using horizontal resolution, vertical levels, inference hardware, and related metadata.GPU and TPU denote graphics and tensor processing units.
  • Baselines: WB2 includes operational IFS forecasts, ERA5 forecasts, climatological forecasts, and probabilistic climatology as baselines.ERA5 forecasts provide a like-for-like baseline for models initialized and evaluated against ERA5.
  • Baselines: The operational IFS baseline is regularly updated and uses HRES forecasts, while its ensemble includes a control run and 50 perturbed members.The exact operational configuration can change during the evaluation period.
  • Data-driven models: The benchmark describes data-driven models including Keisler’s GNN, Pangu-Weather, and GraphCast, with differences in grids, vertical levels, time steps, and forecast construction.Pangu-Weather is evaluated in versions trained for Δt = {1h, 3h, 6h, 24h}; WB2 also includes an IFS-initialized version.

4 Evaluation protocol and metrics

WB2 uses a standardized operational-style evaluation protocol over 2020, with common regridding and explicit handling of initialization, resolution, and below-ground grid points. The protocol also documents temporal and training-data caveats.

  • Protocol: WB2 follows WMO and operational-center forecast verification practices.The notation for the evaluation metrics is listed in Table 4.
  • Evaluation period: The initial WB2 evaluation uses 2020 because it balances recency, future independent testing, and baseline data availability.One year is generally adequate for most metrics, but extreme-event metrics require larger samples.
  • Data caveats: Many AI models were trained only through 2018, but relative score differences between 2018 and 2020 are small for the displayed metrics.The evaluation period can be updated to a more recent or longer period.
  • Evaluation period: Evaluation uses all 00 and 12 UTC initializations in 2020, with 6 hours as the highest-resolution evaluation time step.Some forecasts consequently extend into 2021, without materially changing scores when restricted to forecasts valid in 2020.
  • Spatial evaluation: Forecasts and ground truth are conservatively regridded to 1.5° before metric computation.This common resolution avoids penalizing coarser models, although higher-resolution models can resolve additional small-scale details.
  • Spatial evaluation: WB2 computes results over all global grid points and excludes below-ground pressure-level points from metrics.Regional scores are also provided on the accompanying website.

4.3 Deterministic metrics

WeatherBench 2 evaluates deterministic forecasts with area-weighted metrics, anomaly-based measures, bias diagnostics, and a precipitation-specific categorical score. The framework adapts established operational evaluation practices to account for spatial resolution, variable distributions, and forecast characteristics.

  • Spatial weighting: Area weighting prevents equiangular latitude–longitude grids from overemphasizing smaller polar cells.Latitude weights are computed from each grid cell’s upper and lower latitude bounds, while pressure levels are treated separately.
  • Error metrics: RMSE is computed separately for each variable and level, including a wind-vector version based on forecast and observed u and v components.The formulation agrees with WMO and ECMWF; moving the time mean outside the square root, as in WB1, changed results by less than 2% on average.
  • Anomaly correlation: ACC measures Pearson correlation between forecast and observed anomalies relative to climatology, with values from 1 for perfect correlation to -1 for perfect anti-correlation.A climatological forecast has ACC 0, and ECMWF considers values below 0.6 to lack forecasting value for synoptic-scale feature positioning.
  • Bias metrics: Bias is computed for each location, and globally averaged root mean squared bias provides an additional diagnostic.
  • Precipitation evaluation: SEEPS evaluates precipitation through dry, light, and heavy categories because RMSE and ACC favor unrealistically smooth forecasts for skewed, intermittent precipitation.WB2 uses a 0.25 mm/day dry threshold, excludes very wet and very dry regions, and reports 1 - SEEPS as a positively oriented skill score.

4.4 Probabilistic metrics

WeatherBench 2 evaluates probabilistic forecasts with CRPS and the spread-skill ratio, combining distributional accuracy, ensemble spread, and calibration diagnostics. CRPS supports multidimensional predictions and reduces to weighted mean absolute error for deterministic forecasts.

  • CRPS: CRPS combines a prediction-to-observation error term with a prediction-spread term, penalizing poor forecasts while encouraging ensemble spread.The score is minimized when predictions follow the ground-truth distribution under the relevant conditioning.
  • CRPS: WeatherBench computes multidimensional CRPS by averaging over time, latitude, and longitude components using M ≥2 predictions.For deterministic predictions with M = 1, CRPS reduces to weighted mean absolute error; the ensemble term can be computed in O(M log M) time.
  • Spread-skill ratio: The spread-skill ratio is the ensemble spread divided by the RMSE of the ensemble mean.The ensemble mean is averaged across M predictions, and ensemble variance supplies the spread.
  • Spread-skill ratio: A well-calibrated ensemble should have a spread-skill ratio of 1; smaller values indicate underdispersion and larger values indicate overdispersion.The ratio is only a first-order calibration test, with rank histograms identified as a further diagnostic.

4.5 Energy Spectra

WeatherBench 2 characterizes spatial structure through zonal energy spectra computed along constant-latitude circles. The spectra are normalized to preserve Parseval’s energy relation and averaged over midlatitude bands.

  • Spectrum definition: Zonal spectral energy is computed as a function of wavenumber, frequency, and wavelength along lines of constant latitude.
  • Spectrum construction: The discrete Fourier transform of values sampled around each zonal circle is converted into an energy spectrum using S0 := C|F0|2 and Sk = 2C|Fk|2.The factor of 2 accounts for both negative and positive frequency content at nonzero wavenumbers.
  • Spectrum construction: The normalization is chosen so that Parseval’s relation holds for the sampled continuous function.
  • Aggregation: The final spectrum averages zonal spectra over latitudes satisfying 30° < |lat| < 60°.

4.6 Headline scores

WeatherBench 2’s headline scores cover commonly evaluated upper-level and surface variables, linking atmospheric dynamics, moisture transport, and weather impacts. The listed abbreviations include wind vector and wind speed.

  • Upper-level variables: Headline upper-level variables are selected to capture large-scale atmospheric evolution, with Z500 and T850 tracing extra-tropical dynamics.
  • Upper-level variables: Q700 serves as a proxy for moisture transport and, indirectly, clouds.
  • Surface variables: Surface headline variables include 2 m temperature, 10 m wind speed, and 24 h precipitation accumulation because they align with weather impacts.
  • Score list: Table 3 lists the headline scores and defines WV as Wind Vector and WS as Wind Speed.

5 Results

WeatherBench 2 evaluates contemporary physical and data-driven weather systems using headline deterministic and probabilistic scores, alongside diagnostics of temporal artifacts, biases, spectra, and case-study forecasts. Results show competitive data-driven performance, but interpretation depends on metric choice, forecast representation, and evaluation limitations.

  • Evaluation scope: Results emphasize evaluation metrics and baselines, with less emphasis on directly ranking data-driven models because the snapshot will soon become outdated.The authors show detailed scores for only a smaller selection of models to keep visualizations readable.
  • Headline scores: Data-driven deterministic models achieve errors similar to IFS HRES and lower errors than IFS ENS (mean) through roughly 3–6 days, after which the ensemble mean is lowest.NeuralGCM ENS (mean) roughly matches IFS ENS (mean) scores on RMSE.
  • Headline scores: A strong 6h zig-zag pattern in data-driven scores arises from training and evaluating on ERA5’s 12h assimilation window.The first six forecast hours primarily emulate the IFS version used in ERA5, while the next six hours also require learning the assimilation step.
  • Headline scores: For precipitation, SEEPS ranks IFS HRES as most skillful while blurrier models such as IFS ENS (mean) perform worse; RMSE reverses this ordering.This contrast illustrates how categorical and average scores can favor different forecast characteristics.
  • Probabilistic scores: IFS ENS and NeuralGCM ENS have similar probabilistic upper-level scores, approaching probabilistic climatology near the two-week horizon while IFS ensemble forecasts are generally well calibrated.The precipitation forecast is underdispersive, and some variables retain a gap to climatological skill.
  • Diagnostics and case studies: Model diagnostics reveal shared and model-specific forecast busts, increasing wind-speed bias in GraphCast, scale-dependent smoothing, and strong agreement across models for Storm Alex through four days.GraphCast’s wind-speed bias reflects difficulty representing correlations between separately predicted wind components; Pangu-Weather and NeuralGCM do not show this bias.

6 Discussion

The discussion identifies important limits of the evaluation setup, including ERA5 ground truth, initialization conditions, forecast realism, and incomplete probabilistic metrics. It also outlines extensions involving direct observations, extremes, post-processing, and broader benchmark use.

  • ERA5 as ground truth: ERA5 provides broad temporal and spatial coverage but is a model simulation, and precipitation evaluation against it is only a placeholder for more accurate data.Surface variables, especially precipitation, can differ substantially from observations because precipitation is not directly assimilated into ERA5.
  • ERA5 as ground truth: Direct observations, including station data, are a planned extension, although uneven coverage, varying quality, and quality-control requirements create additional challenges.Operational weather services already use direct observations alongside assimilated ground truths.
  • ERA5 initialization versus operational forecasting: Operational comparisons require distinguishing ERA5-initialized models from models initialized with operational analyses.The reported scores for ERA5-initialized and operational versions of GraphCast and Pangu-Weather are very similar, but the paper still encourages operational initialization for apples-to-apples comparisons.
  • Forecast realism, probabilistic forecasts and extremes: Deterministic mean-error training produces increasingly smooth forecasts that can score well while failing to represent realistic weather states and local feature intensity.The paper connects this behavior to excessive spectral blurring, precipitation-category errors, wind-speed biases, and underrepresented cyclone intensity.
  • Forecast realism, probabilistic forecasts and extremes: Probabilistic metrics such as CRPS and spread/skill remain limited because they are univariate, ignore spatial and temporal correlations, and do not specifically target extreme weather.The paper cautions that better CRPS values do not automatically imply a more useful forecast.
  • Post-processing: WeatherBench 2 can evaluate post-processing models as well as data-driven forecast models, while future comparisons should include post-processed dynamical forecasts.Previous studies cited in the paper suggest probabilistic post-processing can improve CRPS by up to 20%, depending on the variable.

7 Conclusion

WeatherBench 2.0 updates the benchmark for data-driven global weather forecasting, aligning evaluation with operational practice and providing open code and data. It is designed to evolve with new metrics and models while expanding toward observations and extreme-weather evaluation.

  • 7 Conclusion: WeatherBench 2.0 provides a robust framework for evaluating new data-driven methods against operational baselines.Its design stays close to forecast evaluation used by weather centers and supplies evaluation code and data to support reproducible results.
  • 7 Conclusion: The framework is intended to accelerate machine-learning workflows and improve reproducibility through publicly available evaluation code and data.
  • 7 Conclusion: WeatherBench 2.0 will be updated with new metrics and models as data-driven weather research evolves.Discussed extensions include station observations and evaluation of extremes.

Supplement

The supplement provides figures covering global scores, biases, resolution effects, and a Hurricane Laura case study for weather-model evaluation.

  • Supplement: Figure S1 reports global RMSE and SEEPS values for 2018.
  • Supplement: Figure S2 compares IFS HRES and ENS mean evaluations against analysis and ERA5 ground truth.
  • Supplement: Figures S3–S7 show global mean biases for temperature, precipitation, wind speed, and specific humidity at 3-, 5-, and 10-day lead times.Figures S3, S4, S6, and S7 include the stated below-ground masking where applicable.
  • Supplement: Figure S8 compares RMSE scores for IFS HRES evaluated at different resolutions.
Loading 2308.15560v2…