Source-linked AI summary

A Station-Based Evaluation of Machine Learning-based Weather Forecasting Models in Northern Norway

Siyan Chen, Lars Uebbing, Eirik Mikal Samuelsen, Georgios Leontidis, Arnt-Børre Salberg, Sébastien Lefèvre, Robert Jenssen, Kristoffer Wickstrøm

arXiv:2609.10564v1physics.ao-phcs.LG

TL;DR

Local wind forecasting performance of MLWP models remains uncertain in complex terrain because most evidence comes from global gridded benchmarks. This study evaluates FCN3 and GraphCast against HRES using station observations from Northern Norway across lead times, training periods, locations, and wind conditions. HRES has the lowest overall error, MLWP remains comparable beyond training periods, and strong winds remain substantially underestimated.

  • Problem

    MLWP performance for local station-level wind forecasting in complex terrain remains insufficiently understood, especially beyond training periods and during high-wind events.

  • Method

    The study interpolates gridded FCN3, GraphCast, and HRES forecasts to station coordinates and matches them with observations at corresponding valid times.

  • Results

    HRES records the lowest overall RMSE at 2.89 m s^-1, compared with 2.96 m s^-1 for FCN3 and 2.98 m s^-1 for GraphCast; all models underestimate high winds.

  • Takeaways & Limitations

    MLWP models remain comparable to HRES for station-level wind forecasting in Northern Norway, while complex geography and high-wind prediction require further improvement.

  • Takeaways & Limitations

    The evaluation covers only 12 Northern Norway stations, and bilinear interpolation cannot fully represent local elevation, coastline, surface roughness, and unresolved terrain effects.

Abstract

from arXiv · show

Recent machine learning weather prediction (MLWP) models have demonstrated remarkable forecasting skill on global reanalysis-based benchmarks. However, their performance remains unclear in challenging environments such as Northern Norway, where narrow fjords and rapidly changing weather result in highly variable local wind conditions. In this case study, we evaluate FourCastNet3 (FCN3), GraphCast, and ECMWF High Resolution Forecast (HRES) for wind speed forecasting using multi-year station observations from Northern Norway, focusing on their relative performance, generalization beyond the training period, and performance under high-wind conditions. Our results show that HRES slightly outperforms FCN3 and GraphCast, with an overall RMSE of 2.89 $\mathrm{m\,s^{-1}}$, compared to 2.96 $\mathrm{m\,s^{-1}}$ for FCN3 and 2.94 $\mathrm{m\,s^{-1}}$ for GraphCast. Notably, the MLWP models maintain comparable performance beyond their respective training periods, with no clear evidence of noticeable degradation. FCN3 performs best under high-wind conditions, although all models substantially underestimate strong winds. Our findings suggest that MLWP has become competitive with NWP for local wind, but further refinements are still needed to capture complex terrain better.

1 Introduction

Global MLWP models perform strongly on gridded benchmarks, but their accuracy for local wind forecasting in complex terrain remains uncertain. This study evaluates FCN3 and GraphCast against HRES using station observations in Northern Norway, emphasizing lead time, location, training-period generalization, and high-wind events.

  • Global benchmark skill does not establish how MLWP models perform at station level in complex terrain, where fine-scale topography drives local wind variation.
  • The evaluation asks how accurately MLWP and physics-based models predict 10 m wind, how performance changes across training and non-training years, and how it varies with lead time, station, and wind intensity.
  • The study evaluates FCN3 and GraphCast alongside HRES because the MLWP models use different architectures and HRES provides a physics-based reference.
  • The case study provides a station-level comparison across 12 Northern Norway weather stations and characterizes performance across lead times, locations, and wind-force categories.
  • HRES achieves the lowest overall RMSE, while FCN3 has the smallest bias and performs best under high-wind conditions; all models substantially underestimate strong winds.

2 Related Work

Weather forecasting research spans physics-based NWP and increasingly capable MLWP models. Although global evaluations show competitive MLWP skill, station-based studies are needed to assess transfer to local observations and complex environments.

  • NWP combines physical equations, numerical solvers, observations, and data assimilation, with IFS and GFS representing major operational global systems.
  • Deep-learning MLWP models emerged from large reanalysis datasets and use varied architectures, including neural operators, Earth-specific transformers, graph networks, and cascaded systems.
  • Global evaluations against ERA5 and WeatherBench 2 show competitive MLWP skill but also report limitations in temporal generalization and extreme-event prediction.
  • Station-based studies compare model outputs with in-situ observations, including large global evaluations and a Norwegian Pangu-Weather assessment at 183 synoptic stations.

3 Station-Based Evaluation of Weather Forecasting Models

The evaluation framework compares global weather forecasts with station observations across Northern Norway, lead times, locations, years, and high-wind conditions. Forecasts are spatially and temporally matched at 12 stations before error and categorical-skill metrics are calculated.

  • Evaluation Framework: Forecasts are interpolated to each station and matched to observations at exact valid times, retaining common station–time samples across FCN3, GraphCast, and HRES.This produces comparable station-level forecast time series and errors against in situ measurements.
  • Evaluation Framework: Forecasts are initialized at 00:00 UTC and evaluated every 6 h from +6 to +72 h, yielding 12 lead times per forecast cycle.Each forecast is compared with observations at the corresponding valid time.
  • High-Wind Evaluation: High-wind cases are defined by observed station wind speeds of at least 10.8 m s−1, with metrics recalculated for this subset and scatter plots used to assess systematic errors.The threshold follows the high-wind definition adopted by the Norwegian weather service Yr.
  • Experimental Setup: The study evaluates FCN3, GraphCast, and HRES using in situ observations from 12 Northern Norway stations during 2016–2022.Stations were selected for geographical coverage, sufficient observations, and exposure that limits nearby-obstacle effects on 10-m wind measurements.
  • Evaluation Metrics: Performance is quantified with RMSE, bias, normalized RMSE, normalized bias, Pearson correlation, and Equitable Threat Score across overall, temporal, spatial, and wind-speed evaluations.RMSE measures overall error magnitude, bias measures systematic error, and normalized metrics support comparisons across stations with different wind-speed magnitudes.

4 Results

Across Northern Norway, HRES has the lowest overall RMSE, while FCN3 has the smallest bias and performs best during high-wind events. Errors vary substantially by lead time, station, terrain, and wind-force category, but MLWP performance shows no consistent temporal degradation.

  • Overall Forecasting Performance: 2.89 m s−1 overall RMSE makes HRES the best aggregate model, ahead of GraphCast at 2.94 m s−1 and FCN3 at 2.96 m s−1.Only the FCN3–HRES difference was statistically significant; the other pairwise differences included zero.
  • Overall Forecasting Performance: FCN3 has the smallest overall bias at -0.85 m s−1, whereas HRES and GraphCast show larger negative biases of -1.21 m s−1 and -1.26 m s−1.All models systematically underestimate observed wind speeds.
  • High-Wind Analysis: ETS is strongest for lower Beaufort categories, with HRES leading forces 1–4 and FCN3 leading forces 5–8, while skill is negligible above force 10.The strongest-wind categories are also much less frequent in the observations.
  • Lead Time Analysis: RMSE rises from approximately 2.7–2.8 m s−1 at +6 h to 3.1–3.2 m s−1 at +72 h, with HRES lowest at every lead time.FCN3 and GraphCast are similar at shorter leads, while FCN3 increases notably after +48 h.
  • Station-based and Terrain-based Analysis: Station effects exceed most model differences: Bardufoss has RMSE below 1.8 m s−1, while Fakken reaches 4.58 m s−1 for HRES.Model rankings vary by station, although their station-wise error patterns are broadly similar.
  • Training and Non-Training Years: Annual performance shows no consistent increase in error over time, suggesting no systematic degradation across training and non-training years.Similar year-to-year variations across MLWP models and HRES are attributed to differences in annual forecast difficulty rather than systematic drift.
  • High-Wind Analysis: High-wind RMSE is around twice the overall value, with normalized biases from -43.42% to -45.59%; FCN3 performs best but all models substantially underestimate strong winds.FCN3 has the lowest high-wind RMSE and nRMSE and the smallest absolute bias, while HRES has the largest errors.

5 Discussion

At station level, HRES has only a modest numerical advantage over FCN3 and GraphCast, while spatial heterogeneity and unresolved terrain strongly shape forecast errors.

  • 5 Discussion: FCN3 and GraphCast remain comparable to HRES over Northern Norway despite MLWP models performing better against ERA5 in earlier comparisons.The discussion cautions that the choice of reference dataset affects the apparent ranking.
  • 5 Discussion: Forecast performance varies spatially because coarse global grids do not resolve local geography, including fjord terrain, coastal ridges, wind channelling, and downslope wind storms.These effects can produce low RMSE in sheltered valleys without capturing local dynamics, or larger errors at exposed coastal stations.
  • 5 Discussion: HRES slightly outperforms FCN3 and GraphCast at station level, but their RMSE differences are less than 0.07 m s−1 and smaller than variation among stations.The aggregate comparison supports a modest numerical advantage rather than consistent HRES superiority.
  • 5 Discussion: The evaluation covers only 12 stations in Northern Norway, so its conclusions may not extend directly to regions with different climates, terrain, or observation networks.The authors recommend testing additional complex coastal and mountainous regions.
  • 5 Discussion: Bilinear interpolation introduces station-to-grid representativeness error because it cannot fully capture elevation, coastline geometry, surface roughness, or unresolved terrain.Future improvements may require higher-resolution regional models or downscaling that explicitly incorporates local surface characteristics.

6 Conclusion

The study compares FCN3 and GraphCast with ECMWF HRES for station-level wind forecasting in Northern Norway. HRES has the lowest overall errors, while high winds expose substantial weaknesses across all models.

  • 6 Conclusion: 2.89 m s−1 is HRES’s overall error, compared with 2.96 m s−1 for FourCastNet3 and 2.98 m s−1 for GraphCast.HRES achieved the lowest errors, while the MLWP models delivered comparable results.
  • 6 Conclusion: Errors increase with forecast lead time, while spatial differences between stations are more pronounced.
  • 6 Conclusion: All models become less accurate under high-wind conditions and significantly underestimate wind speeds.High-wind degradation appears as increased RMSE and more negative bias.
  • 6 Conclusion: Improving complex-geography representation and high-wind prediction is identified as a priority for future MLWP development.

A Computational Resources

The appendix documents computational requirements and the geographic and meteorological context of the station evaluation. The resources and station metadata support interpretation of station-level differences.

  • A Computational Resources: Each FCN3 and GraphCast annual task requested one GPU, four CPU cores, 192 GB of memory, and up to 40 hours of execution time.Seven annual tasks were submitted in each SLURM job array, with one task per year from 2016 to 2022.
  • A Computational Resources: The Beaufort force scale categorizes wind intensity by wind-speed ranges used in the study.
  • A Computational Resources: The evaluation stations span coastal, fjord, and inland valley environments across Northern Norway, with elevations from near sea level to approximately 76 m.Coordinates, elevation, and terrain category provide context for interpreting station-level forecast differences.

D Evaluation against ERA5

The models were additionally evaluated against ERA5 using 366,948 samples. All three had much lower RMSE values in that comparison, with GraphCast leading RMSE and nRMSE and FCN3 showing the smallest systematic bias.

  • D Evaluation against ERA5: 1.14–1.31 m s−1 is the RMSE range for the three models against ERA5 across 366,948 samples.
  • D Evaluation against ERA5: GraphCast achieves the lowest RMSE and nRMSE against ERA5, while FCN3 has the smallest systematic bias with an nBias of 0.81%.GraphCast and HRES have negative nBias values of −3.30% and −4.95%, respectively.

E Comparison of CARRA1 and ERA5 with Station Observations

CARRA1 represents near-surface winds at the selected stations more accurately than ERA5, consistent with its finer resolution and regional design for complex Arctic terrain.

  • The comparison evaluates how CARRA1 and ERA5 represent local near-surface wind conditions against station observations.
  • CARRA1 is a regional Arctic reanalysis based on HARMONIE-AROME with approximately 2.5 km grid spacing, compared with ERA5’s approximately 31 km resolution.Its higher spatial resolution is designed to better represent regional geography and local atmospheric conditions.
  • 2.41 m s−1 RMSE for CARRA1 versus 2.71 m s−1 for ERA5 shows closer agreement with station observations.CARRA1 also has lower nRMSE: 44.67% versus 50.19%.
  • CARRA1 reduces wind-speed bias to −0.28 m s−1, compared with −0.89 m s−1 for ERA5.Its nBias is −3.78%, versus −16.47% for ERA5.

F Station-level Heatmap

The station heatmap reveals substantial spatial variation in normalized error and bias across the three forecasting models and 12 stations.

  • 33.52%–77.03% nRMSE across stations and models indicates pronounced variation in normalized forecasting error.The highest nRMSE values occur at Bardufoss, Sørkjosen Lufthavn, and Tromsø-Langnes.
  • The heatmap reports RMSE, nRMSE, Bias, and nBias for FCN3, GraphCast, and HRES across 12 stations.
  • Bias ranges from −3.50 to 1.02 m s−1 and is negative for most station-model combinations.The most negative biases occur at Fakken, Banak, and Tromsø-Langnes.
  • Bardufoss has the lowest absolute RMSE but among the highest nRMSE values across the station-model results.The high normalized error partly reflects generally low observed wind speeds.

G Station-Level Scatter Plots

The station-level scatter plots compare observed wind speeds with FCN3, GraphCast, and HRES predictions, providing station-specific views of model performance.

  • The scatter plots compare observed wind speeds with predictions from FCN3, GraphCast, and HRES across the remaining stations.These visualizations provide station-specific views of model performance.
  • Figures G.1 and G.2 show station-specific scatter plots for Skrova Fyr and Bardufoss.
  • Figure G.3 presents scatter plots of observed and predicted wind speeds at stations.
  • Figure G.4 reports station-wise RMSE, nRMSE, Bias, and nBias, while Figure G.5 zooms into the local surroundings of the 12 stations.
Loading 2609.10564v1…