Source-linked AI summary
Skillful Precipitation Nowcasting using Deep Generative Models of Radar
Suman Ravuri, Karel Lenc, Matthew Willson, Dmitry Kangin, Remi Lam, Piotr Mirowski, Megan Fitzsimons, Maria Athanassiadou, Sheleem Kashem, Sam Madge, Rachel Prudden, Amol Mandhane, Aidan Clark, Andrew Brock, Karen Simonyan, Raia Hadsell, Niall Robinson, Ellen Clancy, Alberto Arribas, Shakir Mohamed
TL;DR
Skillful precipitation nowcasting remains important for weather-dependent decision-making, while existing approaches struggle with nonlinear events and can produce blurry longer-lead predictions. The paper develops a conditional deep generative model with spatial and temporal consistency constraints; expert forecasters preferred it most often, and quantitative verification found competitive skill without blurring.
Problem
Existing nowcasting approaches struggle with nonlinear precipitation events, while unconstrained deep learning methods can produce blurry longer-lead predictions that limit operational utility.
Method
The paper uses a conditional generative model that predicts future radar fields from past radar observations, trained with spatial and temporal discriminator losses to discourage blurry and temporally inconsistent predictions.
Results
88% of forecaster judgments ranked the generative approach first for 5 mm/hr nowcasts, while quantitative evaluation found competitive verification and preserved precipitation statistics without blurring.
Takeaways & Limitations
Deep generative nowcasting can provide skillful probabilistic forecasts and physical insight that support operational decision-making where alternative methods struggle.
Takeaways & Limitations
The generative ensemble is under-dispersed compared to PySTEPS, and predicting high precipitation at long lead times remains difficult for all approaches.
Abstract
from arXiv · showhide
Precipitation nowcasting, the high-resolution forecasting of precipitation up to two hours ahead, supports the real-world socio-economic needs of many sectors reliant on weather-dependent decision-making. State-of-the-art operational nowcasting methods typically advect precipitation fields with radar-based wind estimates, and struggle to capture important non-linear events such as convective initiations. Recently introduced deep learning methods use radar to directly predict future rain rates, free of physical constraints. While they accurately predict low-intensity rainfall, their operational utility is limited because their lack of constraints produces blurry nowcasts at longer lead times, yielding poor performance on more rare medium-to-heavy rain events. To address these challenges, we present a Deep Generative Model for the probabilistic nowcasting of precipitation from radar. Our model produces realistic and spatio-temporally consistent predictions over regions up to 1536 km x 1280 km and with lead times from 5-90 min ahead. In a systematic evaluation by more than fifty expert forecasters from the Met Office, our generative model ranked first for its accuracy and usefulness in 88% of cases against two competitive methods, demonstrating its decision-making value and ability to provide physical insight to real-world experts. When verified quantitatively, these nowcasts are skillful without resorting to blurring. We show that generative nowcasting can provide probabilistic predictions that improve forecast value and support operational utility, and at resolutions and lead times where alternative methods struggle.
Generative Models of Radar
The model generates probabilistic future radar fields from past observations using latent variables and spatially and temporally focused training objectives. Four consecutive radar observations condition samples of 18 future frames, while a regularization term improves location accuracy.
- The conditional generative model predicts N future radar fields from M past radar fields using latent variables Z and parameters θ.
- Four consecutive radar observations covering 20 minutes condition multiple realizations of 18 future precipitation frames spanning 90 minutes.
- Spatial and temporal discriminators encourage spatially consistent, non-blurry fields and temporally consistent, non-jumpy sequences.
- A grid-cell regularization term penalizes deviations between real radar sequences and the model predictive mean, improving location-accurate predictions.
- Training uses 256 × 256 radar crops from 110-minute precipitation events, with importance sampling to represent high precipitation more strongly.
Intercomparison Case Study
A challenging convective event exposes distinct failure modes among the baselines: PySTEPS overestimates intensity, while UNet and Axial Attention blur away structure. The generative method better preserves spatial coverage and convection, and 93% of forecasters ranked it first.
- Intercomparison Case Study: The case study compares the generative method with PySTEPS, UNet, and Axial Attention on three meteorologically challenging events selected by an expert forecaster.
- Intercomparison Case Study: The event features convective cells and intense showers making landfall over eastern Scotland, making cell maintenance difficult to forecast.
- Intercomparison Case Study: PySTEPS overestimates rainfall intensity over time and insufficiently covers the rainfall’s spatial extent, whereas UNet and Axial Attention blur away intensity and small-scale structure.
- Intercomparison Case Study: 93% of forecasters chose the generative method as their first choice for the challenging event.
- Intercomparison Case Study: CSI shows the deep learning methods outperform PySTEPS, but similar scores distinguish less than expert judgments between qualitatively different forecasts.
Quantitative Evaluation
Quantitative evaluation finds the generative method competitive or superior across location accuracy, spectral realism, and probabilistic performance. Unlike blurred deep-learning baselines, it preserves precipitation structure across spatial and temporal scales.
- Quantitative Evaluation: All three deep learning systems are significantly more location-accurate than PySTEPS by CSI, with the generative method significantly better than PySTEPS at every precipitation threshold.
- Quantitative Evaluation: The generative method and PySTEPS match radar spectral characteristics, while UNet and Axial Attention lose medium- and small-scale variability as lead time increases.
- Quantitative Evaluation: At T+90 min, UNet has 32 km × 32 km effective resolution and Axial Attention has 16 km × 16 km precipitation variability.
- Quantitative Evaluation: As spatial aggregation increases, the generative method and PySTEPS retain similarly strong CRPS performance, while Axial Attention underperforms all other methods at scale four and above.
- Quantitative Evaluation: Additional analyses report generalization across seasons, rain types, regions, alternative metrics, and data splits.
- Quantitative Evaluation: Overall, the generative method preserves precipitation statistics across spatial and temporal scales without the blurring used by other deep learning methods.
Expert Forecaster Assessment
The study directly assessed operational usefulness through judgments from 56 Met Office forecasters, focusing on accuracy and value for medium- and heavy-rain events. The generative method was strongly preferred, while interviews identified both its strengths and remaining failure cases.
- Evaluation design: 56 Met Office forecasters evaluated nowcasts for decision-making value and operational usefulness.The assessment used a two-phase experimental protocol in the UK national meteorology service’s 24/7 operational center.
- Evaluation design: Medium rain exceeded 5 mm/hr and heavy rain exceeded 10 mm/hr, representing the top 10% and top 1% of rain events, respectively.Forecasters ranked preferences using samples drawn from 2,126 high-intensity precipitation events.
- Results: 88% of 5 mm/hr rankings and 89% of 10 mm/hr rankings preferred the generative approach first, with p<10^-4 for both comparisons.The p-values were computed using permutation tests with 10,000 resamplings, alongside Clopper-Pearson 95% confidence intervals.
- Qualitative assessment: Forecasters described the generative method as having the best envelope, representing risk best, and capturing convection-cell size and intensity most effectively.Alternative methods were described as having positional errors, excessive intensity, or overly bland and unrealistic fields.
- Limitations: The generative method showed intensity decay for heavy rainfall at T+90 min and difficulty predicting isolated showers in cases where alternatives were preferred.Forecasters identified these as important future improvements.
Conclusion
The authors conclude that deep generative models provide fast, accurate short-term precipitation predictions where existing methods struggle. They also argue that operational utility requires evaluation approaches connecting quantitative verification with expert judgment.
- Conclusion: Deep Generative Models provide fast and accurate short-term predictions at lead times where existing methods struggle.The conclusion presents them as a valuable forecasting tool for probabilistic nowcasting.
- Evaluation implications: The work calls for quantitative measurements better aligned with operational utility because standard verification metrics and expert judgments are not mutually indicative of value.This concern is especially relevant for models with few inductive biases and high capacity.
- Conclusion: The approach directly tackles skillful nowcasting for weather-dependent decision-making and improves upon existing solutions.The authors frame operational insight for real-world decision-makers as a central outcome.
Methods
The Methods section provides additional details about the data, models, and evaluation, with references to extended data supporting the main-text results.
- Methods: Additional methodological details about the data are provided in the Methods section.The passage groups these details with model and evaluation information.
- Methods: Additional methodological details about the models are provided in the Methods section.The passage indicates that these details supplement the main text.
- Methods: Additional methodological details about the evaluation are provided in the Methods section and extended data.The extended data add to the results presented in the main text.
Datasets
The experiments use UK radar composites with defined spatial coverage, temporal splits, and preprocessing. The training data are rebalanced because most grid cells contain no rain and medium-to-high precipitation is rare.
- Dataset scope: The main-text experiments use a United Kingdom radar dataset, with additional quantitative results available from a United States dataset.The UK dataset is used for all main-text experiments.
- Radar data: UK radar composites represent surface precipitation rate over 1 km × 1 km grid cells across a 1536 × 1280 composite.Missing precipitation values are masked during training and evaluation, and fields are quantized in increments of 1/32 mm/hr.
- Data splits: Radar observations were collected every five minutes from 1 January 2016 through 31 December 2019, with 2019 reserved for testing.Validation uses the first day of each month from 2016–2018, while other 2016–2018 days form the training set.
- Class balance: Approximately 89% of UK grid cells contain no rain, while medium-to-high precipitation above 4 mm/hr comprises fewer than 0.4% of grid cells.The dataset is rebalanced to include more higher-precipitation observations.
- Training set preparation: Each example contains 24 radar observations over two continuous hours, and preprocessing yields roughly 1.5 million training examples.The data use 256 × 256 crops, extracted every 32 grid cells, with importance sampling and removal of entirely masked examples.
Model Details and Baselines
The model combines a convolutional conditioning stack and recurrent sampler with latent inputs, while spatial and temporal discriminators and regularization guide realistic, non-blurry predictions. The paper compares this generative system with several deep-learning and operational nowcasting baselines.
- Generative architecture: The generator uses four radar observations to create conditioning representations that initialize a four-unit ConvGRU sampler producing 18 future radar frames.The sampler receives one latent representation for each lead time and progressively upsamples recurrent outputs.
- Generative architecture: The conditioning stack is a feed-forward convolutional network that processes four radar observations through downsampling residual blocks and multiscale representations.Its outputs have spatial-channel sizes 64 × 64 × 48, 32 × 32 × 96, 16 × 16 × 192, and 8 × 8 × 384.
- Training objectives: A spatial discriminator promotes spatial consistency and discourages blur, while a temporal discriminator distinguishes observed from generated radar sequences.The temporal discriminator processes sequences containing contextual frames and predictions or targets; the spatial discriminator samples eight of 18 lead times.
- Training objectives: Training combines the two discriminator losses with a grid-cell regularizer that keeps mean predictions close to ground truth and emphasizes higher rainfall.The regularizer is averaged across height, width, and lead-time axes and uses a precipitation-weighting function.
- Generative architecture: The model uses latent random inputs whose spatial scale is 1/32 of the radar field, enabling spatiotemporally consistent predictions over larger evaluation regions.At evaluation, full radar fields are 1536 × 1280 and the latent representation is 48 × 40 × 8.
- Baselines: The evaluation compares the generative method with PySTEPS, UNet, and a radar-only MetNet implementation, while UNet is strengthened with residual blocks and intensity-weighted loss.The UNet modifications target performance at longer lead times and higher precipitation.
Evaluation
The evaluation combines quantitative verification with expert-forecaster comparisons across challenging precipitation events. The generative method is assessed against established baselines, including PySTEPS, UNet, Axial Attention, and NWP forecasts.
- Evaluation: The study evaluates models quantitatively and through a cognitive assessment task with expert forecasters.Models are trained on 2016–2018 and evaluated on 2019 unless otherwise noted.
- Expert Forecaster Study: 71% of forecasters ranked the generative approach first for a challenging high-precipitation front.The comparison involved 56 forecasters; the preference was statistically significant with p<10^-4.
- Expert Forecaster Study: 76% of forecasters ranked the generative approach first for a challenging cyclonic circulation event.The generative method captured the precipitation extent overall, although it slightly overdid coverage between bands.
- Quantitative Evaluation: In nearly all rain-type comparisons, the generative method outperformed competing methods on CRPS and CSI.Its performance was particularly strong for non-frontal rain, which is difficult to predict.
- Baselines: The evaluated NWP forecast was not competitive with the other methods in this nowcasting regime.The comparison used the UKV deterministic forecast.
- Generalization: The generative method showed competitive verification performance under weekly splits and on the United States MRMS dataset.The United States evaluation used two years for training and one year for testing.
A.2. United States Dataset
The United States dataset uses radar composites from the Multi-Radar Multi-Sensor system to train and evaluate precipitation nowcasting models.
- A.2. United States Dataset: The MRMS dataset combines composites from 146 WSR-88D radars and 30 Canadian radars covering the conterminous United States.The composites span 20°–55°N and 130°–60°W.
- A.2. United States Dataset: MRMS composites have spatial dimensions of 3584 × 7168 and 0.01° resolution in both latitude and longitude.This corresponds to approximately 1.11 km north–south, with east–west spacing varying across the image.
- A.2. United States Dataset: The data consists of radar composites collected every 2 minutes from January 1, 2017 through December 31, 2019.The source covers three calendar years for dataset construction.
A.3. Dataset Statistics
The dataset statistics describe rainfall distributions in the UK and US yearly test sets, including missing-rain prevalence, intensity differences, and the effect of importance sampling.
- A.3. Dataset Statistics: Rainfall-distribution statistics are reported for UK and US yearly test sets across locations and dataset types.The statistics are computed across 15 consecutive frames.
- A.3. Dataset Statistics: The datasets contain a high proportion of grid cells with no rain.The statistics also compare high-intensity grid-cell frequencies between the United States and United Kingdom.
- A.3. Dataset Statistics: Importance sampling shifts the sampled data toward higher-intensity grid cells to support learning.The statistics show the effect of this sampling scheme on the rainfall distributions.
B.1. Additional Quantitative Evaluation on Yearly Data Splits
Additional yearly-split analyses examine training variability, loss-component choices, rain-type performance, and whether subsampling changes quantitative results.
- B.1. Additional Quantitative Evaluation on Yearly Data Splits: Six generative-method instances were trained with different training-example sequences to quantify training variability.CSI and pooled CRPS were reported with one-standard-deviation error bars.
- Loss Ablations: The proposed combination of discriminator losses and grid-cell regularization is important for favorable performance across metrics.Using only grid-cell regularization improves per-grid-cell CSI but increases CRPS and worsens PSD characteristics.
- Additional Metrics: Supplementary analyses compare CSI, CRPS, FSS, Pearson correlation, calibration, and rain-type performance across model initializations.The supplementary figures report confidence intervals and multiple precipitation thresholds or spatial scales.
- Subsampling Check: The full-frame and subsampled UK datasets show no quantitative difference in CSI scores under the reported five-sample evaluation.Full-frame prediction is computationally prohibitive, so the comparison uses five samples rather than the twenty used elsewhere.
B.2. NWP Results
This section details baseline implementations, adaptations, and dataset comparisons used to assess precipitation nowcasting at operational lead times.
- NWP reference and baselines: The study compares radar-based methods with an NWP reference using UK rainflux data at 5-minute temporal and 1.5 km spatial resolution.The rainflux variable is upsampled to the OSGB36 1 km reference grid.
- Dataset comparison and computation: The 512 × 512 sub-sampled dataset showed no quantitative difference from the 1536 × 1280 full-frame dataset in the reported comparison.The full-frame comparison used five samples for STEPS and the generative method because full-frame prediction was computationally prohibitive.
- NWP reference and baselines: The axial-attention model uses spatial downsampling, ConvLSTM temporal encoding, and axial-attention aggregation to produce prediction frames.Feature maps are reduced to 64 × 64 spatial dimensions before temporal and attention processing.
- NWP reference and baselines: MetNet is adapted into an axial-attention radar model by reducing its context to four frames covering 20 minutes.The original implementation used seven frames over 90 minutes with 15-minute intervals.
- NWP reference and baselines: Additional elevation, positional, and temporal embeddings did not produce statistically significant CSI changes in the axial-attention model.The authors hypothesize that four input frames already model the relevant precipitation dynamics at nowcasting timescales.
- NWP reference and baselines: The MetNet-style implementation predicts lead times from 5 to 90 minutes in the UK setting with 5-minute temporal resolution.The original MetNet configuration used 15- to 480-minute lead times at 15-minute resolution.
E. Verification Metrics
The verification framework evaluates point, ensemble, neighborhood, calibration, and spatial-structure properties of nowcasts across grid cells and local neighborhoods.
- Evaluation setup: The evaluation conditions on M context frames and assesses predictions over the subsequent N frames using central target grid cells.For UK data, M and N are 4 and 18, respectively.
- Point and threshold metrics: CSI measures binary rainfall-exceedance performance by combining weighted true positives, false positives, and false negatives.Thresholds can represent low, medium, or heavy rain, including 2, 5, and 10 mm/hr.
- Probabilistic metrics: CRPS scores per-grid-cell predictive distributions against observations, with lower values indicating better performance.The evaluation estimates CRPS from ensemble members treated as samples from the predictive distribution.
- Calibration diagnostics: Reliability diagrams compare forecast probabilities with observed frequencies, while sharpness diagrams show how often each probability is forecast.Perfect calibration corresponds to alignment with the diagonal, subject to finite-sample error.
- Calibration diagnostics: Rank histograms assess continuous-rainfall calibration by plotting the pooled rank of each observation among ensemble forecasts.A uniform histogram is expected for perfectly calibrated forecasts.
- Evaluation caveats: The rank-histogram implementation excludes masked cells and cannot incorporate the importance weights used by other metrics.This biases emphasis toward deviations from uniformity associated with higher-rainfall examples.
- Neighborhood and spatial metrics: Neighborhood metrics pool observations over K × K areas, giving partial credit when the broad spatial pattern is correct despite small-scale displacement.FSS evaluates forecast and observed rainfall fractions over these neighborhoods, while pooled CRPS also probes dependence between nearby grid cells.