Source-linked AI summary

Diffusion-Based Refinement for Kilometer-Scale Probabilistic Precipitation Nowcasting

Dohyun Park, Changhoon Song, Tengyuan Chang, Yoo-Geun Ham, Youngjoon Hong

arXiv:2608.30205v1cs.LGphysics.ao-ph

TL;DR

Localized extreme precipitation requires forecasts that combine fine spatial detail with probabilistic uncertainty, but conventional and existing generative approaches can be computationally demanding. exPreCast-ENS extends a deterministic 4 km radar nowcaster with conditional residual diffusion to produce 1 km precipitation ensembles. Across KMA and MeteoNet, the framework improves with ensemble size and provides structured mean corrections alongside plausible fine-scale variability, while calibration and broader applicability remain open limitations.

  • Problem

    Localized extreme precipitation demands fine-resolution, frequently updated forecasts with uncertainty, while conventional ensemble prediction is computationally costly.

  • Method

    exPreCast-ENS conditions a lightweight residual diffusion module on an exPreCast forecast and preceding radar observations to generate 1 km probabilistic ensembles.

  • Results

    Forecast skill improves consistently with ensemble size, while the ensemble mean provides a structured nonzero-mean correction and members represent plausible fine-scale precipitation evolution.

  • Takeaways & Limitations

    Residual diffusion is an effective and computationally practical extension of deterministic precipitation nowcasting across KMA and MeteoNet settings.

  • Takeaways & Limitations

    Some under-dispersion remains, and broader evaluation with additional backbones, radar networks, climates, and longer forecast horizons is needed.

Abstract

from arXiv · show

Localized extreme precipitation is a major trigger of urban flash floods and landslides, yet producing nowcasts that combine fine spatial detail with probabilistic uncertainty remains challenging. Here we introduce exPreCast-ENS, a conditional residual diffusion framework that transforms the deterministic 4 km radar nowcaster exPreCast into a 1 km probabilistic ensemble while correcting systematic forecast errors. Conditioning on both the forecast and preceding radar observations lets the ensemble-mean correct the baseline rather than perturb it, while members represent unresolved fine-scale variability. Over the Korean Peninsula, skill improves with ensemble size. In two high-impact events in 2023, a 30-member ensemble recovers 38-47% of heavy-rain pixels missed by exPreCast while retaining approximately 95% of its correct detections and alarming on under 1% of the pixels it correctly left clear. The method generates a 1-h forecast in 3.4 s on a single GPU and yields consistent improvements on the French regional MeteoNet radar dataset.

1 Introduction

High-resolution probabilistic precipitation nowcasting remains difficult because localized extremes demand fine spatial detail, frequent updates, and uncertainty estimates. exPreCast-ENS addresses this gap by extending a deterministic 4 km nowcaster with lightweight residual diffusion for 1 km ensembles.

  • NWP faces substantial computational demands when producing kilometer-scale precipitation forecasts with ensemble-based uncertainty estimates.
  • Generative radar nowcasting represents multiple plausible future scenarios and quantifies uncertainty, but existing diffusion frameworks can require iterative generation or latent encoding and decoding.
  • exPreCast-ENS conditions an EDM-based residual diffusion module on a deterministic exPreCast forecast and preceding radar observations to generate 1 km probabilistic forecasts from 4 km inputs.
  • The framework recovers localized extreme precipitation structures missed by the deterministic backbone while producing forecasts in seconds on a single GPU.
  • The learned residual distribution has a systematically nonzero ensemble mean that corrects deterministic-backbone biases rather than merely adding random perturbations.
  • Experiments validate the framework on both KMA and MeteoNet, including a setting where residual refinement does not increase spatial resolution.

2 Results

On KMA, exPreCast-ENS is evaluated against the deterministic backbone at matched 4 km resolution while also assessing 1 km distributional fidelity and complementary ensemble products. Forecast skill improves with ensemble size, and individual members better reproduce observed fine-scale rainfall statistics.

  • The two-stage system refines a coarse deterministic forecast into a finer probabilistic forecast using exPreCast as the backbone and EDM training for diffusion.
  • Skill and distribution on KMA: KMA evaluation compares ensemble sizes N = 1, 10, 20, 30 with the exPreCast baseline after pooling 1 km outputs to the native 4 km grid.
  • Skill and distribution on KMA: Performance improves consistently with ensemble size across RMSE, CRPS, CSI, FSS, and BSS, while reliability diagrams move toward the diagonal.
  • Skill and distribution on KMA: BSS is positive for N ≥10 at every lead time, indicating more informative exceedance probabilities as members are added.
  • Skill and distribution on KMA: Individual 1 km SR forecasts more closely reproduce observed rainfall-intensity distributions and spatial power spectra than bilinearly interpolated exPreCast forecasts.
  • Visual characteristics of the forecast products.: A single SR member preserves fine-scale textures, the 30-member mean provides a stable expected field, and the alarm mask highlights threshold-supported heavy-rain signals.

2.2 Non-zero mean correction and spatial redistribution

The diffusion residual is not merely zero-mean noise: its 30-member mean remains a structured correction to exPreCast. The correction largely redistributes rainfall spatially, with local increases and decreases canceling in the domain total.

  • The 30-member mean remains distinct from exPreCast after ensemble statistics stabilize, demonstrating a predictable nonzero-mean residual component.
  • The ensemble-mean correction is spatially coherent and modifies precipitation intensity and placement rather than simply adding uniform rainfall.
  • Approximately 4% net versus 40% absolute correction over the full 2023 test period indicates substantial local changes that largely cancel in the domain-integrated amount.
  • Table 1 separates signed net changes, magnitude-based absolute changes, and their ratio for the 30-member mean relative to the backbone.
  • The figure compares 1 km ensemble means with native 4 km backbone forecasts after mean-pooling, using a diverging difference map for added and removed rainfall.

2.3 Event-level verification through case studies

Case studies compare radar observations, the deterministic backbone, individual SR members, the 30-member mean, and probability-thresholded alarms. These products expose different aspects of fine-scale structure, expected rainfall, and heavy-rain support.

  • The case studies examine two KMA events at T +60 min using ground truth, exPreCast, a single SR member, the 30-member mean, and an alarm product.
  • Individual members restore fine-scale variability, while the ensemble mean spatially stabilizes the forecast and the alarm product emphasizes threshold-exceedance support.

Case 1: Extreme rainfall in Seoul.

In the Seoul case, the 4 km backbone places the rainband too far south, while 1 km ensemble members include northward-shifted realizations and the 30-member products recover central-city rainfall.

  • The backbone places the observed rainband too far south, missing most heavy-rain regions in Seoul.
  • The figure compares ground truth, the 4 km backbone, a single 1 km member, the 30-member mean, and the alarm mask at T+60 min.
  • A single 1 km member places the rainband over central Seoul, illustrating positional variation across ensemble members.
  • The 30-member mean mitigates positional uncertainty, while the alarm mask recovers the rain across the city more sharply.

Case 2: Heavy rainfall associated with the flooding in Cheongju.

In the Cheongju flooding case, the backbone captures the coarse rainband but misses finer extreme-rain structures; residual diffusion resolves localized hazards while conditional evaluation separates recovery from false alarms.

  • The backbone captures the rainband’s location and intensity but misses fine structures because of its coarse grid.
  • A single refined prediction resolves subgrid intensity variation and reveals localized extreme rainfall relevant to the Osong underpass flood.
  • The ensemble mean preserves the overall rainband while reducing spatial uncertainty, producing a smoother and more consistent representation.
  • Conditional recovery: The conditional recovery rate measures heavy-rain pixels missed by exPreCast that the SR alarm subsequently detects.
  • Conditional recovery: The SR alarm recovers missed heavy rain, introduces relatively few new false alarms near the rainband, retains most correct detections, and can remove false alarms.

Case 3: Extreme rainfall associated with Typhoon Khanun.

For Typhoon Khanun, the backbone misses a substantial portion of observed heavy rain, whereas the probability-thresholded SR alarm recovers many missed pixels within or near the precipitation system.

  • The deterministic backbone detects only part of the observed heavy-rain area, leaving a substantial portion undetected.
  • The probability-thresholded SR alarm recovers many heavy-rain pixels missed by exPreCast.
  • Recovered pixels concentrate within or along the margins of the observed precipitation system rather than across unrelated regions.
  • The case studies show spatially coherent recovery, retention of most correct backbone detections, and relatively few new false alarms.

2.4 Application to another radar dataset: MeteoNet

On MeteoNet, residual diffusion is independently trained at the native 1 km resolution and improves probabilistic forecasting while applying structured corrections even without spatial super-resolution.

  • MeteoNet uses 1 km reflectivity observations and backbone forecasts, so diffusion operates at the native resolution rather than refining a coarser grid.
  • Because the backbone already matches the target resolution, individual diffusion forecasts provide only modest additional agreement in distribution and power-spectrum metrics.
  • The MeteoNet experiment is an independent replication, with components trained separately and no weights transferred from KMA.
  • The ensemble mean retains coherent positive and negative residual structures relative to the deterministic backbone after ensemble-size results stabilize.
  • Residual diffusion provides a general mechanism for probabilistic forecasting at either finer or unchanged spatial resolution.

3 Discussion

Residual diffusion extends deterministic precipitation nowcasting into a probabilistic framework whose mean corrects structured forecast errors while members represent fine-scale variability. The approach is computationally compact, but calibration, threshold selection, and broader evaluation remain open considerations.

  • The residual diffusion module gives ensemble members and their mean complementary roles: members represent plausible fine-scale variability, while the mean provides a structured correction.The correction redistributes precipitation spatially rather than merely changing total rainfall.
  • The 30-member mean remains distinct from exPreCast after ensemble statistics stabilize, showing that diffusion does more than add stochastic detail around the deterministic forecast.Preceding radar observations help infer lead-time-dependent displacement and intensity errors in the frozen backbone.
  • The KMA diffusion module contains 6.00 M parameters versus 32.0 M for exPreCast, and generates one 1 h member in 3.4 seconds on a single NVIDIA A6000 GPU.A 30-member ensemble takes approximately 100 seconds sequentially, with parallel sampling possible across GPUs or through batched inference.
  • Operational alarm performance depends on the exceedance-probability criterion, whose skill-maximizing value varies with rainfall intensity and forecasting region.The main experiments fix the alarm threshold at 0.3 for all thresholds and both datasets, whereas regional validation or cost–loss analysis could select different values.
  • The probability field remains imperfectly calibrated, with some under-dispersion unresolved and its source not distinguished among the residual formulation, frozen backbone, and sampler.Calibration-aware training or post-processing is identified as a direction for future work, especially for rare, high-impact rainfall thresholds.
  • Current evidence covers two radar datasets and one deterministic backbone, so additional nowcasters, radar networks, climates, and longer horizons are needed to establish broader applicability.The framework also has not yet been assessed for efficient cross-dataset adaptation instead of separate training.

4 Methods

exPreCast-ENS combines a frozen deterministic radar nowcasting backbone with conditional residual diffusion to generate probabilistic corrections on the target grid. The framework conditions denoising on both forecast and past observations, then forms ensemble members by adding sampled residuals to the interpolated forecast.

  • Framework overview: The backbone predicts future radar fields, while residual diffusion generates probabilistic corrections conditioned on the deterministic forecast and preceding radar observations.This decomposition preserves backbone structure while representing multiple plausible forecast-error realizations.
  • Grid formulation: On KMA, 4 km backbone forecasts are bilinearly interpolated to the 1 km target grid; on MeteoNet, interpolation is the identity.The residual and diffusion model operate on the target grid in both datasets.
  • Conditional diffusion: Past observations and the interpolated deterministic forecast are concatenated as conditioning channels alongside noisy residual channels.Removing the past sequence degrades skill scores, indicating that the conditioning retains temporal information not fully represented by the forecast.
  • Sampling: Each sampled residual is denormalized and added to the interpolated deterministic forecast to form one ensemble member.Thirty realizations are drawn per test case, with reported ensemble sizes of 1, 10, 20, and 30.
  • Network and implementation: The denoising network is a full-domain 1 km U-Net, with K noisy residual channels and K + L conditioning channels.On KMA, bottleneck attention is bypassed because the 1024 × 1024 feature map exceeds the single-GPU memory budget; the configuration has 6.00 M parameters.
  • Alarm mask: The alarm mask flags grid points where the ensemble exceedance probability for threshold τ reaches the fixed criterion palarm = 0.3.The ensemble exceedance probability is computed separately at each grid point and lead time.
  • Data and evaluation: KMA evaluation uses 10-minute radar data with 7-frame inputs, 6-frame forecasts, and a 1 km target domain covering 1024 × 1024 pixels.The KMA split uses 2014–2021 for training, 2022 for validation, and 2023 for testing.
  • Inference: A single KMA stochastic member requires 3.4 s using 20 Heun sampling steps on one A6000 GPU, while a sequential 30-member ensemble requires approximately 100 s.Independent members can be parallelized across members and GPUs.

Data availability

MeteoNet data are publicly available from Météo-France, while the KMA radar composites are distributed under Korea Meteorological Administration terms.

  • The MeteoNet radar dataset is publicly available from Météo-France, and the KMA composites are available through the cited release subject to KMA distribution terms.The passage provides a Google Drive link for the KMA data.

Code availability

The supporting code is available from the corresponding author upon reasonable request, with public repository deposition planned after publication.

  • Supporting code is available from the corresponding author upon reasonable request, and official exPreCast weights are available on GitHub.The code is planned for deposit in a public repository upon publication.

Supplementary Information

Supplementary analyses examine alarm-threshold sensitivity, ensemble calibration and variability, residual corrections, conditional verification, and performance across KMA and MeteoNet evaluation settings.

  • Alarm criterion: palarm = 0.3 is adopted throughout, while the CSI-maximizing alarm probability shifts lower at higher intensity thresholds and differs between KMA and MeteoNet.The criterion depends on both intensity threshold and radar-network or climatological regime.
  • Residual correction: The ensemble-mean residual decreases as ensemble size grows but levels off at a non-zero value, leaving a spatially structured correction.This persistence is shown for both KMA and MeteoNet at T +60 min.
  • Residual correction: 40–50% absolute corrections coexist with net corrections of only a few percent, indicating substantial local rainfall redistribution through offsetting increases and decreases.The same pattern holds over both the full domain and the observed rain area.
  • Additional analyses: Supplementary evaluations include full-period KMA conditional rates, MeteoNet threshold-based verification, seed variability, rank histograms, and past-context ablation.These analyses cover calibration, individual-member behavior, and the role of preceding radar observations in conditioning.
Loading 2608.30205v1…