Source-linked AI summary
Residual Corrective Diffusion Modeling for Km-scale Atmospheric Downscaling
Morteza Mardani, Noah Brenowitz, Yair Cohen, Jaideep Pathak, Chieh-Yu Chen, Cheng-Chin Liu, Arash Vahdat, Mohammad Amin Nabian, Tao Ge, Akshay Subramaniam, Karthik Kashinath, Jan Kautz, Mike Pritchard
TL;DR
Expensive km-scale numerical simulations motivate machine-learning downscaling from coarse global weather inputs. The paper introduces CorrDiff, which predicts a conditional mean and then generates a residual with diffusion, achieving realistic multivariate regional weather structure and substantial computational efficiency while retaining calibration challenges.
Problem
Km-scale hazard prediction requires expensive regional numerical simulations driven by coarser global inputs, motivating alternatives for probabilistic high-resolution downscaling.
Method
CorrDiff uses a two-step architecture in which UNet regression predicts the conditional mean and a diffusion model generates residuals for multivariate 2-km downscaling from 25-km inputs.
Results
CorrDiff produces reasonably realistic spectra, distributions, and coherent frontal and typhoon structures, while inference is 652 times faster and 1,310 times more energy efficient than CWA-WRF on CPUs.
Takeaways & Limitations
CorrDiff supports a potential global-to-regional multiscale machine-learning simulation approach for probabilistic high-resolution weather downscaling.
Takeaways & Limitations
CorrDiff predictions are generally under-dispersive and not yet optimally calibrated, while typhoon cases can produce overly contracted cyclone morphology.
Abstract
from arXiv · showhide
The state of the art for physical hazard prediction from weather and climate requires expensive km-scale numerical simulations driven by coarser resolution global inputs. Here, a generative diffusion architecture is explored for downscaling such global inputs to km-scale, as a cost-effective machine learning alternative. The model is trained to predict 2km data from a regional weather model over Taiwan, conditioned on a 25km global reanalysis. To address the large resolution ratio, different physics involved at different scales and prediction of channels beyond those in the input data, we employ a two-step approach where a UNet predicts the mean and a corrector diffusion (CorrDiff) model predicts the residual. CorrDiff exhibits encouraging skill in bulk MAE and CRPS scores. The predicted spectra and distributions from CorrDiff faithfully recover important power law relationships in the target data. Case studies of coherent weather phenomena show that CorrDiff can help sharpen wind and temperature gradients that co-locate with intense rainfall in cold front, and can help intensify typhoons and synthesize rain band structures. Calibration of model uncertainty remains challenging. The prospect of unifying methods like CorrDiff with coarser resolution global weather models implies a potential for global-to-regional multi-scale machine learning simulation.
1 Introduction
Kilometer-scale weather downscaling is valuable but computationally expensive, motivating probabilistic machine-learning alternatives. CorrDiff uses a two-step approach and shows realistic coherent-weather improvements, sample efficiency, and large computational savings in a Taiwan proof of concept.
- Motivation: Kilometer-scale forecasts are needed for risk assessment and local topographic and land-use effects, but global km-scale simulation is computationally challenging.Training costs grow superlinearly with resolution, while regional dynamical downscaling is expensive and limits ensemble sizes.
- Prior approaches: Statistical and deterministic machine-learning downscaling can be inexpensive, but probabilistic outputs require additional interventions or distributional assumptions.Diffusion models offer a generative alternative, while GANs face mode collapse, instability, and long-tail modeling challenges.
- Approach: CorrDiff targets simultaneous multivariable stochastic downscaling and channel synthesis from coarse global predictions to regional high-resolution weather.The paper demonstrates this proof of concept for the region surrounding Taiwan.
- Contributions: CorrDiff adds physically realistic improvements to under-resolved frontal systems and typhoons while learning effectively from just 3 years of data.These are reported as contributions alongside the two-step physics-inspired design.
- Efficiency: 22 times faster and 1,300 times more energy efficient on a single GPU than the numerical model used to produce its high-resolution training data.The numerical model runs on 928 CPU cores.
2 Generative downscaling: Corrector diffusion model
CorrDiff maps coarse 25-km meteorological inputs to higher-resolution regional targets by predicting a conditional mean and stochastically generating a residual. Residual modeling reduces the distributional variance the diffusion model must learn, addressing challenges observed with direct conditional diffusion.
- Problem setup: CorrDiff takes 25-km global meteorological inputs and higher-resolution aligned regional targets, with target dimensions larger than the input grid.The Taiwan proof of concept uses ERA5 inputs and radar-assimilating WRF targets.
- Diffusion modeling: Diffusion models learn the conditional target distribution through forward noising and backward sequential denoising.The denoising network guides samples toward representations of the target data distribution.
- Two-step method: CorrDiff first predicts the conditional mean with UNet regression, then uses a diffusion model to generate a correction residual.This decomposition is motivated by poor convergence and incoherent structures when directly learning the full conditional distribution.
- Residual formulation: Assuming accurate mean regression, the residual is approximately zero mean and retains the conditional variance of the target while reducing the modeled target variance.The variance reduction is especially pronounced when the conditional mean varies strongly, such as during typhoons.
- Experimental setting: The study trains on 2018–2020 data and tests on 2021, with additional 2022 frontal and 2023 typhoon case studies.The proof-of-concept dataset uses hourly WRF targets and corresponding ERA5 inputs.
3 Results
Across 205 out-of-sample times, CorrDiff improves probabilistic skill and restores missing multiscale variability, while case studies show sharper coherent fronts and stronger typhoon structure. Its predictions remain imperfect, especially for radar statistics, ensemble calibration, and rare extreme events.
- 3.2 Skill: CorrDiff exhibits the most skill by CRPS, followed by UNet, random forest, and interpolated ERA5.The comparison uses 205 randomly selected 2021 validation times and 32 ensemble members for CorrDiff.
- 3.3 Spectra and distributions: CorrDiff significantly improves power-spectrum realism for 10-meter kinetic energy, 2-meter temperature, and synthesized radar reflectivity over deterministic baselines.The corrective diffusion restores variance especially for radar reflectivity across length scales, with additional gains for kinetic energy and temperature at specified scales.
- 3.3 Spectra and distributions: CorrDiff matches target radar-reflectivity distributions between 0 and 43 dBZ, whereas UNet and RF fail to produce realistic radar statistics.Temperature tails improve only incrementally, and the windspeed PDF is virtually unchanged relative to UNet.
- 3.4 Calibration: CorrDiff remains under-dispersive for most channels, with ensemble spread too small relative to mean error and rank histograms showing observations outside predicted ranges.The authors identify stochastic calibration as a priority for future development.
- 3.5.1 Frontal system case study: For a cold front, CorrDiff sharpens temperature and wind gradients while concentrating generated radar reflectivity near the frontal boundary.The generated morphology differs from the target but remains consistent across winds and temperature.
- 3.5.2 Tropical Cyclone case study: Typhoons remain a limitation because they are rare in training data and only partially resolved at the 25-km input resolution.These conditions produce cyclonic structures that are too wide and too weak before downscaling.
- 3.5.2 Tropical Cyclone case study: For typhoon Haikui, CorrDiff partially restores the high-wind tail, reduces the radius of maximum winds from 75 km in ERA5 to about 50 km, and raises maximum windspeed from 22 m s−1 to 33 m s−1.The corresponding WRF values are 25 km and 45 m s−1; CorrDiff predicts winds up to 40 m s−1 versus 50 m s−1 in the target.
- 3.5.2 Tropical Cyclone case study: Extended typhoon analysis suggests the highlighted Haikui case is qualitatively representative when typhoons are far from the domain boundaries, but limitations remain.The supplied passage identifies this as an extended analysis of additional dates, storms, and a 600-member ensemble.
4 Discussion
CorrDiff combines fast mean prediction with stochastic residual correction to downscale multivariate weather fields, while offering substantial computational savings. Its main remaining challenges are uncertainty calibration, temporal coherence, regional data assimilation, and broader geographic or climatic deployment.
- Method: CorrDiff first predicts the conditional mean with a UNet, then uses diffusion to stochastically generate residual fine-scale details.The two-step design trades fast mean inference against probabilistic residual generation.
- Results: CorrDiff produces reasonably realistic power spectra and probability distributions across target variables in extensive Taiwan testing.The diffusion component is especially important for synthesizing radar reflectivity.
- Results: CorrDiff improves the representation of under-resolved frontal systems and typhoons, including sharper co-located gradients and qualitatively realistic rainband details.Typhoon corrections remain partial, while frontal-event improvements include winds, temperature, and radar reflectivity.
- Limitations: CorrDiff’s uncertainty is not yet optimally calibrated, and useful km-scale forecasting additionally requires temporal coherence and regional data assimilation.The current demonstration inherits global ERA5 assimilation but effectively bypasses regional assimilation.
- Efficiency: 652 times faster and 1,310 times more energy efficient than CWA-WRF on CPUs, although the comparison between dynamical and statistical downscaling is limited.The comparison uses current hardware and code and focuses on generation quality rather than optimal inference speed.
- Future directions: Future extensions target coarse-resolution forecasts, other geographic regions, future climate predictions, and sub-km sensor synthesis.These directions require addressing forecast-error dependence, scarce kilometer-scale data, scalability, climate sensitivity, and sensor-generation challenges.
5 Methods
The method models diffusion as forward noise addition and backward denoising, then applies this framework to residual downscaling. CorrDiff separates conditional-mean regression from stochastic residual generation and recombines them into high-resolution samples.
- Diffusion background: Forward diffusion adds Gaussian noise to the data distribution until sufficiently large noise approaches pure Gaussian noise.The noise level is controlled by σ, with the forward process transforming pdata(x) into pdata(x; σ).
- Diffusion background: Backward diffusion starts from Gaussian noise and denoises through descending noise levels toward the original data distribution.The sequence follows σ0 = σmax > σ1 > . . . > σN = 0.
- Diffusion background: The backward SDE combines a deterministic probability-flow component with stochastic noise injection, requiring a score function for sampling.The score is ∇x log p(x; σ), and σ̇(t) denotes the derivative of the noise schedule.
- CorrDiff formulation: CorrDiff writes the target state as a mean plus residual, allowing diffusion to learn a smaller, near-zero-mean distribution.The residual formulation reduces target variance, especially when var(E[x|y]) is large.
- CorrDiff formulation: A UNet regression model estimates the conditional mean, after which diffusion is trained on r = x − ˆµ.The regression model uses MSE loss, while the residual’s smaller departure from the target permits smaller diffusion noise levels.
- CorrDiff formulation: An EDM conditional diffusion model samples p(r|y), and the sampled residual is added to ˆµ to produce ˆµ + r.The method uses a second-order EDM stochastic sampler to solve the reverse SDE.
5.3 Experimental setup
The experiment downscales 25-km ERA5 reanalysis over Taiwan to 2-km regional weather-model data using a two-stage CorrDiff architecture. Evaluation examines probabilistic calibration and forecast quality, including CRPS, alongside the data and training configuration.
- Dataset: The study uses Taiwan and surrounding ocean, whose meteorology includes typhoons, mid-latitude weather fronts, steep topography, and land-sea contrast.
- Dataset: ERA5 provides 12 conditioning channels at approximately 25-km spatial and 1-hour temporal resolution, interpolated onto the CWA grid.
- Dataset: The target is 2-km, hourly RWRF data from 2018–2021 on a 448 × 448 nested domain, incorporating radar and surface-data assimilation.
- Network architecture and training: CorrDiff uses an 80-million-parameter UNet for both regression and diffusion, with six encoder and six decoder layers, Fourier timestep embeddings, and spatial positional embeddings.
- Network architecture and training: The regression network receives the 12 ERA5 channels, while diffusion additionally uses four noise channels and the first-stage regression mean.
- Evaluation criterion: Evaluation assesses calibration through spread-error relationships and rank histograms, and uses CRPS as a proper scoring rule for probabilistic forecasts.
1 Our position with respect to existing works
The work positions CorrDiff within machine-learning weather downscaling research while emphasizing joint prediction of multiple atmospheric variable types and coherent structures. Its distinctive capability is evaluating these variables together for physically consistent regional fields.
- Existing works: Prior downscaling studies addressed state-vector inflation, large spatial domains, large resolution ratios, and precipitation in tropical cyclones.
- Our position: CorrDiff jointly predicts dynamical, thermodynamical, and microphysical variables, enabling analysis of their combined downscaling in coherent structures.
2 Descriptions of the architecture and the training data
CorrDiff uses one UNet architecture in two roles: regression and diffusion denoising. The model maps 12 coarse-resolution input channels to outputs that include channels absent from the input, using a dataset split with a final-year validation period.
- Architecture: The UNet serves as both the regression model and diffusion denoiser, while the input and output datasets differ in pixel size and channel composition.
- Channels: The inputs include total-column water vapor and pressure-level variables, whereas maximum radar reflectivity appears only among the outputs.
- Training data: Training uses 2018–2020 data, with the last 12 months of 2021 serving as validation; corrupted and missing periods were removed.
3 Localization by two-step formulation
The two-step formulation separates large-scale structure from smaller-scale residual variation. The regression stage captures larger spatial scales, leaving a more localized residual for diffusion modeling.
- Localization by two-step formulation: The regression step learns larger spatial scales, while the diffusion step models remaining smaller-scale structure.
- Localization by two-step formulation: The residual is more localized and varies less overall, especially for temperature, which is strongly driven by elevation changes.
- Localization by two-step formulation: The residual has reduced large-scale variance and removes long-range spatial autocorrelation relative to the target.
4 Examining sample diversity of CorrDiff
CorrDiff generates 20 diverse radar-reflectivity samples alongside the target and UNet predictions, enabling visual comparison of sample realism and variability.
- Figure S4 compares target data, UNet predictions, and 20 CorrDiff samples of maximum radar reflectivity across diverse cloud regimes.The visualization is provided as an animation.
5 Pooled metrics
Pooled evaluation over coarse-resolution grid boxes preserves the main performance ordering: CorrDiff remains strongest, followed by UNet, RF, and ERA5 interpolation.
- Pooled CRPS and MAE computed over 14 points reproduce the main-text performance pattern.
- CorrDiff has superior CRPS to every baseline, followed by UNet, RF, and ERA5 interpolation.The pooled results support the main-text comparison.
6 Additional Case Study Analysis
Additional analyses show that CorrDiff generally preserves frontal structure, improves some typhoon properties, and substantially reduces inference cost, while revealing limitations for difficult storm inputs and size correction.
- Additional frontal analysis: CorrDiff synthesizes frontal reflectivity consistently with other variables and keeps the warm sector cloud-free, although it does not always sharpen fronts to the target.
- Typhoon case studies: For Chanthu, CorrDiff fails to recover typhoon intensity on most days and improves over ERA5 only after the storm passes the island.
- Typhoon case studies: CorrDiff improves Haikui by correcting about 50% of the intensity error and contracting its radius of maximum winds.
- Historical typhoon analysis: Across 648 historical typhoon instances, CorrDiff improves windspeeds up to 50ms−1 and increases the probability of windspeeds exceeding 33ms−1 five-fold.It also reduces storm size, including for some storms whose input size was already correct or too small.
- Inference efficiency: CorrDiff inference takes 0.18 sec per sample and is about 500 times faster and 10,000 times more energy efficient than CWA-WRF downscaling.The comparison uses one H100 GPU versus 928 CPUs and excludes additional regional data-assimilation compute.
- Statistical significance: CorrDiff has lower CRPS than RF on all 205 evaluation times, with formal tests yielding p-values below 10−30.