Source-linked AI summary
Steering Diffusion Priors with Sparse Observations for High-Resolution Temperature Downscaling
Anirudh Avireddy, Manmeet Singh, Shivanshi Singh, Ayush Raj, Saptarishi Dhanuka, Parthasarathi Mukhopadhyay, Sandeep Juneja
TL;DR
The paper addresses fine-scale temperature downscaling when stations are sparse and ERA5 misses local terrain and land-surface contrasts. It uses a conditional diffusion emulator guided by sparse observations at inference, achieving its main advantage at 1% observation density in controlled synthetic-grid evaluation while requiring further real-station and held-out-year validation.
Problem
Sparse ground stations and coarse reanalysis limit high-resolution temperature fields needed for local heatwave hazard assessment.
Method
A conditional diffusion emulator uses geography, climatology, exact-time ERA5, and solar-temporal features, with a differentiable Gaussian likelihood steering the score toward sparse observations without retraining.
Results
At 1% observations, SDA reduces hidden-cell RMSE from 0.431 K to 0.318 K versus nearest-observation filling and wins all 32 cases, while sparser regimes favor nearest filling.
Takeaways & Limitations
The guided fields provide a temperature layer for downstream heatwave analyses such as threshold exceedance, duration, and exposure assessment.
Takeaways & Limitations
Evidence is limited to a controlled 32-case 2022 synthetic-grid protocol, so station representativeness, held-out-year transfer, humidity-aware metrics, and threshold outcomes remain to be evaluated.
Abstract
from arXiv · showhide
Local heatwave hazard depends on fine-scale air temperature, but ground stations are sparse and reanalysis products such as ERA5 cannot resolve the terrain and land-surface contrasts that shape real heat exposure. We present a conditional diffusion emulator for high-resolution 2-m temperature downscaling, conditioned on static geography, a training climatology, exact-time ERA5 temperature, and solar and temporal features, guided at inference by score-based data assimilation (SDA): a differentiable Gaussian observation likelihood steers the diffusion score toward sparse revealed temperature observations without any retraining. On a controlled 32-case synthetic-grid protocol over AORC, guidance improves hidden-cell reconstruction over both ERA5 and a strong observation-proximal nearest-neighbor baseline once observation density reaches 1\% (RMSE 0.318 vs.\ 0.431~K, winning all 32 cases), while sparser regimes still favor direct interpolation. We further map the full guidance-strength landscape across three observation densities, showing that the optimal strength shifts systematically with density and that over-guiding causes sharp, predictable degradation -- giving a concrete operating recipe rather than a single untuned setting. The resulting fields are intended as a temperature layer for downstream heatwave-hazard products such as threshold exceedance and cumulative heat-burden. The present evidence is a controlled synthetic-grid validation; station-network and held-out-year evaluations are the next steps toward deployment.
1 Introduction
The paper addresses sparse observations and coarse reanalysis by combining a conditional diffusion temperature prior with inference-time observation guidance. In controlled synthetic evaluation, guidance improves reconstruction at 1% observation density, while sparser regimes favor nearest-observation filling.
- Unlike fixed covariance data assimilation, the learned score is nonlinear, spatially adaptive, and conditioned on exact-time covariates.The generative prior also supports ensemble uncertainty signals and preserves fine-scale texture.
- The method conditions a diffusion prior on geography, climatology, exact-time ERA5, and solar-temporal features, then steers its score using sparse observations at inference.This avoids retraining for observation injection.
- At 1% revealed cells, guidance reduces hidden-cell RMSE from 0.431 K to 0.318 K against nearest-observation filling and wins all 32 cases.The evaluation uses a controlled synthetic-grid protocol over AORC.
- At lower observation densities, nearest-observation filling remains stronger than guided diffusion.The reported operating regimes therefore depend on observation density.
2 Problem formulation and data
The study models high-resolution AORC temperature anomalies from training-only climatology and multiple spatial, temporal, and meteorological conditioning fields. ERA5 supplies exact-time information rather than the climatological baseline.
- The target is hourly AORC 2-m temperature on an approximately 800 m grid, represented as an anomaly relative to a monthly-hourly per-pixel training climatology.Training uses CONUS AORC from 2020–2021, with 2022 reserved for validation and evaluation.
- Normalization statistics and climatology are computed from training years only, while patches with more than 5% invalid target pixels are excluded.
- The conditioning inputs include topography, sky-view factor, climatological statistics, exact-time ERA5 temperature, solar angle, coordinates, and day-of-year and hour-of-day encodings.ERA5 is bilinearly interpolated onto the AORC grid.
3 Conditional diffusion emulator
The emulator uses an EDM-parameterized conditional U-Net to sample plausible fine-scale temperature anomaly fields. Sampling proceeds from Gaussian noise through a sequence of noise levels using the EDM Heun update.
- The conditional U-Net uses EDM parameterization with 69,324,125 trainable parameters, multiscale channel widths, residual blocks, and attention at 32×32.Training runs for 250,000 steps with AdamW, EMA weights, mixed precision, and gradient clipping.
- Prior sampling integrates 32 rho-spaced noise levels from σmax = 80 to σmin = 0.02 with the EDM Heun update.The result is an ensemble of plausible high-resolution anomaly fields conditioned on shared covariates.
4 Score-based data assimilation
Score-based data assimilation steers the conditional diffusion sampler toward revealed temperatures through a differentiable Gaussian observation likelihood. Hidden cells are updated indirectly through learned spatial correlations and the network Jacobian.
- At each noise level, the denoiser supplies a prior score and a Gaussian observation likelihood modifies that score at revealed cells.The mechanism adapts score-based data assimilation to a conditional temperature emulator.
- The likelihood evaluates only revealed cells, while hidden cells change through the model’s learned spatial correlations and Jacobian.This allows sparse observations to reshape an entire high-resolution patch.
- Guidance strength and assumed observation error control the pull between observations and the learned prior.Main results use a fixed moderate guidance setting and assumed physical observation error of 0.5 K, with sensitivity explored separately.
5 Synthetic SDA downscaling evaluation
The evaluation uses a fixed 32-case AORC 2022 synthetic-grid protocol to compare sparse-observation reconstruction methods across three observation densities.
- Protocol: The protocol evaluates a step-205,000 EMA teacher on 32 fixed AORC 2022 patches with nested 0.01%, 0.1%, and 1% valid-cell masks.Mean observed counts are 6.94, 67.06, and 655.25 cells per 256×256 patch.
- Metrics and baselines: RMSE and MAE are computed only over valid hidden cells, with comparisons against ERA5, a prior-only sampler, nearest observed anomaly plus local climatology, Shepard interpolation, and ordinary kriging.The classical interpolators use eight revealed neighbors, while the nearest baseline preserves local anomalies through climatology reconstruction.
6 Results
At 1% observations, SDA outperforms the strong nearest-observation baseline across all 32 cases, while lower densities favor nearest filling; the representative posterior follows the reference field and exposes uncertainty.
- Results: 0.318 K versus 0.431 K RMSE at 1% observations, SDA improves on nearest observed anomaly plus local climatology in all 32 cases.The improvement is 0.1125 K in macro-mean RMSE, with a bootstrap 95% interval of 0.076–0.156 K.
- Results: The representative 1% posterior mean follows the AORC field while sparse observations steer the learned prior and ensemble spread marks spatially uncertain regions.
- Results: Nearest filling is strongest at 0.01% and 0.1% observations, whereas SDA is strongest at 1%, showing complementary operating regimes by density.The prior-only sampler indicates that the gain comes from combining the learned prior with observations.
7 From guided temperature fields to heatwave-risk analysis
The reconstructed temperature field is positioned as a fine-scale physical layer for heatwave-hazard analysis, while broader health-risk interpretation remains a downstream step requiring additional inputs and validation.
- Fine-scale hazard layer: At 1% observations, SDA beats the observation-proximal nearest baseline on all 32 patches, reducing hidden-cell RMSE by 0.112 K and recovering fine-scale structure.Such structure can expose threshold and duration contrasts smaller than a reanalysis grid cell.
- Downstream analysis: Reconstructed 2-m temperature fields can supply maps of threshold exceedance or cumulative degree-hours for combination with humidity, urban morphology, population exposure, and social vulnerability.The work establishes the reconstruction component; event-based heatwave and health-risk studies are the next step.
- Scope and limitations: Operational deployment remains bounded by the controlled 32-case 2022 synthetic-grid evaluation, uncalibrated four-member spread, and unevaluated station, held-out-year, humidity-aware, and threshold outcomes.These extensions are identified as future evaluations of the validated physical layer.
8 Conclusion
The paper concludes that conditional diffusion with denoiser-guided likelihood conditioning can downscale sparse-grid temperatures and provide a temperature layer for heatwave analyses, while broader deployment evaluations remain future work.
- Conclusion: At 1% revealed cells, SDA improves over ERA5 and the strong nearest-observation baseline on all 32 AORC 2022 patches, reducing hidden-cell RMSE by 0.112 K.The conclusion attributes this result to learned spatial structure beyond direct pointwise propagation.
- Conclusion: The method injects observations through a denoiser-differentiated likelihood without retraining and produces fields usable for heatwave threshold, duration, and exposure analyses.
- Conclusion: The evaluation diagnostics are not pooled as one benchmark because the main table, guidance stress test, and model-only diagnostic use different checkpoints and protocols.
A.1 Data, target transform, and conditioning
The emulator models standardized temperature anomalies relative to a training climatology and conditions generation on geography, exact-time weather, and temporal and solar features. EDM diffusion preconditioning and sampling produce high-resolution fields while masking invalid pixels during training and evaluation.
- Target transform: The target is a standardized anomaly relative to a per-pixel monthly–hourly climatology computed from 2020–2021 training data, with 2022 reserved for evaluation.The physical temperature field is reconstructed by adding the climatology and inverse standardization.
- Conditioning: Each patch uses twelve conditioning channels spanning topography, sky-view factor, climatology statistics, exact-time ERA5 temperature, location, and solar-temporal encodings.The noisy target is concatenated as a thirteenth model input channel.
- EDM preconditioning: The EDM denoiser combines a noise-dependent skip path, output scaling, input scaling, and noise encoding before predicting the clean target.The conditioning stack includes the static and exact-time covariates together with the noisy target.
- Validity masking: Training and evaluation use a valid-pixel indicator so missing or fill-value cells are excluded from the weighted denoising objective and reported hidden-cell errors.This masking applies to both the loss and metrics.
- Sampling and guidance: The prior sampler uses 32 noise levels and Euler proposals with Heun correction steps, while SDA progressively strengthens observation constraints as diffusion noise decreases.The production configuration uses γ = 0.001, λobs = 1, and an assumed physical observation error of 0.5 K.
C Guidance-strength sweep
The guidance sweep shows that intermediate observation forcing can improve hidden-cell reconstruction, but the best strength decreases as observation density increases and excessive forcing causes rapid degradation.
- 0.414 ± 0.123 K RMSE and 0.304 ± 0.086 K MAE are achieved at 1% observations with λ = 5, versus 0.530±0.189/0.333±0.110 for nearest-observation filling.At 5% and 10%, nearest-observation filling is stronger than the best SDA setting.
- λ = 5 at 1%, λ = 1 at 5%, and λ = 0.5 at 10%, showing that the best guidance strength shifts downward as observations become denser.
- At 1% observed cells, hidden RMSE rises from 0.414 K at λ = 5 to 1.905 K at λ = 8, 5.844 K at λ = 20, and 106.233 K at λ = 50.The same pattern appears at 10%, where RMSE reaches 4.538, 10.090, and 1.338 × 10^5 K at λ = 8, 20, and 50.
- The sweep supports tuning guidance jointly with observation density and likelihood scaling because observation fit can improve while hidden-cell performance collapses.
- Errors initially decrease and then rise sharply as observation force overwhelms the learned prior across the 1% and 5% guidance sweeps.The sweeps use logarithmic axes and compare against matched λ = 0 prior-only baselines.
- For 10% observed cells, λ = 0.5 is best, while stronger guidance rapidly degrades hidden-cell performance.
F Follow-up robustness checks
Follow-up checks support the stability of SDA’s advantage across masks and observation layouts, while revealing sensitivity to observation noise and guidance strength. The evaluation remains synthetic, with station and held-out-year tests still pending.
- Robustness results: SDA retains a macro advantage on a regular approximately 10-cell-spacing grid and under 0.5 K observation noise, although noise degrades its performance.The random-mask gain is also stable across two new seeds.
- Scope and next steps: Station representativeness, density-aware guidance, and genuinely held-out-year model selection remain necessary next evaluations toward deployment.The follow-up checks extend 2022 validation evidence but do not replace station or blind-year tests.
- Guidance-strength sensitivity: At 1% observation density, guidance-strength sweeps show conspicuous spatial artifacts from λ = 8 and extreme artifacts at λ = 50.The matched case displays posterior means and absolute errors on common scales.
- Evaluation protocol: Table 5 summarizes hidden-cell macro RMSE/MAE and per-case SDA wins across 32 cases for the follow-up 1% robustness checks.
- Observation-density behavior: Across 0.01%, 0.1%, and 1% observed cells, more observations recover finer structure, reduce hidden-cell error, and contract ensemble spread in better-constrained regions.