Source-linked AI summary

UHI-Bench: Benchmarking Dual-Source Urban Heat Island Modeling Across Cities in Diverse Climate Regimes

Wanyun Ling, Chenxi Liu, Yi Xie, Aopu Xu, Zhuoqi Zeng, Ziyue Li

arXiv:2608.23857v1cs.LG

TL;DR

Existing UHI studies often rely on one of two physically distinct signals, while data gaps and sparse AirT observations complicate joint modeling and cross-city evaluation. UHI-Bench builds a standardized dual-source benchmark with environmental context and evaluates models across cities, climate classes, and tasks. Results show that foundation models are competitive and stable, covariate utility is task- and source-dependent, and transferability aligns more with UHI-regime overlap than climate-label similarity.

  • Problem

    Most UHI studies use one observation source, while LST-UHI and AirT-UHI differ physically and remain difficult to align with dynamic and static environmental data.

  • Method

    UHI-Bench co-registers dual-source UHI, hourly meteorology, and static urban morphology, then evaluates over 20 baselines from four model families on five tasks across cities and climate classes.

  • Results

    Foundation models remain competitive and stable across tasks, while environmental covariate utility varies by source and task and transferability is better explained by UHI-regime overlap than Köppen climate-label distance.

  • Takeaways & Limitations

    Dual-source UHI signals are complementary rather than interchangeable, supporting task-aware environmental integration and UHI-regime-aware cross-city evaluation.

Abstract

from arXiv · show

Urban heat islands (UHIs) are intensifying under climate change, exacerbating thermal exposure risks. Their two primary observations, land surface temperature UHI (LST-UHI) and near-surface air temperature UHI (AirT-UHI), capture physically distinct aspects of urban heat. However, most studies rely on a single source, and substituting one for the other can substantially bias the magnitude and spatial variability of human heat exposure. Accurate UHI modeling also requires dynamic meteorological drivers and static urban morphology features, but spatiotemporal incompatibilities hinder their alignment. Cloud gaps in LST observations and sparse AirT station networks further limit dual-source UHI modeling, motivating cross-city transfer across diverse climates. To bridge these gaps, we introduce UHI-Bench, the first UHI benchmark for dual-source UHI modeling that integrates dynamic and static environmental context. Following a unified signal, mechanism, and transfer framework, it evaluates over 20 baselines from four model families on five tasks across 20 cities and nine Köppen climate classes. Results show that no model is uniformly best, although foundation models remain consistently competitive and stable. Environmental covariates generally improve performance, but their utility varies across sources and tasks. Cross-city transferability is better explained by overlap in UHI regimes than by climate-zone similarity. With the dataset and standardized pipeline, our work provides practical guidance for urban heat modeling, promotes climate data equity, and supports future advances in climate research.

1 Introduction

UHI-Bench addresses the need to model physically distinct LST-UHI and AirT-UHI signals jointly with environmental context and to assess transfer across cities and climates. It organizes this evaluation around signal validity, mechanism understanding, and cross-city transfer.

  • Motivation: LST-UHI and AirT-UHI respond differently to surface radiation, land cover, and atmospheric mixing, so substituting one for the other can bias human heat-exposure estimates.Existing studies often use satellite-derived LST-UHI instead of near-surface air temperature because it has broader spatial availability.
  • Motivation: Cloud gaps in LST-UHI and sparse representative station networks for AirT-UHI limit dual-source modeling, especially in data-poor regions.These constraints motivate testing whether models can transfer to unseen cities across diverse climate regimes.
  • Benchmark design: UHI-Bench integrates LST-UHI, AirT-UHI, hourly meteorological drivers, and static urban morphology on a standardized 1 km hourly grid.The main dataset covers 20 cities across 9 Köppen climate classes, including paired and source-specific city coverage.
  • Benchmark design: The benchmark evaluates over 20 baselines from four model groups on five tasks within a signal–mechanism–transfer framework.The framework covers dual-source signal validity, environmental mechanism understanding, and cross-city and cross-climate transfer.
  • Contribution: UHI-Bench is presented as the first benchmark supporting both LST-UHI and AirT-UHI modeling while integrating dynamic and static environmental data.Its stated goal is to support standardized evaluation across cities and climate classes.

2 Related Work

Prior urban-temperature resources expanded coverage and resolution but generally provide a single UHI source, absolute temperature, or geographically limited data. UHI-Bench is positioned to address the lack of multi-source alignment across cities and climate zones.

  • Existing datasets: Early UHI datasets typically focused on either satellite-based LST or station and gridded meteorological AirT observations.Because LST and AirT represent different physical layers, single-source datasets cannot systematically compare them.
  • Existing datasets: UrbClim provides hourly simulated AirT at 100 m for 100 European cities but lacks paired UHI targets and aligned covariates.Its geographic scope is limited to Europe.
  • Existing datasets: Global UHII and Global Heat Wave Exposure provide broad city- or daily heat estimates, but they do not supply the benchmark’s paired dual-source modeling context.The cited resources support large-scale analysis through monthly surface- and canopy-UHII or daily 1 km MODIS-based estimates.
  • Existing datasets: Other resources support fine-grained or city-specific temperature modeling, yet generally provide absolute temperature or only one UHI source.They do not align LST-UHI and AirT-UHI across multiple climate zones.
  • Benchmark gap: UHI-Bench differs by combining dual-source UHI with aligned environmental context across multiple cities and climate regimes.The comparison is framed against urban heat datasets with narrower source, city, or climate coverage.

3 UHI-Bench Data

UHI-Bench Data combines multi-source UHI observations, meteorological drivers, and urban morphology across cities and climate regimes on a shared 1 km hourly grid. Its construction includes downscaling, rural-reference UHI derivation, and standardized spatiotemporal alignment.

  • Coverage: The dataset spans 20 cities plus two supplementary station-based cities across 10 Köppen climate classes, four continents, and 15 countries.The main benchmark includes 16 dual-source core cities and four LST-UHI-only cities.
  • Coverage: Approximately 81,755 urban pixels are organized with LST-UHI, AirT-UHI, hourly meteorology, and static morphology on a standardized 1 km hourly grid.The four LST-UHI-only cities support surface-UHI transfer evaluation, while two supplementary cities provide station-format AirT observations.
  • LST-UHI construction: LST-UHI is produced by downscaling native approximately 3 km MSG/SEVIRI land-surface temperature to 1 km with an RF-TsHARP approach.Quality checks include round-trip consistency, comparison with MODIS 1 km LST, and expected relationships with built-up coverage, NDVI, and water bodies.
  • LST-UHI construction: LST-UHI is defined as local LST minus the contemporaneous mean LST of surrounding rural reference pixels.Cloud-contaminated observations are retained as missing values rather than filled.
  • Environmental context: Hourly meteorology comes from ERA5-Land, while static features include buildings, roads, POIs, nighttime lights, vegetation, water, elevation, building height, and wind exposure.These variables capture complementary aspects of surface cover, human activity, terrain, heat retention, and airflow.
  • Evaluation and release: The default split uses 2015–2022 for training and 2023–2025 for evaluation, with the LST-UHI-only cities using the same evaluation period.The released pipeline includes standardized dataloaders and baseline implementations and can be extended to other cities in the MSG/SEVIRI full-disk domain.

4 Experiments

Experiments evaluate signal validity, mechanism stability, imputation, forecasting, and climate-aware transfer across UHI sources and cities. Results show source-specific behavior, task-dependent covariate utility, and transfer governed more by UHI-regime overlap than climate labels alone.

  • Signal validity: LST-UHI and AirT-UHI are related but not interchangeable: spatial agreement exceeds same-hour temporal agreement, and lag correction is unreliable.AirT-UHI can support LST-UHI imputation in some cities, whereas LST-UHI does not consistently improve sparse AirT-UHI imputation.
  • Heat-risk timing and detection: AirT-UHI extremes are easier to detect than LST-UHI extremes, with XGBoost best for AirT-UHI and Chronos-2 best for LST-UHI.Meteorological features improve detection for both sources.
  • Imputation: LST-UHI imputation becomes difficult under cloud gaps, where GraphWaveNet is the strongest non-foundation baseline and spatial missingness matters beyond missing fraction.Chronos-2 achieves the best overall performance with meteorological features, while covariate benefits vary by model and climate.
  • Imputation: AirT-UHI sparse imputation is generally easier than LST-UHI cloud-gap imputation, with model preferences differing between tropical and temperate cities.XGBoost with static morphology performs best across missing-rate bins in Lagos, while Chronos-2 remains stable in Cfb cities.
  • Forecasting: XGBoost with meteorological and static features provides the most robust overall forecasting, although DLinear variants perform best for LST-UHI in some Cfb cities.Meteorological inputs generally outperform static features alone, but added covariates can degrade models that cannot exploit heterogeneous exogenous information.
  • Climate-aware transfer: Cross-city transfer benefits primarily from climate-diverse source sets, while directed transfer is better explained by UHI-regime overlap than Köppen climate-label distance.ChronosAdapter with Diverse-7 obtains average MAE 0.960, closely followed by XGBoost with Diverse-7 plus static inputs at 0.963.

5 Conclusion

UHI-Bench unifies dual-source UHI modeling with environmental context and evaluates it across diverse cities, climates, tasks, and model families. Its results emphasize stable foundation-model performance, task-dependent covariate utility, and UHI-regime overlap as a stronger transfer guide than climate-zone similarity.

  • UHI-Bench integrates LST-UHI, AirT-UHI, hourly meteorology, and static urban morphology features.
  • Over 20 baselines from four model families are evaluated on five tasks across 20 cities and nine Köppen climate classes.
  • Foundation models remain competitive and stable across tasks, although no model is uniformly best.
  • Environmental covariate utility varies across UHI sources and tasks, requiring task-aware integration of meteorology and urban morphology.
  • Cross-city transferability is better explained by overlap in UHI regimes than by climate-zone similarity, supporting source selection based on target UHI characteristics.

A Task 3b Transfer Diagnostics

Task 3b presents transfer-map diagnostics and directed climate-pair transfer results.

  • Task 3b includes transfer-map diagnostics and UHI-regime summaries.
  • Task 3b reports complete results for directed climate-pair transfer.

B Full Dataset Summary

The dataset summary documents UHI-Bench’s city coverage, climate indicators, meteorological variables, and supporting environmental data sources.

  • Table 7 provides the full dataset summary for UHI-Bench.
  • Grid meteorological inputs include wind components, total cloud cover, 2-m dew point, boundary-layer height, surface variables, and solar radiation.
  • The dataset summary includes cities such as Hamburg, Munich, and Dortmund, with records spanning 2015/01/01 to 2025/12/31.
  • Supporting data sources include DWD, ERA5, HOSTRADA, MSG/SEVIRI, GHSL, OpenStreetMap, and JRC Global Surface Water products.
  • Figure 10 compares Köppen–Geiger climate-classification indicator distributions for additional representative cities against benchmark-wide context.

C Related UHI Modeling Methods

UHI modeling has progressed from classical spatial-statistical methods to machine learning and spatiotemporal forecasting. Recent models capture nonlinear, temporal, and spatial structure, but are often tied to specific cities or target definitions.

  • Early UHI studies used correlation, regression, geostatistics, inverse-distance weighting, and kriging to quantify or interpolate urban–rural temperature differences.
  • Machine-learning methods such as RF and XGBoost were adopted for UHI prediction, imputation, and driver analysis.
  • Multivariate time-series and spatiotemporal models capture temporal patterns and spatial dependencies for forecasting and gap filling.
  • Recent model families include DLinear, Autoformer, PatchTST, iTransformer, STGCN, STID, DeepUHI, and Earth-observation models such as Prithvi-EO.
  • Existing methods are usually tied to specific cities or target definitions.

D Data Processing

UHI-Bench constructs LST-UHI and AirT-UHI on a common 1 km grid using rural-reference anomalies, source-specific temperature products, and consistent UTC alignment. LST is thermally sharpened from 3 km to 1 km, while international AirT uses residual correction constrained by German observations.

  • LST-UHI: LST-UHI uses 1 km downscaled LST minus contemporaneous mean LST over rural reference pixels.The native MSG/SEVIRI product is downscaled from nominal 3 km using RF-TsHARP; quality is assessed through round-trip consistency and MODIS comparison.
  • LST-UHI: Rural LST reference pixels lie 15–25 km from each city center, excluding built-up and water pixels and primarily selecting cropland and grassland.The reference construction uses great-circle distance and land-cover filtering.
  • AirT-UHI: German AirT-UHI uses HOSTRADA, whereas international AirT is generated with an ERA5-constrained residual model trained on German HOSTRADA−ERA5 residuals.The residual model uses meteorological, morphological, and temporal predictors and achieves held-out R^2 ≈0.79.
  • AirT-UHI: AirT-UHI is local air temperature minus mean air temperature over a surrounding 15–25 km rural annulus.The annulus uses the same distance definition as the LST reference set but without its land-cover restriction.
  • Co-registration: All source matching, model inputs, splits, and ERA5 joins use exact UTC hourly timestamps on a fixed 1 km city grid.This common spatiotemporal registration supports consistent integration of LST, AirT, and meteorological data.

E.1 Experimental Setup Details

The experiments organize UHI-Bench around signal, mechanism, and transfer questions, evaluating diverse model families and environmental-input configurations across controlled task settings. Tasks span source comparison, missing-data recovery, forecasting, driver attribution, and cross-city transfer.

  • Evaluation design: Experiments use representative city subsets selected by data availability, quality, missingness, and climate diversity, with temporal splits and leakage control.The benchmark is described as extensible to additional cities and climate zones.
  • Signal: The benchmark first tests whether LST-UHI and AirT-UHI offer consistent or complementary signals through association, imputation, extreme timing, and event detection.These analyses comprise Analysis 1a, Analysis 1b, and Task 1.
  • Mechanism: Tasks 2a–2d evaluate operational learning for cloud-gap LST imputation, station-sparse AirT imputation, 1–96 h forecasting, and driver attribution.Driver attribution removes location-, hour-, and meteorology-related effects as specified in the task design.
  • Transfer: Task 3 compares climate-homogeneous and climate-diverse source sets and tests directed transfer asymmetry among seven cities.The comparisons use matched source-set sizes and address transfer across climate and urban contexts.
  • Baselines and inputs: The benchmark compares statistical/geostatistical, classical machine-learning, deep spatiotemporal, and time-series foundation-model baselines.Input configurations are base, +m, +s, and +s+m, corresponding to UHI history, meteorology, static morphology, or both.

E.2 Evaluation Metrics by Task

UHI-Bench combines task-specific metrics, feature-gain analysis, static-feature aggregation, and cross-city evaluation protocols. Its evaluation materials cover event detection, timing, imputation, forecasting, driver consistency, and climate-diverse transfer.

  • Extreme-event detection: F1, MissRate, and FAR evaluate binary extreme-event detection using true-positive, false-positive, and false-negative counts.F1 is defined from TP, FP, and FN; MissRate and FAR quantify missed and false detections.
  • Extreme timing: Extreme-hour analysis measures each local hour’s share, the nighttime fraction, and the peak hour with the largest extreme-hour share.The analysis uses source- and city-specific extreme-hour sets and a predefined nighttime-hour set.
  • Covariate gain: ΔMAE compares feature-augmented and base models, with negative values indicating improvement after adding the feature block.The comparison is defined as MAE+feature − MAEbase.
  • Static features: Static urban morphology features are aggregated onto standardized 1 km reporting cells from raster and geometric source data.The feature definitions include intersection areas, centroids, building heights, distances, points of interest, and aligned raster values.
  • Benchmark reporting: Tables document meteorological drivers, model-family consistency, climate-diverse transfer, baseline groups, task protocols, city coverage, and metrics.Additional tables report complete event-detection, imputation, and forecasting results for multiple cities and baselines.
  • Forecasting: Forecasting tables report AirT-UHI/LST-UHI MAE across six displayed horizons and their average for multiple cities.The reported city coverage includes Munich, Berlin, Cologne, Lagos, Johannesburg, and Riyadh.
  • Driver attribution: Task 2d visualizes per-city signed top-three meteorological and static drivers separately for AirT-UHI and LST-UHI.Green and orange bars denote the two targets, with direction showing raw high-minus-low response sign and lengths normalized within target groups.
Loading 2608.23857v1…