Source-linked AI summary
AIMIP Phase 1: systematic evaluations of AI weather and climate models
Brian Henn, Christopher S. Bretherton, Nikolay Koldunov, Christian Lessig, Maria J. Molina, Troy Arcomano, Oliver Watt-Meyer, Guillaume Couairon, Renu Singh, Robert Brunstein, Yana Hasson, Antonia Jost, Noah Brenowitz, Peter Manshausen, Nathaniel Cresswell-Clay, Dale Durran, Kyle Joseph Chen Hall, Janni Yuval, Dmitrii Kochkov, Stephan Hoyer, Ignacio Lopez-Gomez
TL;DR
AIMIP Phase 1 addresses the challenge of predicting climate trends from historical information and physical knowledge by defining a common AI weather and climate model intercomparison. The project finds similarities across models in reproducing ERA5 climate patterns and simulating ENSO responses, while physically plausible out-of-sample responses remain a challenge.
Problem
Predicting future climate trends from historical information and reliable physical knowledge remains a key challenge.
Method
AIMIP Phase 1 defines specifications for an intercomparison and trains AI weather and climate models on ERA5 reanalysis data from 1979 onward.
Results
AIMIP models more faithfully replicate ERA5 climate patterns than a conventional CMIP6 model and strongly simulate the ENSO response to specified conditions.
Takeaways & Limitations
A common experiment and data format enables the community to leverage rapid advances in AI model development.
Takeaways & Limitations
Physically plausible responses in extreme out-of-sample perturbed SST experiments remain difficult to produce reliably from historical information.
Abstract
from arXiv · showhide
We present the AI weather and climate model intercomparison project (AIMIP), phase 1. Drawing from the rich tradition of intercomparisons in climate model development, we specify a common experiment, output data format, and training constraints (namely, training against historical reanalysis data) for AIMIP Phase 1 models. We aim to identify differences in modeling frameworks and AI architectural choices that influence model behavior, and build trust in AI weather and climate models through open data and evaluation. AIMIP Phase 1 models must simulate the atmosphere given specified historical sea surface temperatures over 1979-2024. We evaluate the models' performance using five major evaluation criteria: biases, trends, response to El Niño-related sea surface temperature anomalies, temporal variability, and out-of-sample generalization tests. We find that the AI models are able to simulate the historical climate and response to forcing as well as a conventional physically-based model, but some AI models underestimate historical warming trends, and their predictions diverge in the out-of-sample generalization tests. We describe the AIMIP Phase 1 dataset that is publicly available for additional evaluations.
1 Introduction
AIMIP Phase 1 establishes a coordinated intercomparison for AI weather and climate models, modeled on AMIP and using shared historical forcing, formats, and evaluation tools. It targets longer-term climate simulations and broader climate-community engagement while preparing for future coupled AI models.
- AIMIP focuses on longer-term simulations than weather-oriented evaluations and seeks to socialize AI climate models within the broader climate-science community.
- AIMIP Phase 1 is an intercomparison of AI weather and climate models that describes a common protocol and initial results.
- Models are trained on global reanalysis and forced by historical sea-surface temperature and sea-ice patterns during 1979–2024.
- The project adapts the AMIP and CMIP tradition to evaluate AI models using established climate-science analysis methodologies and output formats.
- Future AIMIP phases may evaluate interactively coupled AI atmosphere, ocean, and sea-ice models across future climates.
- Standardized outputs are intended to enable use of CMIP evaluation software and open comparison across models.
2 AIMIP Phase 1 goals and protocol
AIMIP Phase 1 defines goals, training constraints, simulations, standardized outputs, and evaluation boundaries for comparing AI-driven atmospheric models. The protocol emphasizes ERA5-only training, historical SST and SIC forcing, common formats, and explicit out-of-sample tests.
- Goals: AIMIP compares time-mean climate, trends, variability, and weather phenomena in multi-decadal simulations trained exclusively on ERA5 and forced by historical SST and SIC.
- Goals: The project seeks CMIP7-compatible outputs so the broader climate-science community can comprehensively evaluate AI weather and climate models.
- Scope: The protocol permits autoregressive, hybrid, and conditional-sampling architectures rather than prescribing one AI modeling framework.
- Scope: AIWCM evaluation targets long-term climate statistics that emerge from aggregates of many atmospheric snapshots, making direct training on them computationally expensive.
- Training and evaluation: Models train on ERA5 from 1979–2014 and are evaluated against ERA5, whose precipitation and some climate trends are acknowledged to be poorly constrained by observations.
- Experiments: Standard simulations run from late 1978 through 2024, including a training period and a 2015–2024 out-of-sample test period, with prescribed monthly SST and SIC.
- Constraints: The protocol excludes CO2 and other anthropogenic radiative forcings as training inputs to reduce proxy-clock overfitting, while acknowledging that SST trends can still encode forcing-related effects.
- Experiments: Uniformly warmer SST experiments test generalization, but they lack definitive ground truth and may not predict near-historical utility.
3 Participating Models
AIMIP Phase 1 includes eight AI weather and climate model submissions from six groups, spanning autoregressive, probabilistic, diffusion, hybrid, and latent-diffusion approaches. Models differ in temporal resolution, conditioning, grids, ensemble construction, and methods for preserving variability.
- Participation: AIMIP Phase 1 received eight AIWCM submissions from six model-development groups.
- ACE2.1-ERA5: ACE2.1-ERA5 is a 6-hourly autoregressive SFNO emulator with learned secondary diagnostics and five ensemble members.
- ACE2.1-ERA5: ACE2.1-ERA5 uses historical SST forcing by overwriting prognostic surface temperature over majority-ocean grid cells.
- Arches models: ArchesWeather and ArchesWeatherGen use daily averaged ERA5 data, add SST and SIC as prognostic variables, and differ in deterministic versus probabilistic forecasting.
- Arches models: ArchesWeatherGen generates ensemble members by changing noise seeds, while ArchesWeather uses varied initial-condition dates.
- cBottle-1.3: cBottle-1.3 is a conditional diffusion model that generates atmospheric snapshots on an approximately 0.9° HEALPix grid with eight snapshots per day.
- cBottle-1.3: cBottle-1.3 uses independent samples at each inference, so its monthly means are expected to have lower variance than those from an autoregressive model.
- cBottle-1.3: Correlated-noise diffusion inference introduces temporal correlation through an AR1 latent process whose autocorrelation halves over 24 hours.
4 AIMIP Phase 1 Evaluation Results
AIMIP Phase 1 evaluates AI weather and climate models across biases, trends, ENSO response, temporal variability, and perturbed-SST responses using standardized regridded comparisons. AIWCMs often match or outperform GFDL-CM4 on historical climate metrics, but models differ in trend fidelity, variability, and out-of-sample responses.
- E1: Biases: AIWCMs typically achieve lower RMSB versus ERA5 than GFDL-CM4, although performance varies across models and variables.Several AIWCMs consistently produce lower RMSB than others at the surface and in the upper atmosphere.
- E1: Biases: Test-period biases are generally higher than training-period biases, while models differ in how well they generalize to the test period.Some models better capture the elevated 2015–2024 warming than others, including NeuralGCM, ArchesWeather, ArchesWeatherGen, and DLESyM.
- E2: Trends: Most AIWCMs reproduce the sign of training-period trends but tend to underestimate their magnitude, especially for Arctic and land warming.Trend error patterns differ by model, and many AIWCMs do not reproduce ERA5 precipitation trends associated with an intensified central-Pacific convergence zone.
- E3: ENSO response: AIWCM ENSO coefficient errors are generally small for tropical-ocean temperature responses, with precipitation errors also small in training and modestly larger in testing.Most AIWCMs produce slightly smaller ENSO coefficient errors than CM4 relative to ERA5.
- E4: Daily variability: AIWCMs usually underestimate daily anomaly magnitudes for non-precipitation variables, whereas precipitation variability errors differ across models and often involve overestimation over wet regions.Temperature anomaly errors are smaller over ocean than land because SST is specified and slowly varying; DLESyM is the exception to the broad precipitation overestimation pattern.
- E5: Perturbed SST response: Perturbed-SST responses vary much more across AIWCMs than historical forced comparisons, although NeuralGCM-HRD and DLESyM closely match CM4’s expected temperature-warming response.The perturbation experiments constitute a strongly out-of-sample test, and AIWCM responses span a broad range.
5 Discussion
AIMIP Phase 1 models perform similarly across several historical-climate metrics and often capture ENSO responses and variability despite diverse frameworks. However, models diverge in warming trends and extreme perturbed-SST responses, underscoring challenges for unseen-scenario guidance while demonstrating the feasibility of common intercomparison protocols.
- AIMIP Phase 1 models perform quite similarly across multiple metrics, including ENSO responses and daily variability, despite using diverse modeling frameworks and AI architectures.The evaluated frameworks include autoregressive full emulation, hybrid physics/AI models, conditional diffusion, U-Nets, Fourier operators, and vision transformers.
- Most AIWCMs produce lower climate biases against ERA5 than the physically based GFDL-CM4 model, which was not directly tuned against ERA5.
- Warming-trend performance varies substantially: most AIWCMs underestimate trends to some degree, while models differ in capturing training- and test-period warming.ACE2.1-ERA5’s inability to capture warming differs from a similar AMIP-type experiment, likely because AIMIP excludes CO2 forcing, whose historical rise correlates strongly with warming trends.
- Physically plausible responses to extreme out-of-sample perturbed-SST experiments vary widely between models, and SST preprocessing choices alone do not explain the differences.Models use constant filling, SST interpolation, or merged ocean and land/sea-ice fields; architecture and sensitivity to large-scale input perturbations may also contribute.
- AIWCMs may struggle with the primary climate-model goal of providing guidance for unseen scenarios, so their ability to achieve this goal remains a work in progress.
- AIMIP demonstrates that a common experimental specification and output format can be achieved across diverse AIWCM training methods, supporting future intercomparison phases and model development.Future phases may include coupled Earth-system experiments, more extensive outputs, and additional evaluation metrics.
6 Conclusions
AIMIP Phase 1 defines a common intercomparison experiment, data format, and evaluation framework for AI weather and climate models. Across five metrics, models show shared strengths but also substantial divergence, especially in warming trends and perturbed-SST responses; the resulting dataset is released for further evaluation.
- Evaluation: Eight AIWCM submissions from six groups are evaluated across biases, trends, ENSO response, daily variability, and out-of-sample perturbed-SST response.The models span multiple simulation frameworks and AI architectures.
- Main findings: AIWCMs more faithfully reproduce ERA5 climate patterns than a conventional CMIP6 model and strongly simulate ENSO responses to specified SSTs.These similarities are reported across the intercomparison, rather than for one individual architecture.
- Main findings: The models nevertheless show a wide range of behavior in both in-sample and out-of-sample warming trends, with major disagreement in strongly perturbed-SST experiments.The submissions also differ in their ability to capture training-period and out-of-sample warming trends in ERA5.
- Implications: The authors present the AIMIP Phase 1 dataset and common format to support further evaluation and experimentation.They frame the submissions as snapshots of ongoing development rather than fixed model references.
Appendix A: AIMIP Phase 1 monthly SST and SIC dataset
The AIMIP Phase 1 monthly SST and SIC dataset was created for 1979–2024 inference runs because existing AMIP forcing was unsuitable. It uses centered monthly averaging and linear interpolation, avoiding overshoots while preserving annual means, and produces a climate nearly identical to daily forcing in ACE2.1-ERA5.
- Limitations: The conventional AMIP algorithm can create SIC overshoots and negative intermediate values, producing annual-mean SIC biases that affect temperature in some seasonal-ice grid cells.ERA5 also contains small SIC and land-mask incompatibilities near polar coastlines and lake boundaries.
- Dataset motivation: The compact AIMIP Phase 1 forcing dataset covers 1979–2024, extending beyond the existing AMIP forcing period to support longer observational comparisons.It was created because the conventional AMIP dataset did not extend far enough for the AIMIP simulations.
- Construction: The dataset derives monthly SST and SIC forcing from daily ERA5 data using centered averaging between monthly midpoints.The forcing is based on ERA5’s 0.25° latitude–longitude grid.
- Construction: SST and SIC values at intermediate times are obtained by linear interpolation, avoiding overshoots and preserving each forcing field’s annual time mean.Individual monthly means are not exactly preserved.
- Validation: In ACE2.1-ERA5 inference runs, the monthly forcing produces a climate nearly identical to daily forcing, including in seasonal sea-ice zones.Modeling groups may spatially interpolate the forcing to their native grids.
Appendix B: cBottle1.3 physics indices checkpoints
Appendix B documents the checkpoints and ensemble design used for cBottle1.3 physics realizations. cBottle1.3 uses an Ensemble-of-Experts denoising setup, with different networks assigned to noise ranges to reduce overfitting at high noise levels.
- Checkpoint selection: The cBottle physics realizations use multiple listed training checkpoints to generate the AIMIP Phase 1 simulation ensemble.The appendix enumerates checkpoints such as training-state-000512000 and training-state-009984000.
- Model design: cBottle1.3 is an Ensemble-of-Experts model in which different networks perform different parts of denoising.The setup uses three networks associated with noise levels below 10, between 10 and 100, and above 100.
- Model design: Less-trained or early-stopped networks handle higher noise levels to avoid overfitting at large noise levels.The appendix identifies this as the rationale for assigning different network versions across noise ranges.
- Latent-space variant: One listed realization uses temporally uncorrelated latent space, making every sample fully independent.This configuration is identified as p5.
C1 Biases
Appendix C1 reports model bias diagnostics for 2-meter air temperature over training and test periods and RMSB diagnostics for several atmospheric variables and pressure levels.
- 2-meter temperature: The appendix compares 2-meter air-temperature biases across each model ensemble during the training and test periods.These diagnostics are shown in Figs. C1 and C2.
- Pressure-level diagnostics: RMSB is evaluated at 1° resolution across seven pressure levels for air temperature, specific humidity, and eastward and northward wind.These diagnostics are shown in Figs. C3–C6.
C2 Trends
The appendix extends trend and bias diagnostics across model ensemble members, atmospheric pressure levels, and training versus test periods. It also includes globally averaged ENSO coefficient errors.
- Annual- and global-mean 2-meter air-temperature series are shown for each AI weather and climate model, including ensemble members.
- Temperature trends are evaluated over pressure levels during both the training and test periods.
- Specific-humidity trends are likewise evaluated over pressure levels during the training and test periods.
- The appendix includes 2-meter air-temperature bias diagnostics over the 1979–2014 training period for each model ensemble member.
- Globally averaged ENSO coefficient errors are reported as an additional diagnostic.
C4 Daily variability
The appendix examines daily variability through dry-day fraction errors, ENSO coefficient errors, and atmospheric-variable diagnostics across pressure levels and periods. These evaluations compare model outputs with ERA5 where specified.
- Dry-day fraction errors are evaluated against ERA5 over 1979 at 1° resolution for models submitting daily surface data.A 0.1 mm precipitation cutoff defines a dry day.
- RSMB diagnostics cover eastward and northward wind across pressure levels and training and test periods.
- Additional diagnostics show ensemble-member annual and global mean 2-meter air-temperature series and pressure-level trends for air temperature and specific humidity.
- Globally averaged ENSO coefficient errors are included as a variability-related diagnostic.
- The dry-day analysis uses ERA5 error maps alongside model-versus-ERA5 error panels.
Appendix D: Selected results at 2.8◦resolution
Selected appendix results repeat the main diagnostics at 2.8° resolution, using NeuralGCM instead of NeuralGCM-HRD. They include biases, RSMB, trends, ENSO coefficient errors, and perturbation responses.
- The selected results use 2.8° resolution and substitute NeuralGCM for NeuralGCM-HRD.
- Bias maps and RSMB diagnostics are shown at 2.8° resolution, corresponding to the main bias and RSMB figures.
- Global-mean trend maps are presented at 2.8° resolution.
- ENSO coefficient errors are shown at 2.8° resolution.
- Perturbation responses are also shown at 2.8° resolution.