Source-linked AI summary
Missing the Butterfly and Predicting the Past: Features or Bugs of Accurate AI Weather Models?
Pedram Hassanzadeh, Weidong Li, Y. Qiang Sun, Jiangdi Wang, Alexander Wikner, Justin Finkel, Jonathan Q. Weare
TL;DR
The paper asks why AI weather models are highly accurate despite unclear physical fidelity and missing butterfly effects. Across reanalysis, a general circulation model, and Lorenz 96, it tests forecasting and backcasting and finds that coarse-grained training data jointly explain forecast skill, skillful backcasting, and missing butterflies. Reducing coarse-graining makes predictions more physics-like but reduces forecast accuracy.
Problem
AIWP models have surprising forecast accuracy and unclear physical fidelity, including skillful backcasting and missing butterfly effects that conflict with established chaotic-dynamics understanding.
Method
The study compares forecasting and time-reversed backcasting across ERA5, an intermediate-complexity GCM, and multi-scale Lorenz 96 while varying the coarse-graining of training data.
Results
Coarse-graining removes fast, small scales, yielding slower large-scale error growth and higher forecast accuracy while producing skillful backcasting and suppressing butterfly-like effects.
Takeaways & Limitations
Reducing coarse-graining makes AI predictions more physics-like, with stronger forecast–backcast asymmetry and emerging butterfly-like effects, but forecast accuracy declines.
Takeaways & Limitations
ERA5 remains highly coarse-grained, missing state variables and subgrid processes, and butterfly analyses require caution about hardware artifacts and nonphysical forecasts.
Abstract
from arXiv · showhide
AI weather prediction (AIWP) models rival physics-based models, yet the sources of their unexpected forecast accuracy and the degree of their physical fidelity remain unclear. Here, across a hierarchy spanning observation-based reanalysis, a general circulation model, and the multi-scale Lorenz system, we show that AI models can be trained to skillfully predict the past (backcast), though backcasts are systematically less accurate than forecasts. However, skillful backcasting appears to violate the second law of thermodynamics, and all these forecasting and backcasting models miss the butterfly effect. We trace the surprising forecast accuracy, missing butterfly, and skillful backcasting to a single cause: inevitable coarse-graining of training data, which removes fast, small scales and/or some variables. From the Lorenz system to official Pangu-Weather models, reducing coarse-graining makes AI predictions more physics-like (arrow of time and butterfly-like effects emerge), but forecast accuracy declines. Results offer an explanation for AIWP models' forecast skill: unlike physics-based models, they implicitly learn how fast, small scales affect large scales without inheriting their rapid error growth. Broader implications are that AI models' proliferation calls for revisiting predictability theories and long-term climate emulation strategies, and backcasting offers a useful, new lens for such analyses.
1 Introduction
AI weather models achieve strong and efficient forecasts, but their physical fidelity and sources of accuracy remain unclear. The paper investigates these issues through backcasting and links forecast skill, missing butterfly effects, and skillful backcasting to coarse-grained training data.
- Motivation: AIWP models outperform leading physics-based weather models across common skill metrics at much lower computational cost.They forecast from global three-dimensional atmospheric states and can generate longer predictions autoregressively.
- Motivation: Despite their forecast skill, AIWP models struggle with unprecedented gray-swan extremes, small-scale statistics, and the butterfly effect.The missing butterfly effect concerns rapid ensemble-spread growth from vanishing initial-condition perturbations.
- Approach: The paper trains time-reversed neural networks to predict the past, providing a new dimension for studying AIWP behavior and atmospheric predictability.Backcasting uses the forecasting model’s reversed input–output pairing with independent learnable parameters.
- Scope and findings: Across reanalysis, a general circulation model, and multi-scale Lorenz 96, the study finds skillful backcasting, missing butterfly effects, and surprising forecast accuracy.These behaviors are presented as symptoms of one underlying cause and are examined as potential features or bugs of AIWP models.
2 Results
Across ERA5, PlaSim, and Lorenz 96, AI models skillfully backcast and miss the butterfly effect, patterns linked to coarse-grained training data. Reducing coarse-graining makes predictions more physics-like but reduces forecast skill and can produce nonphysical behavior.
- 2.1 Skillful AI backcasting and persistent asymmetry: ERA5 and PlaSim (GCM): 6.5 days versus 9.1 days: ERA5 backcasts remain skillful less long than forecasts, despite backward numerical integrations failing catastrophically.The forecast/backcast asymmetry is 9.1/6.5 > 1, while atmospheric dissipation makes time reversal unstable.
- 2.1 Skillful AI backcasting and persistent asymmetry: ERA5 and PlaSim (GCM): PlaSim shows larger forecast-backcast asymmetries than ERA5 across deterministic, generative, and neural-operator AIWP models, variables, pressure levels, and skill metrics.The asymmetry is also overall larger in the tropics, where dissipative processes are stronger.
- 2.2 Missing butterflies in AI forecasting and backcasting alike: ERA5 and PlaSim (GCM): Physics-based ICON and PlaSim integrations show amplitude-dependent butterfly growth, whereas ERA5 and PlaSim AIWP models lack this effect in both forecasting and backcasting.Small-amplitude numerical ensembles grow by orders of magnitude within a day or two, while AI models do not show the corresponding accelerating spread.
- 2.3 Training-data coarse-graining dramatically alters forecasts, backcasts, and butterflies: Lorenz 96: Perfect-data Lorenz 96 models have forecast skill worse by a factor of ∼7 and asymmetry larger by a factor of ∼4, while backcasting blows up and rapid error growth emerges.These models see all scales in the numerical ground truth and become more consistent with multi-scale chaotic dynamics.
- 2.3 Training-data coarse-graining dramatically alters forecasts, backcasts, and butterflies: Lorenz 96: Coarse-graining—filtering small or fast scales and/or ignoring state variables—is proposed as the common cause of unexpected forecast skill, missing butterflies, and skillful backcasting.The trends persist from the real-world-regime corner to the perfect-data-regime corner, while complete data restore comparable large- and small-scale Lyapunov exponents and butterfly-like growth.
- 2.4 As coarse-graining is reduced, a butterfly-like effect emerges and forecast skill degrades: Pangu-Weather: Reducing Pangu-Weather’s time step from 24 h to 1 h changes small-amplitude spread from slow, amplitude-independent growth to rapid, amplitude-dependent growth.At 1 h, perturbation energy grows fastest at small scales before progressively filling larger scales, although late-time small-scale variance overshoots ERA5.
- 2.4 As coarse-graining is reduced, a butterfly-like effect emerges and forecast skill degrades: Pangu-Weather: The 1-h Pangu-Weather behavior should not be identified with the real butterfly effect because ERA5 remains highly coarse-grained and omits state variables and unresolved processes.The authors also caution that hardware artifacts and nonphysical forecasts require control and assessment.
3 Discussion
Across ERA5, PlaSim, and Lorenz 96, the paper identifies spatial and temporal coarse-graining in training data as the common source of AI models’ forecast accuracy, missing butterfly effect, and skillful backcasting. More complete training data make predictions more physics-like but reduce large-scale forecast skill, exposing a trade-off between practical accuracy and physical fidelity.
- Coarse-graining is the common cause of surprising forecast accuracy, the missing butterfly effect, and skillful backcasting across ERA5, PlaSim, and Lorenz 96.The paper frames these as H1, H2, and H3, respectively.
- As training data more completely represent the true system, backcasting skill disappears, forecast–backcast asymmetry grows, and butterfly-like effects emerge while practical forecast accuracy declines.The same architecture and loss can therefore become more physics-like without remaining equally accurate for large-scale forecasting.
- Removing fast, small scales slows large-scale error growth while allowing AI models to learn their averaged effects implicitly as subgrid-scale parameterizations.This differs from numerical models, where higher resolution can reduce discretization and explicit-parameterization errors.
- The missing butterfly effect depends on training-data content rather than architecture, loss function, or deterministic-versus-generative design.Lorenz 96 networks with the same architecture and MSE loss differed only according to their training data, and Pangu-Weather showed the same temporal dependence.
- Skillful backcasting arises because coarse-graining produces a smoother, less irreversible effective system with smaller leading Lyapunov exponents and fewer fast scales.Residual irreversibility remains visible through forecast–backcast asymmetry and eventual unstable or unphysical long-term backcasts.
- For weather forecasting, coarse-graining can be a feature because it supports accurate large-scale predictions; for atmospheric physics and long-term emulation, it is a bug because information is irreversibly lost.The paper points toward representing unresolved effects through memory and stochasticity, including Mori–Zwanzig-inspired approaches.
A Methods and Data
The paper documents its datasets, models, supplementary materials, and evaluation framework across observation-based, numerical, and climate-model settings.
- The section begins by stating that detailed dataset, model, and metric descriptions follow.
- The Methods and Data section covers ERA5, ICON, PlaSim, two-scale Lorenz 96, AI models, and evaluation metrics.
- Supplementary Materials discuss time reversal, Lorenz 63 phase-space visualizations, and Transformer-, SFNO-, and Diffusion-based architectures.
A.1 Observation-based data: ERA5
ERA5 experiments use reanalysis-driven AIWP Transformers, with separate forecasting and backcasting models trained at 1° resolution and evaluated through autoregressive ensembles and perturbation analyses.
- ERA5 supplies five upper-air variables on 17 pressure levels plus selected surface variables as the prognostic state.
- The experiments train separate Pangu-Weather-based Transformer forecasters and backcasters on 40 years of 6-hourly data, with 6- and 24-hour prediction intervals.
- Backcasting reverses the forecast input–output pairing and differs from meteorological hindcasting, reforecasting, and non-autoregressive backward sampling.
- Rollouts are generated autoregressively after single-time-step training, with each predicted state fed back as the next input.
- Training uses latitude-weighted MAE with differential upper-air and surface weighting, while optimization uses AdamW and OneCycleLR selected by one-step validation MAE.
- Sixteen-member ensembles use independent Perlin-noise perturbations of initial conditions sampled monthly during 2020 and 2021.
- All experiments use FP32 with TF32 disabled, and DKE growth is tested across averaging domains and perturbation types.
- High-latitude instabilities in the 1-hour Pangu-Weather model complicate clean separation of learned dynamics from unphysical behavior.
A.2 Numerical weather prediction (NWP) model: ICON
ICON provides the physics-based NWP reference, using an operational global model that solves discretized governing equations on an icosahedral grid.
- ICON is an operational global model from the German Weather Service and Max Planck Institute for Meteorology.
- The study uses ICON ensemble DKE data at 2.5 km and 20 km horizontal resolutions.
A.3 Atmosphere-only GCM: PlaSim
PlaSim supplies an intermediate-complexity atmosphere-only GCM setting for comparing Transformer, SFNO, and Diffusion forecasts and backcasts under shared perturbation and training procedures.
- PlaSim couples a spectral atmospheric dynamical core solving primitive equations to a simplified land model with annually repeating SST and sea-ice boundary conditions.
- The study trains separate PlaSim models at 6 and 24 hours for Transformers and at 24 hours for SFNO and Diffusion.
- Transformer and SFNO models receive forcings and static fields by channel concatenation, whereas Diffusion uses cross-attention context plus calendar-time conditioning.
- Diffusion models represent predictions as samples from a learned conditional distribution for both forecasting and backcasting.
- Transformer training minimizes latitude-weighted MAE, SFNO minimizes latitude-weighted MSE, and Diffusion uses latitude-weighted noise-prediction MSE.
- Diffusion training injects noise across 100 diffusion steps using a cosine noise schedule.
- PlaSim AI ensembles use Perlin perturbations, while quasi-deterministic Diffusion fixes the sampling seed so spread comes from initial conditions.
- Physics-based PlaSim ensembles perturb spectral surface-pressure coefficients with 10^-8 white noise, advance at 20-minute steps, and discard six hours for spin-up.
A.4 Canonical multi-scale, chaotic system: Two-scale Lorenz 96
The two-scale Lorenz 96 system provides a chaotic, multiscale testbed with slow large-scale X variables and faster small-scale Y variables. Independent forecast and backcast AI models predict the full state across several prediction intervals and are evaluated against numerical dynamics.
- System and data: Lorenz 96 uses 8 slow variables X_i and 32 fast variables Y_i,j, forming a 264-dimensional chaotic state.Standard parameters include K=8, J=32, F=20, h=1, and β=γ=10; one model time unit is approximately five days.
- System and data: The forward dynamics are chaotic and dissipative, with dissipation 10 times stronger at the smaller scales.This scale-dependent dissipation makes the system suitable for studying multiscale predictability and time reversal.
- System and data: Training data are generated by fourth-order Runge–Kutta integration at δt=0.005 MTU after a 25-MTU spin-up, with 10^7 subsequent states saved.The data are partitioned sequentially into training and validation segments, followed by an independent test trajectory.
- AI models: Independent multilayer perceptrons forecast or backcast the full state x at t±Δt, using Δt=nδt for n∈{1,5,10}.The models use ReLU activations and are trained separately for the two time directions.
- AI models: Models vary whether they receive only X or the concatenated state (X,Y), and all are trained by minimizing mean squared prediction error.The study also varies the prediction interval and uses extensive hyperparameter optimization and ensemble perturbations for evaluation.
- Evaluation: Predictability analysis estimates finite-time Lyapunov exponents separately for X and Y using perturbation rescaling over finite windows.The projected growth rate uses only the selected variable components, allowing large- and fast-scale predictability to be distinguished.
A.5 Evaluation metrics
Forecast and backcast skill are assessed with trajectory-accuracy metrics, while ensemble spread is measured separately through difference kinetic energy. The study averages these quantities over many test-set initial conditions to obtain lead-time-dependent curves.
- Evaluation procedure: All metrics are averaged over evaluation cases from many test-set initial conditions to produce lead-time-dependent skill curves.Deterministic metrics use predicted and true states, while DKE uses ensemble-member variances.
- Trajectory accuracy: ACC measures area-weighted pattern similarity between predicted and target anomalies relative to climatology.ACC primarily reflects large-scale forecast accuracy, and the conventional weather-prediction skill threshold is ACC>0.6.
- Trajectory accuracy: Forecast-backcast asymmetry is defined as the ratio of forecast to backcast lead times at ACC=0.6 for ERA5 and PlaSim.An asymmetry greater than 1 indicates that backcasting is less skillful than forecasting.
- Trajectory accuracy: RMSE quantifies prediction error between predicted and ground-truth values.For Lorenz 96, RMSE uses uniform weights over the eight large, slow X variables, with asymmetry defined as the backcast/forecast RMSE ratio at a fixed 1.5-day lead.
- Ensemble spread: DKE measures ensemble spread rather than error against a target by averaging half the ensemble variance of horizontal wind components.It is computed about the instantaneous ensemble mean and is reported globally at 300 hPa unless otherwise specified.
- Ensemble spread: For Lorenz 96, DKE is defined analogously as half the sum of ensemble variances of the large-scale variables.The metric uses the instantaneous ensemble mean of the large-scale state.
B.1 Time reversal in multi-scale, dissipative, chaotic systems and the 2nd law of thermodynamics
Time reversal is intrinsically unstable for dissipative, multiscale systems because discarded small scales and reversed dissipation amplify errors. The paper argues that coarse-grained AI data can nevertheless support skillful backcasting because the learned dynamics operate on resolved, attractor-supported states.
- Time reversal: For the heat equation, forward perturbations decay most rapidly at small scales, whereas reversing time amplifies them as εe^(κk^2t*).The smallest scales therefore dominate the backward solution almost immediately, making the backward heat problem Hadamard ill-posed.
- Time reversal: Lorenz 63 reverses its Lyapunov spectrum, changing approximately from (0.91,0,-14.57) forward to (14.57,0,-0.91) backward.The reversed leading exponent is therefore about 16 times larger in magnitude than the forward leading exponent.
- Time reversal: In Lorenz 96, reversed anti-dissipation causes the fast Y variables to grow 10 times more strongly and drag the large-scale X variables through coupling.Numerical backcasting consequently diverges because the fast variables blow up first.
- Coarse-graining and irreversibility: Skillful AI backcasting does not violate the second law because coarse-graining removes scales where the strongest dissipation occurs and leaves a dataset closer to reversible.The models operate on resolved, attractor-supported states rather than reconstructing the discarded microscopic degrees of freedom.
- Coarse-graining and irreversibility: The residual forecast-backcast asymmetry measures the irreversibility that survives coarse-graining and remains in the training data.This residual includes dynamical dissipation and, for the atmosphere, thermodynamic entropy production.
- Coarse-graining and irreversibility: Mori–Zwanzig formalism describes coarse-graining deterministic multiscale dynamics as producing resolved-scale memory and stochastic forcing.This connects coarse-grained deterministic data with stochastic modeling and parameterization approaches.
B.2 More details on neural network architectures
The supplementary material documents the architectures, variable inventories, hyperparameters, and diagnostic figures used across ERA5, PlaSim, Pangu-Weather, and Lorenz systems. These materials examine forecast-backcast asymmetry, butterfly-effect behavior, scale-dependent errors, and one-step spectral fidelity.
- Architectures: The Transformer uses a 3D patch embedding and hierarchical Swin-style shifted-window attention encoder-decoder to reconstruct the prognostic state.Its tokens represent 2×2×2 blocks across pressure level, latitude, and longitude.
- Architectures: The SFNO performs learned transformations in spherical spectral space using spherical harmonic transforms and degree-dependent filters.This design respects Earth’s spherical geometry rather than operating directly on a latitude-longitude grid.
- Architectures: The Diffusion model learns to reverse Gaussian noising through iterative denoising with a conditional score-based vision-transformer architecture.Its forward noising process uses 100 diffusion steps with a cosine noise schedule.
- Data and diagnostics: Table S2 reports slow- and fast-variable leading finite-time Lyapunov exponents estimated over 20 initial conditions and a 0.1–0.3 MTU window.The exponents are computed with perturbation rescaling and projected separately onto X and Y variables.
- Data and diagnostics: Table S3 inventories upper-air and surface prognostic variables, time-varying boundary forcings, and static fields across Pangu-Weather, ERA5, and PlaSim models.The pretrained Pangu-Weather model is used for inference on its native 0.25° grid, while the other columns describe trained AIWP datasets.
- Supplementary comparisons: Figures S1–S4 test missing butterfly behavior and forecast-backcast asymmetry across SFNO, Diffusion, and Transformer models, variables, and prediction intervals.Figures S6–S8 additionally examine scale-dependent Pangu-Weather errors and one-step kinetic-energy spectra.
- Butterfly-effect diagnostics: Figure S5 compares Lorenz 63 DKE growth for numerical and AI models across perturbation amplitudes using 10,000 initial conditions and 100-member ensembles.The AI model closely tracks the numerical DKE curves, while both lack the real butterfly effect.
Caption for Movie S1. Numerical forecasts and backcasts on the Lorenz 63 attractor.
The animation numerically integrates Lorenz 63 states forward and backward from 10 slightly different initial conditions on the attractor. Forecast trajectories remain on the attractor but diverge significantly over the long term.
- 10 slightly different initial conditions on the Lorenz attractor are integrated numerically both forward and backward.The animation uses the same numerical methods as those used for Lorenz 96.
- Forecast trajectories remain on the attractor but diverge significantly in the long term.The trajectories remain there due to phase-space contraction.