Source-linked AI summary

Predictability of El Niño from Delayed Observations

Francisco J. Beron-Vera

arXiv:2608.24428v1physics.ao-phcs.LGmath.DSnlin.CD

TL;DR

How much predictive information is contained in delayed Niño-3.4 observations, and does added model complexity improve forecasts? The paper compares explicit delayed-coordinate models with recurrent networks through July 2026, finding that compact representations improve skill while complexity provides no systematic gain.

  • Problem

    The study asks whether a small collection of past Niño-3.4 anomalies provides useful predictive coordinates, what memory ranges emerge, and whether forecast skill saturates as delays increase.

  • Method

    Using monthly Niño-3.4 anomalies through July 2026, the study compares ridge, MLP, SINDy, GRU, and LSTM forecasts using explicit or internally learned temporal representations.

  • Results

    Forecast skill saturates with a small number of delayed coordinates, while nonlinear and recurrent complexity provides no systematic improvement over compact explicit representations.

  • Takeaways & Limitations

    The results support a compact predictive representation of Niño-3.4 evolution in which representing recent history matters more for prediction than increasing model complexity.

  • Takeaways & Limitations

    The existence of a predictive map for the finite lag sets considered is not implied by Takens’ embedding theorem and is assessed empirically.

Abstract

from arXiv · show

Using monthly Niño-3.4 anomalies through July 2026, we investigate how much predictive information is contained in delayed observations of the index. Ridge regression identifies informative delays, while multilayer perceptron and sparse identification of nonlinear dynamics (SINDy) models test whether nonlinear complexity provides additional direct forecast skill; gated recurrent unit (GRU) and long short-term memory (LSTM) networks provide a complementary test in which the temporal representation is learned internally. Delayed observations substantially improve forecasts over persistence and climatology at leads of up to six months, but increasing model complexity provides no systematic improvement. Historical recursive experiments favor a simple explicit SINDy recurrence and select shallow recurrent architectures, with no appreciable gain from learning the temporal representation internally. These results support a compact predictive representation of Niño-3.4 evolution in which the representation of past information is more consequential than model complexity. As a prospective application, the selected models are used to forecast the developing 2026 event beyond the last available observation and to compare its predicted evolution with completed historical El Niño events.

Plain Language Summary · Key Points · 1 Introduction

The study examines whether delayed Niño-3.4 observations contain predictive information about ENSO and whether nonlinear or recurrent model complexity adds forecast skill. Using monthly observations through July 2026, it finds that a compact delayed-coordinate representation is more consequential than increasing model complexity.

  • 1 Introduction: ENSO is a dominant source of interannual tropical climate variability and worldwide climate-anomaly predictability, shaped by coupled ocean–atmosphere feedbacks across time scales.Conceptual ENSO models emphasize memory and delayed negative feedback.
  • 1 Introduction: The study tests whether past Niño-3.4 anomalies provide predictive information without assuming that finite delays reconstruct the ENSO state or correspond to specific oceanic processes.Delayed scalar observations can under appropriate conditions provide coordinates for an underlying dynamical state.
  • 1 Introduction: The study’s data-driven approach complements prior ENSO prediction research that often exploits spatially distributed oceanic and atmospheric information for long-lead forecasts.Related work includes convolutional neural networks trained on tropical Pacific sea-surface-temperature fields and deep-learning approaches identifying oceanic precursors.
  • 1 Introduction: The forecasting framework uses ridge regression to identify informative delays, then tests nonlinear maps with MLP and SINDy models.The approach follows a delayed-coordinate forecasting framework.
  • 1 Introduction: GRU and LSTM models provide a complementary comparison in which the temporal representation is learned internally rather than specified through selected delays.This comparison distinguishes complexity in state representation from complexity in the forecasting map.
  • 1 Introduction: Forecast skill saturates with a small number of delayed coordinates, while nonlinear and recurrent complexity provides no systematic improvement using monthly Niño-3.4 data through July 2026.The result supports focusing on the representation of past information rather than simply increasing model complexity.

2 Methods

The study models monthly Niño-3.4 anomalies from December 1949 through July 2026 using explicit delayed coordinates or internally learned recurrent representations. It evaluates linear, nonlinear, and recurrent models against persistence and zero-anomaly climatology in direct and recursive forecasting settings.

  • Data: Monthly Niño-3.4 sea-surface-temperature anomalies N(t) span December 1949 through July 2026 and are averaged over 5°S–5°N, 170°W–120°W.The index comes from NOAA Climate Prediction Center ERSSTv6 data and is used directly rather than the three-month-averaged ONI.
  • Delayed-coordinate models: Ridge regression searches all admissible combinations of candidate delays {0, 1, 2, 3, 4, 6, 9, 12, 15} months, always retaining the instantaneous coordinate.The best-performing lag set is retained for each number of coordinates, and delayed-coordinate usefulness is assessed empirically rather than inferred from Takens’ theorem.
  • Nonlinear models: MLP and SINDy test nonlinear dependence using ridge-selected coordinates, with a single-hidden-layer MLP and polynomial SINDy libraries through degree three.SINDy constrains its candidate-polynomial coefficient vector to be sparse, while deeper MLPs are formed by composing additional layers.
  • Recurrent models: GRU and LSTM learn temporal dependence internally from consecutive past observations, carrying information through hidden states rather than selected explicit delays.GRUs use update and reset gates, whereas LSTMs additionally maintain a cell state controlled by input, forget, and output gates.
  • Evaluation and forecasting protocol: Forecast skill is evaluated with chronological 70%/30% training/test splits for ridge and SINDy, 70%/15%/15% splits for neural networks, and comparisons against persistence and zero-anomaly climatology.Values below unity indicate improvement over the corresponding reference; direct open-loop forecasts use observed delayed coordinates, whereas closed-loop forecasts recursively replace unavailable observations with predictions beyond July 2026.

3 Results

Forecast skill is driven primarily by a compact representation of delayed Niño-3.4 observations rather than by nonlinear or recurrent model complexity. Historical recursive tests favor a simple linear recurrence and shallow recurrent architectures, whose prospective forecasts closely agree through October 2026.

  • At forecast leads ∆= 1, 3, and 6 months, the best ridge models use 4, 6, and 5 coordinates, respectively, with skill saturating after a few delays.Lag sets include short delays and, at longer leads, additional delays of 9–15 months.
  • An MLP using the same delayed coordinates does not systematically improve forecasts, while an instantaneous MLP yields Rp = 1.009, 0.958, and 0.863 at leads ∆= 1, 3, and 6 months.The gain therefore comes from the delayed representation rather than nonlinear flexibility in the forecast map.
  • Similarity-weighted three-month RMSEs for degree-one, degree-two, and degree-three SINDy recurrences are 0.514 ◦C, 125.6 ◦C, and 12.3 ◦C, respectively, favoring the linear recurrence under iteration.The larger nonlinear-map errors result from occasional recursive instabilities, despite similar direct skill at ∆= 1 month.
  • Shallow GRU and LSTM models select h = 32 and h = 16, with RMSE 0.530 ◦C and 0.524 ◦C, close to degree-one SINDy’s 0.514 ◦C.Shallow models have lower aggregate errors and perform better for five of six origins nearest the July 2026 initialization.
  • For July 2026, SINDy and selected shallow recurrent models keep the anomaly near its July value through August–September before weakening in October, while deep models weaken faster.The preferred forecast most resembles the 1965–66 El Niño among the 20 completed historical events.

4 Conclusions

The Niño-3.4 record contains substantial predictive information in a compact set of delayed observations spanning short and longer memory scales, while added nonlinear complexity and internally learned temporal representations provide no systematic improvement. Recursive tests favor simple SINDy and shallow recurrent models, whose closely agreeing forecasts are applied prospectively to the developing 2026 event.

  • Predictive representation: Ridge regression identifies substantial predictive information in a compact set of delayed observations spanning short and longer memory scales.The predictive information saturates with this compact coordinate set.
  • Model complexity: MLP and SINDy models show no systematic forecast gain from adding nonlinear complexity.The conclusion contrasts increased nonlinear complexity with the compact delayed-observation representation.
  • Temporal representation: Historical recursive experiments show that GRU and LSTM networks learning temporal representations internally do not improve appreciably over explicit delayed coordinates.The experiments therefore favor explicit representation of delayed observations rather than internally learned temporal representations.
  • Recursive robustness: Recursive robustness tests favor a simple linear SINDy recurrence over higher-degree polynomial maps.This preference is supported by recursive robustness testing.
  • Recurrent architecture selection: Pseudo-real-time hindcasts select shallow GRU and LSTM architectures, which have lower similarity-weighted historical errors than selected deep architectures.The comparison is based on historical pseudo-real-time hindcast errors.
  • Prospective 2026 application: SINDy and shallow recurrent models give closely agreeing short-term forecasts for the developing 2026 event, while deep recurrent models illustrate forecast sensitivity.These models are applied prospectively beyond the available observations.
Loading 2608.24428v1…