Source-linked AI summary

A guide to state-space modeling of ecological time series

Marie Auger-Méthé, Ken Newman, Diana Cole, Fanny Empacher, Rowenna Gryba, Aaron A. King, Vianney Leos-Barajas, Joanna Mills Flemming, Anders Nielsen, Giovanni Petris, Len Thomas

arXiv:2002.02001v3stat.MEq-bio.PEq-bio.QM

TL;DR

Ecologists face complex fitting, comparison, and validation challenges when using state-space models for ecological time series. This paper reviews formulation, fitting, estimation, and validation tools, providing examples intended to help researchers develop and apply SSMs to their own data.

  • Problem

    The variety and complexity of SSM inference procedures can limit ecologists’ ability to formulate, fit, and evaluate their own models.

  • Method

    The paper reviews tools for formulating, fitting, diagnosing, and validating SSMs, including posterior predictive checks and guidance on addressing estimability problems.

  • Results

    The review and accompanying examples provide an extensive foundation for applying SSMs to ecological data and facilitating their formulation and fitting.

  • Takeaways & Limitations

    Researchers can use the guide’s tools to formulate SSMs, assess estimation problems, and select study-design or model-simplification strategies such as replication or covariates.

  • Takeaways & Limitations

    Cross-validation methods specifically designed for SSM dependency structures are few, restricted to some SSMs, and rarely used.

Abstract

from arXiv · show

State-space models (SSMs) are an important modeling framework for analyzing ecological time series. These hierarchical models are commonly used to model population dynamics, animal movement, and capture-recapture data, and are now increasingly being used to model other ecological processes. SSMs are popular because they are flexible and they model the natural variation in ecological processes separately from observation error. Their flexibility allows ecologists to model continuous, count, binary, and categorical data with linear or nonlinear processes that evolve in discrete or continuous time. Modeling the two sources of stochasticity separately allows researchers to differentiate between biological variation (e.g., in birth processes) and imprecision in the sampling methodology, and generally provides better estimates of the ecological quantities of interest than if only one source of stochasticity is directly modeled. Since the introduction of SSMs, a broad range of fitting procedures have been proposed. However, the variety and complexity of these procedures can limit the ability of ecologists to formulate and fit their own SSMs. We provide the knowledge for ecologists to create SSMs that are robust to common, and often hidden, estimation problems, and the model selection and validation tools that can help them assess how well their models fit their data. In this paper, we present a review of SSMs that will provide a strong foundation to ecologists interested in learning about SSMs, introduce new tools to veteran SSM users, and highlight promising research directions for statisticians interested in ecological applications. The review is accompanied by an in-depth tutorial that demonstrates how SSMs models can be fitted and validated in R. Together, the review and tutorial present an introduction to SSMs that will help ecologists to formulate, fit, and validate their models.

1 Introduction

SSMs are hierarchical models for ecological time series that separate hidden process variation from observation error. This review addresses the growing complexity of fitting, validating, and selecting such models.

  • They are widely used for population dynamics, fisheries, biodiversity, movement ecology, and biologging data.
  • SSMs represent hidden ecological states and separate process variation from observation error in observed time series.Their hierarchical structure includes an unobserved state series and an observation series linked to those states.
  • Advances in MCMC and computing expanded SSMs beyond linear Gaussian models to non-Gaussian and nonlinear formulations.
  • Fitting some SSMs is computationally burdensome, while their complex structures complicate model comparison, validation, and diagnostics.
  • Fitting-procedure variety and potential estimability problems can limit ecologists’ ability to formulate, fit, and evaluate SSMs independently.
  • The review surveys SSM flexibility, inference methods, estimability, model selection, diagnostics, and R-based fitting and validation tools.

2 Examples of ecological SSMs

SSMs provide a flexible framework for ecological time series with hidden states and imperfect observations. Examples span observation structures, process dynamics, time scales, and statistical distributions.

  • SSMs can model univariate or multivariate observations and biological processes evolving in discrete or continuous time.
  • The framework supports linear or nonlinear processes and distributions including normal, Poisson, and multinomial forms.
  • Examples drawn from population and movement ecology illustrate structural flexibility, while SSMs can be applied across ecological fields.

2.1 A toy example: normal dynamic linear model

The toy example introduces a normal dynamic linear model for univariate observations and hidden states at evenly spaced times. It establishes SSM assumptions, equations, parameters, states, and inference targets.

  • The toy NDLM models univariate observations and hidden states at discrete, evenly spaced times using linear equations and Gaussian distributions.
  • The state process is generally first-order Markov, so each state depends on the previous state.
  • Observations are conditionally independent once their corresponding hidden states are accounted for.
  • Process variation and observation error are modeled separately with normal distributions and distinct standard deviations.
  • Fitting estimates fixed parameters and unobserved states, with frequentist and Bayesian approaches available.Bayesian inference adds prior distributions for the fixed parameters.
  • The toy model serves as a teaching and R-fitting example but is not particularly useful for ecology because of its simplicity.

2.2 Handling nonlinearity

Ecological SSMs can represent nonlinear population dynamics and incorporate covariates while retaining separate process and observation components. Examples show that accounting for observation error can improve inference about density dependence.

  • A stochastic density-dependent population model uses nonlinear exponential dynamics for the hidden abundance state.
  • The model parameters describe baseline growth, density dependence, and random annual changes in growth rate.
  • Log-scale reconfiguration can produce a linearized Gompertz model, but alternative formulations may remain nonlinear.
  • External covariates such as annual pond availability can be added to the process equation.
  • Accounting for observation error decreased bias in density-dependence estimates and reduced erroneous conclusions about density dependence.

2.3 Joining multiple data streams

SSMs can combine independent data streams that observe shared ecological states, while also linking those states to otherwise difficult-to-estimate processes. The stock-assessment example integrates catches and survey indices with demographic and mortality dynamics.

  • Multiple data streams can offset individual limitations and reveal more complex ecological relationships when integrated in one SSM.
  • The stock-assessment state vector combines log-transformed stock sizes and fishing mortality rates for each age class.
  • Recruitment, survival, and mortality are represented through process equations linking current states to previous-year states.
  • Fishing-mortality variation is correlated across age classes using an AR(1) covariance structure, whereas recruitment and survival variation is assumed uncorrelated.
  • Catch and scientific-survey observations use separate equations and can have distinct covariance structures across ages and surveys.
  • Catch data provide direct information about fishing mortality, while both catch and survey data inform common stock-size states.

2.4 Accounting for complex data structure

SSMs can represent ecological observations with irregular timing, heterogeneous quality, and outliers rather than discarding difficult data. The DCRW example adapts movement processes and measurement errors to Argos location data.

  • Argos locations have large errors, irregular observation times, and quality ratings that the DCRW incorporates directly.Reported mean observation errors range from 0.5–36 km, and outliers are common.
  • The DCRW models true locations at regular time steps and makes each location depend on the previous location and previous displacement.
  • The process equation uses γ to control step-to-step correlation, with values near 1 indicating persistence in movement speed and direction.
  • The observation equation linearly interpolates the true location to irregular observation times and uses t-distributed measurement errors for outliers.
  • SSMs can account for unexplained outliers and differing data quality within the model instead of arbitrarily discarding observations.

2.5 Accommodating continuous-time processes

Continuous-time SSMs represent biological processes evolving between observations and naturally accommodate irregularly timed data. The movement example uses a continuous-time correlated random walk, integrates velocity to location, and remains Kalman-filterable in its linear Gaussian form.

  • Continuous-time process equations facilitate modeling biological processes that occur continuously and observations collected at irregular times.
  • The continuous-time movement model represents velocity autocorrelation with β and adds a random perturbation over interval ∆.
  • As the time interval increases, velocity depends less on its previous value and more on random perturbation; short intervals support continued speed.
  • Integrating velocity over each interval links the continuous-time process to location observations and permits state tracking at any nonnegative time interval.
  • The resulting linear Gaussian SSM can be fitted with a Kalman filter.
  • Continuous-time models are useful for unequal observation intervals and intrinsically continuous ecological processes, though they can be more complex to understand.

2.6 Integrating count and categorical data streams

SSMs can jointly model count, categorical, movement, health, and survival data for individual animals. In the right-whale example, this integration supports inference about health metrics, survival differences among zones, and seasonal movement.

  • A joint SSM integrates count and categorical data to model right-whale health, monthly movement, and survival.
  • Photographic observations provide whale sightings by geographic zone and month alongside ordinal visual health metrics.
  • Three process equations describe each whale’s health, survival, and monthly movement, with health depending on previous health and age.
  • Survival probability depends on health and occupied zone, allowing identification of geographic zones associated with reduced survival.
  • Seasonally varying transition probabilities model migration among nine geographic zones.
  • The model supports inference about which visual metrics track underlying health, how regions relate to survival, and when whales move to particular zones.

2.7 Capturing heterogeneity with random effects

State-space models use random effects to represent individual and temporal heterogeneity in capture and survival probabilities, while retaining a mechanistic capture-recapture structure.

  • Model structure: Capture-recapture SSMs represent observations as capture indicators and states as whether each individual is dead or alive.At first capture, the alive state is fixed; subsequent states follow process and observation equations.
  • Random effects: The survival transition probability is zero after death, while survival and capture probabilities are multiplied by the relevant state values.This mechanistic structure prevents an individual from surviving after its state becomes dead.
  • Model structure: Survival and capture probabilities can vary across individuals and sampling occasions rather than being constrained to single overall values.Such variation may reflect differences in environmental and body conditions.
  • Random effects: Random effects model heterogeneity in survival and capture probabilities with only two additional parameters, whereas fixed temporal effects require many more.The fixed-effects formulation requires (T −1) + (T −2) additional parameters.
  • Implications: SSMs are commonly used for capture-recapture data because their mechanistic structure accommodates additional complexity and random-effects variation.The same random-effects strategy can be applied in other ecological applications.

2.8 Modeling discrete state values with hidden Markov models

Hidden Markov models are SSMs with discrete states, enabling ecological analyses of categorical latent conditions such as survival status or behavioral modes.

  • Definition and applications: Hidden Markov models are a special class of SSMs whose states are discrete, usually categorical with finitely many possible values.They have been used for capture-recapture data and for animals switching between behavioral modes.
  • Transition structure: In the dead-alive example, staying dead has probability 1 and resurrecting has probability 0.For living individuals, dying has probability 1 −φ_i,t−1 and surviving has probability φ_i,t−1.
  • Transition structure: Other discrete-state SSMs can permit transitions among all states and combine discrete and continuous states.A movement model used two behavioral modes for animals tracked with Argos data.

3 SSMs as a framework for ecological time series

SSMs provide a flexible framework for ecological time series by separating process variation from observation error and supporting diverse ecological applications.

  • Framework: SSMs support ecological time-series analysis by modeling temporal dependence and distinguishing process variation from observation error.Their hierarchical structure represents hidden process states separately from observations.
  • Ecological applications: 27,000 individual trees were modeled with SSMs for growth, fecundity, and survival, illustrating their use in plant ecology.Other plant applications include plant cover and canopy-process estimation from imperfect monitoring.
  • Ecological applications: SSMs have been used in paleoecology to model temporal changes in mammal mass and stable isotopes alongside climate and community covariates.Tomé et al. used three separate linear Gaussian SSMs.
  • Cyclicity: SSMs can outperform cyclicity tests: simpler tests produced Type I error rates as high as 79%, whereas the SSM had an appropriate 5% rate.The approach also helped identify mechanisms underlying ecological cycles.
  • Ecosystem applications: The best SSM formulation improved accuracy and reduced bias in estimates of gross primary production, respiration, and gas exchange.In some cases, its bias was half as large as that of simpler models.
  • Scope and terminology: SSM terminology is broad: some occupancy models resemble capture-recapture SSMs but lack the temporal autocorrelation typically associated with the process equation.Some complex ecological models also contain essential SSM structure without being labeled SSMs.
  • Model complexity: SSMs are particularly advantageous over simpler models when both process variance and observation error are large.Simpler models can be adequate when one stochastic source is small and the dynamics are not too complex, provided the model is correctly specified.

4 Fitting SSMs

Fitting SSMs requires methods for estimating parameters and hidden states under frequentist or Bayesian inference, with algorithm choice governed by model structure and computational demands.

  • Goals and approaches: SSM analyses commonly estimate both model parameters and hidden states, whose uncertainty can be summarized with point, interval, or variance measures.State estimates are often interpreted as better representations of true locations or ecological states than raw observations.
  • Goals and approaches: Frequentist methods maximize likelihood, whereas Bayesian methods focus on posterior density; both confront high-dimensional integration.The many fitting tools are different solutions to this integration problem.
  • Frequentist fitting: Frequentist fitting can estimate parameters by maximizing the marginal likelihood after integrating out hidden states, then estimate states conditionally.Marginal-likelihood estimates have consistency and asymptotic-normality properties, unlike direct joint optimization in expanding state dimensions.
  • Frequentist fitting: The Kalman filter is fast and easy to calculate, supporting broad ecological use for linear Gaussian SSMs.It does not work directly with nonlinear and non-Gaussian SSMs, though approximate variants cover some cases.
  • Monte Carlo methods: Sequential Monte Carlo methods approximate filtering distributions by simulating and weighting particles, offering flexibility for any SSM at substantial computational cost.Simple sequential importance sampling can suffer particle depletion and large state-estimate variances.
  • Bayesian fitting: Hamiltonian Monte Carlo can improve ecological SSM efficiency and diagnostics, but it does not easily handle discrete parameters or latent states.Stan is a prominent HMC implementation available in R through rstan.
  • Bayesian fitting: Bayesian prior selection affects posterior distributions, and nominally noninformative priors can substantially influence ecological SSM parameter and state estimates.Informative priors can instead incorporate parameter knowledge or previously collected data.

5 Formulating an appropriate SSM for your data

Appropriate SSM formulation requires matching model complexity to the available data, checking identifiability and estimability, and reformulating confounded models. The review presents diagnostics and remedies for detecting and addressing these problems.

  • 5 Formulating an appropriate SSM for your data: SSM flexibility can produce models too complex for the available data, making parameters unreliable and state estimates inaccurate.The model structure and dataset jointly determine whether parameters can be estimated reliably.
  • 5.1 Identifiability, parameter redundancy and estimability: A model is identifiable when its parameterization has a unique representation; global identifiability requires M(θ1) = M(θ2) to imply θ1 = θ2.Local identifiability instead concerns whether uniqueness holds within a neighborhood of θ.
  • 5.1 Identifiability, parameter redundancy and estimability: Parameter redundancy creates non-identifiability when parameters appear only through combinations such as products that can be replaced by a single parameter.For example, y = αβx can be reparameterized as y = γx, where γ = αβ.
  • 5.1 Identifiability, parameter redundancy and estimability: Structural identifiability does not ensure estimability: sparse or missing data can cause practical non-identifiability or statistical inestimability.Practical non-identifiability has an infinite confidence interval, whereas statistical inestimability has an extremely large but finite one.
  • 5.1 Identifiability, parameter redundancy and estimability: Non-identifiability can yield nonunique maximum-likelihood estimates, undefined standard errors, and invalid complexity penalties in model selection.Numerical optimization may still converge to one apparent estimate, while numerical Hessians can produce explicit but incorrect standard errors.
  • 5.1 Identifiability, parameter redundancy and estimability: Identifiability and estimability should be checked during fitting using likelihood profiles, parameter-correlation checks, and simulations comparing estimates with known values.Simulations generate state and observation series from a specified SSM and then compare estimated parameters and states with their true values.
  • 5.1 Identifiability, parameter redundancy and estimability: Advanced diagnostics include data cloning, Hessian eigenvalue analysis, symbolic methods, and hybrid symbolic-numerical rank calculations.Data cloning uses K copies of the data, while the Hessian method flags zero or near-zero eigenvalues as evidence of non-identifiability or parameter redundancy.
  • 5.2.1 Reformulate the SSM: Reliable fitting requires both a structurally identifiable model and a dataset appropriate for that model; reformulation should remove confounded products, sums, differences, and fractions.The symbolic method can identify confounded parameters and help select estimable parameter combinations, including among additive variability sources.

6 Computationally-efficient model comparison methods

The paper reviews information criteria and related methods for comparing ecological state-space models, emphasizing computational trade-offs and limitations caused by latent states and temporal dependence.

  • Model comparison evaluates competing hypotheses and can refine state and parameter estimates because different structures produce different estimates.
  • Frequentist approach: AIC penalizes model likelihood by the number of estimated parameters to reduce overfitting, but latent states complicate the effective parameter count.These concerns are less problematic when compared models have the same number of states and no additional random effects.
  • Frequentist approach: For small samples, AICb can outperform AIC and AICc but requires fitting the model repeatedly, whereas AICc remains a practical fallback for computationally demanding models.AICc may still favor overly complex models more often than AICb.
  • Bayesian approaches: WAIC uses the full posterior distribution and is generally favored over DIC for Bayesian comparison, but naive data partitioning can be problematic for dependent time series.The paper identifies marginal-likelihood WAIC as the best current information-criterion option for Bayesian SSMs, while noting that cross-validation may avoid several shortcomings when models are inexpensive to fit.
  • Bayesian approaches: WAIC and related information criteria can provide unreliable predictive-ability estimates for ecological models, so further research is needed on appropriate use and data partitioning.

7 Diagnostics and model validation for SSMs

SSM diagnostics are essential because relative model selection does not establish absolute fit, while temporal dependence and unobserved states complicate conventional checks. The review presents posterior predictive checks, residual diagnostics, and cross-validation tools, while noting important limitations for dependent ecological time series.

  • Why validate SSMs: Relative model selection does not establish absolute fit, so selected SSMs may still poorly represent ecological or measurement processes.AIC and similar relative measures do not quantify how closely a model matches the data.
  • Challenges with SSMs: Diagnostics are difficult because observations are temporally dependent and the hidden states generally lack direct observations.These features complicate residual-based checks and direct comparisons between predicted states and their true values.
  • Residual diagnostics: Response residuals can remain serially dependent, impairing detection of misspecification and potentially producing inflated goodness of fit.They are based on predicted observations conditioned on all observations, which contributes to their dependence.
  • Posterior predictive checks: Posterior predictive checks compare observed data with replicate time series simulated from the fitted model using test quantities or discrepancy functions.The procedure uses simulated data, observed data, model parameters, and a statistic such as the mean or a χ2 measure.
  • Residual diagnostics: One-step-ahead residuals should be temporally independent when the model is adequate because each prediction uses observations only up to the previous time.This distinguishes them from response residuals, which use all observations.
  • Cross-validation: Cross-validation directly assesses predictive ability and can rank models by prediction error, but SSM dependence complicates data removal and error estimation.Removing few observations can underestimate prediction error, whereas removing many can propagate error; refitting also makes cross-validation computationally demanding.
  • Cross-validation: Few cross-validation methods specifically handle SSM dependence, and existing approaches apply only to a restricted set of models.Time-series and block cross-validation methods are identified as starting points for developing better SSM-specific procedures.

8 Conclusion

SSMs offer flexible ecological time-series modeling while separating process variation from observation error, but effective use requires careful formulation, fitting, selection, and validation. The review provides guidance and examples for these tasks and identifies further research needs.

  • Model flexibility: SSMs can represent univariate or multivariate, linear or nonlinear, discrete- or continuous-time series with normal or non-normal stochasticity.This flexibility supports continuous, count, binary, and categorical data.
  • Fitting methods: Fitting tools now support models that better represent ecological processes and data structures in both frequentist and Bayesian frameworks.The available methods differ in computational efficiency and restrictions.
  • Model formulation: Model formulation should address estimability problems, with replication or covariates potentially reducing common estimation difficulties.The review emphasizes that complex fitting tools do not replace appropriate model design.
  • Selection and validation: Model selection and validation should become part of every SSM workflow because misspecification can affect ecological inferences and state-estimate accuracy.AIC, WAIC, posterior predictive measures, and one-step-ahead residuals are discussed as useful tools.
  • Future research: Future research should improve understanding of when simpler alternatives suffice, optimize data-collection designs, and develop computationally efficient procedures for selecting and validating SSMs.The authors also call for cross-validation methods that account for dependencies in the data.
  • Contribution: The review and Appendix S1 provide methods and examples intended to facilitate formulation and fitting of SSMs for ecological data.The authors hope this guide supports application and further development across ecology.

Figure captions

The figures illustrate SSM dependence structures, state estimation, movement-model validation, identifiability profiles, and diagnostic comparisons. Together, they connect model components with practical checks of estimation and fit.

  • Figure 1: Panel a shows observations conditionally independent after accounting for states, while panel b compares simulated observations and states with estimated states and 95% confidence intervals.The true states usually fall within the intervals, unlike the observations.
  • Figure 2: The DCRW figure maps Argos observations, estimated locations, and GPS locations, then displays longitude and latitude subsets to highlight temporal clustering.The caption notes that clustering likely helped the state-estimation procedures.
  • Figure 3: The log-likelihood profiles contrast globally identifiable, locally identifiable, and non-identifiable models for parameter θ.These scenarios correspond to distinct forms of parameter uniqueness or non-uniqueness.
  • Figure 4: Figure 4 contrasts diagnostic plots for correctly specified and misspecified models, including predictive checks, residual autocorrelation, and residual distributions against a standard normal density.The misspecified model assumes σo = 0.5 instead of the simulated σo = 0.1.
Loading 2002.02001v3…