Source-linked AI summary

Avoidable errors in the modeling of outbreaks of emerging pathogens, with special reference to Ebola

Aaron A. King, Matthieu Domenech de Cellès, Felicia M. G. Magpantay, Pejman Rohani

arXiv:1412.0968v5q-bio.PEstat.AP

TL;DR

The paper asks how outbreak transmission models can estimate key quantities without producing misleadingly precise parameters and forecasts when early clinical data are scarce. It combines simulation experiments with stochastic SEIR modeling of early West Africa Ebola incidence, finding that raw-data stochastic approaches better quantify uncertainty and diagnose model misspecification than deterministic models fitted to cumulative data.

  • Problem

    Early outbreaks often lack sufficient clinical and household transmission data, increasing reliance on incidence-model estimates whose modeling and data choices can produce serious errors.

  • Method

    The paper uses simulations and fits a stochastic SEIR partially observed Markov model to raw Ebola incidence data using maximum likelihood, sequential Monte Carlo, and iterated filtering.

  • Results

    Stochastic models fit to raw data reduced bias, better quantified parameter and forecast uncertainty, and more readily diagnosed lack of model fit than deterministic models fit to cumulative data.

  • Takeaways & Limitations

    Models should use raw, disaggregated data when possible, prefer stochastic formulations, and compare simulated with actual data to detect misspecification.

  • Takeaways & Limitations

    The paper does not assert that the SEIR model adequately captures all epidemic features needed for accurate forecasts.

Abstract

from arXiv · show

As an emergent infectious disease outbreak unfolds, public health response is reliant on information on key epidemiological quantities, such as transmission potential and serial interval. Increasingly, transmission models fit to incidence data are used to estimate these parameters and guide policy. Some widely-used modeling practices lead to potentially large errors in parameter estimates and, consequently, errors in model-based forecasts. Even more worryingly, in such situations, confidence in parameter estimates and forecasts can itself be far over-estimated, leading to the potential for large errors that mask their own presence. Fortunately, straightforward and computationally inexpensive alternatives exist that avoid these problems. Here, we first use a simulation study to demonstrate potential pitfalls of the standard practice of fitting deterministic models to cumulative incidence data. Next, we demonstrate an alternative based on stochastic models fit to raw data from an early phase of 2014 West Africa Ebola Virus Disease outbreak. We show not only that bias is thereby reduced, but that uncertainty in estimates and forecasts is better quantified and that, critically, lack of model fit is more readily diagnosed. We conclude with a short list of principles to guide the modeling response to future infectious disease outbreaks.

Introduction

Early outbreak response increasingly relies on transmission models fitted to incidence data because key epidemiological quantities are difficult to estimate from rare clinical and household transmission data. Model structure and data choice can nevertheless produce serious errors, making model fit and uncertainty important for science and policy.

  • Introduction: Transmission models are increasingly used to estimate key epidemiological quantities and support real-time public health policy.Examples of model-based forecasts have been produced within weeks of an index case.
  • Introduction: Clinical and household transmission data are rare early in outbreaks, so incidence reports become an important basis for estimating epidemiological quantities.These quantities include total outbreak size and the durations of incubation and infectious phases.
  • Introduction: The choice of fitted model and data can lead to potentially serious errors in estimates and forecasts.The paper examines these issues using simulated data and early West Africa Ebola Virus Disease data.
  • Introduction: Because model-based conclusions hinge on model fit, evidence of model misspecification should be sought alongside uncertainty quantification.The paper presents stochastic modeling as a straightforward way to diagnose misspecification and quantify forecast uncertainty.

Deterministic models fit to cumulative incidence curves: a recipe for error and overconfidence

Fitting deterministic transmission models to cumulative incidence can make estimates appear precise while misrepresenting measurement noise, parameter uncertainty, and forecast uncertainty. A simulation study shows that nominal confidence intervals can have severely inadequate coverage despite limited average bias in R0.

  • Deterministic models fit to cumulative incidence curves: a recipe for error and overconfidence: Deterministic least-squares fitting treats discrepancies from model trajectories as independent, normally distributed measurement errors with constant variance.Cumulative-data fitting additionally creates dependence among successive measurement errors.
  • Deterministic models fit to cumulative incidence curves: a recipe for error and overconfidence: R0 estimates showed considerable error but little evidence of bias when either raw or cumulative incidence data were used.For early exponentially growing outbreaks, both raw and cumulative curves carry growth-rate information determined by R0.
  • Deterministic models fit to cumulative incidence curves: a recipe for error and overconfidence: Cumulative-data fitting produced extreme downward bias in estimated measurement-noise overdispersion because smooth accumulated incidence appears to require little noise.Raw-incidence fitting recovered the true observation variability.
  • Deterministic models fit to cumulative incidence curves: a recipe for error and overconfidence: Nominal 99% confidence intervals for R0 achieved coverage closer to 25% when cumulative data were used.The intervals were narrower with cumulative data, creating an illusory appearance of precision.
  • Deterministic models fit to cumulative incidence curves: a recipe for error and overconfidence: Deterministic cumulative-data fitting underestimates uncertainty by ignoring correlated errors, underestimating measurement-noise variance, and omitting demographic and environmental stochasticity.The omitted stochasticity also makes forecast uncertainty grow unrealistically slowly with forecast horizon.

Stochastic models fit to raw incidence data: feasible and transparent

The paper fits a stochastic SEIR model to raw Ebola incidence data using likelihood-based sequential Monte Carlo methods. This approach is feasible, yields more credible uncertainty than cumulative-data fitting, and enables direct diagnostics revealing discrepancies between model simulations and observed temporal structure.

  • Stochastic models fit to raw incidence data: feasible and transparent: A stochastic SEIR model formulated as a partially observed Markov process was fit to early 2013–2015 West Africa Ebola data.Parameters were estimated by maximum likelihood using sequential Monte Carlo and iterated filtering.
  • Stochastic models fit to raw incidence data: feasible and transparent: Raw-data fitting produced somewhat larger confidence intervals than cumulative-data fitting, with a substantial increase in forecast uncertainty.The paper interprets the narrower cumulative-data intervals as probably illusory because the true parameters are unknown.
  • Stochastic models fit to raw incidence data: feasible and transparent: The stochastic model was fit and its likelihood profiles computed without being particularly difficult or time-consuming.This supports using stochastic models even during rapidly unfolding emerging-disease outbreaks.
  • Stochastic models fit to raw incidence data: feasible and transparent: Likelihood profiles showed that cumulative-data fitting changed the R0 maximum-likelihood estimate relatively little but created an illusion of high precision.Profile curvature represents estimated precision, and the horizontal line marks the 95% likelihood-ratio threshold.
  • Stochastic models fit to raw incidence data: feasible and transparent: Stochastic-model diagnostics compared observed incidence and lag-1 autocorrelation with distributions from simulations of the fitted model.The observed ACF(1) for Guinea, Liberia, and West Africa lay in the extreme right tail, while simulations showed greater high-frequency variability.
  • Stochastic models fit to raw incidence data: feasible and transparent: Localized transmission hotspots suggested spatial heterogeneity inconsistent with the well-mixed SEIR mass-action assumption.The paper links this heterogeneity to possible movement restrictions and spatial coordination of control measures.

Discussion

The paper argues that fitting deterministic models to cumulative incidence can produce biased estimates, underestimated uncertainty, and overconfident forecasts. It recommends stochastic models fitted to raw, disaggregated data, with explicit checks for assumption violations and model misspecification.

  • Deterministic models fit to cumulative incidence can bias parameters and substantially underestimate their uncertainty.
  • Models should use raw, disaggregated data rather than temporally accumulated data whenever possible.
  • When assumptions such as independent errors are violated, their effects should be checked explicitly.
  • Early Ebola incidence displayed erratic spatial and temporal dynamics contrary to a well-mixed model, while forecast envelopes depended on model and fitting data.
  • Forecast uncertainty is overconfident when deterministic models ignore environmental and demographic stochasticity.
  • Stochastic models improve accounting for real variability and provide greater opportunity to quantify uncertainty and diagnose misspecification.

Methods

The analysis uses weekly Ebola reports from Guinea, Liberia, and Sierra Leone, alongside a regional epidemic curve, and fits deterministic and stochastic SEIR variants. The formulation represents incubation with stages and defines transmission through the force of infection λ(t).

  • Weekly case reports from Guinea, Liberia, and Sierra Leone were digitized from a WHO situation report and also aggregated into a regional epidemic curve.
  • The transmission models were variants of the basic SEIR model, using stages to represent a more realistic Erlang incubation-period distribution.
  • The deterministic model uses R0, α, m, γ, and N to represent reproduction number, incubation, distribution shape, infectious period, and population size.
  • The force of infection is λ(t) = R0 γ I(t)/N, interpreted as the per-susceptible rate of infection.
  • The stochastic variant was implemented as a continuous-time Markov process approximated with a multinomial τ-leap algorithm using ∆t = 10^-2 wk.
  • Methods for the simulation study and empirical inference are described in the Supplementary Materials.

Appendix A. Simulation study

The simulation study compares deterministic-model fitting to raw versus cumulative incidence data under three levels of observation overdispersion. It evaluates estimation accuracy and confidence-interval performance across 1,500 simulated 39-week datasets.

  • Simulation design: 1,500 simulated 39-week datasets were generated from the stochastic model for k ∈ {0, 0.2, 0.5}, with R0 = 1.4 and fixed incubation and infectious periods.The deterministic model was fit to both raw and cumulative incidence data.
  • Estimation: The study estimated R0, reporting probability ρ, and negative binomial overdispersion k while fixing all other parameters at their generating values.Trajectory matching in pomp produced likelihood profiles, maximum-likelihood estimates, and likelihood-ratio confidence intervals for R0.
  • Uncertainty assessment: Least-squares fitting to cumulative data produced confidence intervals that were too narrow and achieved coverage far below the nominal value.This second simulation study examined the common deterministic fitting procedure and compared raw with accumulated simulated incidence.
  • Simulation results: Figure A2 compares estimates of R0 and ρ, nominal 99% confidence-interval widths for R0, and achieved coverage against the 99% target.Raw and accumulated incidence results are shown separately, with generating values marked by dashed lines.

Trajectory matching

The trajectory-matching workflow estimates model parameters by profiling likelihoods while addressing reporting-rate non-identifiability. It uses structured initialization and substantial computational effort before subsequent stochastic estimation.

  • Parameter identifiability: A flat likelihood profile for ρ indicated non-identifiability caused by a trade-off between reporting probability and initial conditions, so ρ was fixed at 0.2.The authors state that this assumption did not affect fit quality because the profile was flat.
  • Trajectory matching: Likelihood profiles over R0 and k were computed by initializing other parameters and initial conditions at 40 Sobol’ design points.The trajectory-matching calculations required approximately 21 CPU hours.
  • Model specification: Table B1 lists model parameters, their interpretations, assumed values, and the evidence sources underlying those assumptions.The table distinguishes parameters estimated from incidence data from parameters fixed by assumption.
  • Stochastic estimation: The workflow used trajectory-matching estimates to initialize 60-iteration IF2 runs with 2×10^3 particles for each country and data type.IF2 was implemented as mif in the R package pomp, with hyperbolic cooling and parameter random walks.

Results

The results identify cumulative-incidence fitting as fundamentally problematic because accumulated measurement errors are dependent. They compare model-data combinations and use stochastic simulations and summary-statistic probes to assess fit and forecast uncertainty.

  • Model-data comparison: The study fit deterministic and stochastic models to both raw incidence and accumulated incidence to separate stochasticity effects from cumulative-data fitting effects.The accumulated-data stochastic fit is described as incompatible with the monotonicity of the data and model simulations.
  • Cumulative-data problem: Cumulative-incidence errors are not independent because each accumulated observation contains measurement errors from multiple reporting intervals.This dependence is the fundamental problem with fitting cumulative data for either deterministic or stochastic transmission models.
  • Model checking: Figure B1 compares data and model-simulation probes, including autocorrelation, log-scale standard deviation, percentiles, detrended autocorrelation, and exponential growth rate.The probes are shown as simulated probability densities alongside their values on the data.
  • Parameter estimates: Table B2 reports maximum-likelihood estimates with nominal 95% confidence intervals for stochastic and deterministic models fitted to raw data.These estimates complement the likelihood-profile and probe-based comparisons.
  • Likelihood comparison: Figure B2 presents R0 likelihood profiles for four SEIR model-data combinations fitted to the Sierra Leone outbreak.The comparison varies both model type and whether raw or cumulative incidence data were used.
  • Forecast uncertainty: Figure B3 shows forecast medians and 95% simulation envelopes for models fit to raw and cumulative Sierra Leone incidence data.Observed fitting data are marked with black triangles.

Appendix C. Data and codes

The appendix documents the data, software, package versions, and cluster configuration used to reproduce the paper’s analyses. It identifies Dryad as the repository for the referenced files.

  • Reproducibility materials: All data and codes needed to reproduce the results are provided through Dryad at DOI:10.5061/dryad.r5f30.The appendix describes the associated files and reproducibility workflow.
  • Software dependencies: The appendix installs required R packages when they are absent, including pomp, ggplot2, foreach, doMPI, and iterators.The package list and installation commands are given in the reproducibility code.
  • Software versions: The study used R version 3.1.3 and records the versions of the analysis packages used to perform the study.The appendix lists versions for packages such as pomp, ggplot2, foreach, doMPI, and iterators.
  • Computing environment: Most appendix code was designed for execution on an OpenMPI cluster specified through an MPI hostfile.The example configuration includes localhost and four compute nodes, each with at least 10 cores.

Model implementation

The implementation constructs an Ebola pomp model with configurable country, data type, and least-squares options, combining deterministic and stochastic model components. It specifies SEIR dynamics, measurement models, parameter transformations, and simulation-study workflows for trajectory matching.

  • Model construction: The ebolaModel function builds a pomp object containing Ebola data, model code, parameters, state names, and simulation settings.It supports country selection, raw or cumulative data, timestep and exposed-stage configuration, missing-value handling, and least-squares measurement models.
  • Measurement model: The measurement model uses negative-binomial observations with mean rho*H_t and variance rho*H_t*(1+k*rho*H_t), while an ordinary-least-squares alternative uses normal errors.The model can generate cases and deaths through reporting probability, case-fatality ratio, and overdispersion parameters.
  • Process model: The process model represents susceptible, exposed, infectious, and recovered compartments with transitions driven by transmission, incubation, and recovery rates.The model tracks new E-to-I and I-to-R transitions through N_EI and N_IR.
  • Parameterization: The default assumptions set a Gamma-distributed incubation period with mean 11.4 days, a 7-day infectious period, rho=0.2, and case-fatality ratio 0.7.The corresponding discrete-time rates are calculated from the timestep and the number of exposed stages.
  • Simulation study: The simulation study varies raw versus cumulative data and overdispersion settings, then estimates R0, k, and rho using trajectory matching and likelihood profiles.The workflow records parameter estimates, likelihoods, convergence, and profile results across simulated datasets.

Diagnostics

The diagnostics workflow evaluates fitted models and forecasts using simulated probes, likelihood-based parameter ranges, effective sample sizes, and calibration-versus-projection comparisons. It runs separate deterministic and stochastic forecast pipelines and summarizes weighted predictive quantiles.

  • Additional diagnostics: Additional diagnostics simulate 2,000 probe datasets and compare observed and simulated exponential growth trends after log-scale detrending.The custom trend probe fits log1p(cases) against week, while residual probes remove the fitted exponential trend.
  • Forecast summaries: Forecast outputs are summarized with lower, median, and upper weighted quantiles over calibration and projection periods.The weighted quantile function uses probabilities 0.025, 0.5, and 0.975.
  • Parameter diagnostics: Likelihood-based parameter ranges retain fits within six log-likelihood units of the maximum and summarize minimum and maximum values by country, data type, model, and parameter.Maximum-likelihood estimates are rounded and organized alongside these ranges.
  • Forecasting: Forecast diagnostics generate weighted case trajectories for calibration and projection periods from deterministic and stochastic forecast models.Deterministic forecasts use model trajectories and simulated observations, whereas stochastic forecasts use particle-filtered states before simulation.
  • Forecast diagnostics: Effective sample size is computed from normalized likelihood weights at the final observed time for both deterministic and stochastic forecast simulations.The diagnostics report separate ESS values using 1/sum(ess^2).
Loading 1412.0968v5…