Source-linked AI summary

Estimation in emerging epidemics: biases and remedies

Tom Britton, Gianpaolo Scalia Tomba

arXiv:1803.01688v1stat.AP

TL;DR

Emerging outbreaks require rapid estimation of transmission and disease-course parameters from limited data, but common analyses can be biased by contact tracing, interval substitution, multiple infectors, and truncation. The paper uses statistical modelling, theoretical analysis, numerical examples, and Ebola-inspired simulations to identify and reduce these biases, finding that R0 may be underestimated by up to 20%.

  • Problem

    Early outbreak data are limited, yet estimates of transmission, disease-course, and fatality parameters are needed for prediction, countermeasure planning, and monitoring; standard analyses can introduce bias.

  • Method

    The paper combines stochastic epidemic modelling, theoretical analysis, heuristics, numerical examples, and simulations to study and correct biases in emerging-outbreak inference.

  • Results

    R0 could be underestimated by as much as 20% in the Ebola-inspired parameter setting, with backward generation-time observation and multiple possible infectors identified as major error sources.

  • Takeaways & Limitations

    Proper statistical modelling is needed to avoid biases that can propagate from generation-time and fatality estimation to R0 and growth-rate estimates.

  • Takeaways & Limitations

    The multiple-exposures problem remains difficult because assigning earliest, latest, or all potential infectors produces biased estimates under the discussed approaches.

Abstract

from arXiv · show

When analysing new emerging infectious disease outbreaks one typically has observational data over a limited period of time and several parameters to estimate, such as growth rate, R0, serial or generation interval distribution, latent and incubation times or case fatality rates. Also parameters describing the temporal relations between appearance of symptoms, notification, death and recovery/discharge will be of interest. These parameters form the basis for predicting the future outbreak, planning preventive measures and monitoring the progress of the disease. We study the problem of making inference during the emerging phase of an outbreak and point out potential sources of bias related to contact tracing, replacing generation times by serial intervals, multiple potential infectors or truncation effects amplified by exponential growth. These biases directly affect the estimation of e.g. the generation time distribution and the case fatality rate, but can then propagate to other estimates, e.g. of R0 and growth rate. Many of the traditionally used estimation methods in disease epidemiology may suffer from these biases when applied to the emerging disease outbreak situation. We show how to avoid these biases based on proper statistical modelling. We illustrate the theory by numerical examples and simulations based on the recent 2014-15 Ebola outbreak to quantify possible estimation biases, which may be up to 20% underestimation of R0, if the epidemic growth rate is fitted to observed data or, conversely, up to 62% overestimation of the growth rate if the correct R0 is used in conjunction with the Euler-Lotka equation.

Significance statement

Early outbreak analyses must estimate key disease and transmission parameters from limited, incomplete data to guide countermeasures and monitor epidemics. The paper identifies statistical biases in these estimates and develops modelling-based ways to reduce them.

  • Motivation: Early outbreak estimation uses limited case, disease-course, and contact-tracing data while key parameters guide intervention planning and progress monitoring.Available data commonly include reported cases, case histories, and possible infector-infectee durations, with little information on uninfected exposure.
  • Bias sources: The study examines biases from backward generation-time observation, replacing generation times with serial intervals, and multiple potential infectors.These problems arise directly from how early outbreak observations and contact tracing are typically collected.
  • Results: The combined biases can underestimate R0 by about 10% to 20% in the Ebola-inspired illustration, where the true R0 was 1.7 and an estimate could be 1.36.The paper notes that the magnitude may be larger in other settings and that such differences affect control planning.
  • Approach: The paper combines stochastic epidemic modelling, theoretical analysis, heuristics, numerical examples, and simulations to identify and reduce estimation biases.Its analyses use assumptions and parameter settings inspired by the 2014 West Africa Ebola outbreak.
  • Underlying model: The model represents incidence through a renewal equation in which infections generated by individuals infected s time units earlier contribute at rate R0fG(s).The incidence is assumed to approach exponential growth, and simulations illustrate this feature on a log scale.
  • Underlying model: The generation-time distribution links the basic reproduction number R0 and exponential growth rate r, making its estimation central to the model.The average infection rate also determines both R0 and r in the underlying epidemic framework.

3 Looking backwards rather than forwards in time

Contact tracing commonly observes generation times backward from infectees rather than forward from infectors, producing a biased distribution during exponential epidemic growth. Using that backward distribution can distort estimates of both growth and reproduction.

  • Observed generation times: Backward contact tracing observes the interval from an infectee to an identified infector, rather than the forward infection interval generated by an infector.The paper defines the generation-time distribution through infection-to-infection times and shows how tracing direction changes the observed distribution.
  • Bias mechanism: The backward generation-time density is fB(t) = e^-rtR0fG(t), and its mean is smaller than that of the forward distribution fG(t).The density integrates to 1 under the Euler-Lotka relation.
  • Consequences: Using fB(t) with the correct R0 always produces a growth-rate estimate rB larger than the true r.In a Markovian SIR example, rB = R0r; for R0 between 1.5 and 2, growth is overestimated by 50% to 100%.
  • Consequences: Using fB(t) when r is known causes compensatory underestimation of R0 because the backward distribution gives excessive weight to recent, more numerous infections.This follows from incidence being R0 times a weighted sum of previous incidence.

4 Replacing Generation times with Serial intervals

Serial intervals are often used as proxies for generation times because infection times are rarely observed, but their distributions can differ and bias estimates of outbreak parameters. The section also addresses multiple potential infectors and shows that likelihood-based modelling is needed to avoid biased incubation-period inference.

  • Replacing generation times with serial intervals: Serial intervals measure symptom-onset differences, whereas generation times measure infection-time differences, making the former a practical surrogate when infection times are unavailable.The relation between infectiousness and symptom onset is disease dependent.
  • Replacing generation times with serial intervals: Under independent, identically distributed individual periods, generation time and serial interval have equal expected values but generally different variances.The serial interval variance lacks the covariance term present in the generation-time variance, and existing models examined typically give S equal or larger variance than G.
  • Replacing generation times with serial intervals: Using serial intervals instead of generation times typically underestimates R0 when r is given and overestimates r when R0 is given.These effects are model dependent but follow from the typically larger variance of serial intervals.
  • Multiple exposures: Multiple potential exposures create an uncertain infection origin, so treating the earliest, latest, or all contacts as infectors produces biased incubation estimates.Earliest-contact assignment overestimates incubation, latest-contact assignment underestimates it, and treating all contacts as independent observations also underestimates incubation while overstating precision.
  • Multiple exposures: Correct analysis of multiple exposures conditions on exposure number and times and uses the likelihood contribution containing the infection probability and incubation distribution.A parametric form for the incubation distribution, such as a two-parameter gamma distribution, can then be used for inference.

6 Case fatality rate

Early case-fatality estimates are biased because deaths occur after notification and the observation window truncates outcomes during exponential case growth. Delayed-event modelling can estimate or correct this bias, while recovery-based estimators depend on the relative timing of death and remission.

  • Case fatality rate: The naive estimator Dobs/K is biased downward because some deaths among observed cases have not occurred by the observation time.The bias is amplified by the finite time horizon and exponential increase of cases.
  • Case fatality rate: Knowledge of the notification-to-death delay distribution and exponential growth rate can be used to correct the naive case-fatality estimate.For an exponential delay distribution with mean µ, the correction uses π(∞) = 1/(1+rµ).
  • Case fatality rate: The estimator Dobs/(Dobs + Robs) is approximately unbiased only when the eventual observed death and recovery fractions are equal.If recovery takes stochastically longer than death, the case-fatality rate is overestimated; the reverse ordering can produce the opposite relationship.

7 Results

Using Ebola-inspired parameter values and simulations, the paper quantifies biases from backward generation times, serial intervals, restricted contact tracing, and delayed outcome observations. These biases can substantially distort estimates of growth, R0, case fatality, and epidemic timing.

  • 7.1 Looking backwards: 19% overestimation of the growth rate resulted from using backward rather than forward generation times, while R0 was underestimated by 8%.The backward distribution had mean 12.6 days and standard deviation 7.26, compared with 15 days and 8.66 for the true generation-time distribution.
  • 7.2 Serial intervals: Under favorable assumptions, replacing generation times with serial intervals changed R0 by 0.2% and the growth rate by 0.5%.The assumed incubation-to-latent-period variation corresponded to c = 1.026, producing only a small increase in serial-interval variability.
  • 7.2 Serial intervals: The serial-interval result depended on a specific favorable model, and larger latent-incubation differences or different event ordering could produce larger discrepancies.The authors selected the model to avoid statistical complications, so the small estimated effect is not general across natural-history assumptions.
  • 7.3 Multiple exposures: Restricting observations to infectees with one potential infector shortened the estimated incubation period by 3.3 days and underestimated the mean generation time as 12 rather than 15.3 days.With R0 = 1.7, the resulting growth-rate estimate was 0.0522 versus the true 0.0383, a 36% overestimate; using the true growth rate yielded R0single = 1.50, a 12% underestimate.
  • 7.5 Case fatality rate: With a 20-day doubling time, the simple case-fatality estimator underestimated CFR by about 24%, while a known-final-destiny estimator overestimated a 70% CFR by approximately 5%.The first result used a notification-to-death exponential mean of 9 days; the second used ρ(∞) = 0.63.
  • 7.6 A simulation study: Reaching 4500 notified cases took approximately 200 days on average and showed a range of 123–358 days, with most variability occurring before the first 100 cases.The first stage averaged 102 days with standard deviation 28, whereas the second averaged 98 days with standard deviation 10.

8 Discussion

The paper identifies several early-outbreak inference biases and shows that they can compound, materially distorting R0 and preventive planning. It also outlines how modelling and simulation can diagnose or reduce these problems, while noting important unstudied data and model limitations.

  • Three bias sources—backward generation-time estimation, serial intervals, and multiple infectors—can all underestimate R0 when growth rate r is correctly estimated.The converse is growth-rate overestimation when a correct R0 is used, though this is likely less common in practice.
  • Up to 20% underestimation of R0 is illustrated using Ebola-inspired parameters, with a true R0 of 1.7 potentially estimated as 1.36.The third effect is largest, the serial-interval effect is smallest, and their combined direction can make the bias substantial.
  • The corresponding critical immunization level falls from 41% for the true R0 to 26% under the biased estimate.The paper states that this underestimation may suggest preventive measures insufficient to stop spread.
  • Correctly modelling the sampling situation can help explain and correct many biases, while simulation can assess procedures when simple analytical results are difficult.The authors also report that accurate inference may be possible for the multiple-infector problem, although more research is needed.
  • The paper does not treat underreporting, reporting delays, batch reporting, changing behaviour, or spatial effects on parameter estimates.It assumes no individual- or society-induced behaviour change during data collection and ignores social or spatial spreading effects.
  • The results are intended to improve future analyses and support health-authority efforts to predict outbreaks and identify preventive measures.

Supplementary information

The supplementary material specifies an Ebola-like stochastic epidemic simulation and derives benchmark growth quantities from a Gamma-distributed SEIR model. Its renewal-equation solution approaches exponential growth across infected, reported, recovered, and dead populations.

  • Model structure: Each infected individual follows a stochastic SEIR model with Gamma-distributed latent and infectious periods calibrated to the recent Ebola outbreak.The latent period is Γ(2, 1/5), with mean 10 and standard deviation 7.1, while the infectious period is Γ(1, 1/5), with mean and standard deviation 5.
  • Duration of simulations: Only exploding simulations reaching 4500 reported cases are retained, after which statistics are collected and trajectories continue for six weeks.The continuation supports prediction of the final level six weeks after the first 4500 cases.
  • Theoretical results: The generation-time distribution is Γ(3, 1/5), with mean 15 and standard deviation 8.7, while the Malthusian parameter is r = 0.0387.The deterministic doubling time is 17.9 days.
  • Theoretical results: The expected incidence satisfies a renewal equation driven by R0 and the generation-time distribution fG.The stochastic epidemic is represented as a Crump-Mode-Jagers branching process.
  • Theoretical results: The renewal-equation solution quickly approaches exponential growth, which also describes total infected, reported, recovered, and dead populations with different constants.

2 Delayed observations

The delayed-observation analysis expresses the observed-event fraction through the delay distribution and shows that exponential growth makes this fraction depend on the delay transform rather than simply approaching one.

  • For exponentially growing event intensity λ(s) ∼ e^rs, the observed-event fraction approaches E(e^−rD), where D follows the observation-delay distribution.For constant or polynomially growing intensity, the fraction instead approaches 1 as T grows large.
  • For D distributed as Γ(α, λ), the limiting fraction is π(∞) = (1/(1+rE(D)/α))^α.
  • At fixed growth rate r and mean delay E(D), π(∞) decreases as the Gamma shape parameter α increases.The exponential case α = 1 therefore gives the largest value among Gamma distributions with the same mean.

3 "Backward" observation of generation times

Backward observation of generation times reweights the true generation-time distribution during exponential growth. Solving Euler–Lotka with this backward distribution therefore yields a larger growth-rate estimate, with an explicit Gamma-distribution relation.

  • During exponential growth, backward-observed generation times have density fB(t) = e^−rGtR0fG(t), reweighting the true distribution fG.
  • Using fB instead of fG in the Euler–Lotka equation produces a growth-rate estimate rB at least as large as the true rG.The inequality follows from the monotonicity and covariance properties described in the analysis.
  • For a Gamma generation-time distribution, the backward-based estimate satisfies rB = R0^(1/α) rG.
  • The derivation compares expectations under backward and forward generation-time distributions, whose relevant functions have different monotonicity.

4 The Euler-Lotka equation and G ∼Γ(α, λ)

The Euler–Lotka framework links the generation-time distribution with exponential growth and R0, while distributional misspecification can systematically bias these quantities.

  • For a Gamma(α, λ) distribution, r = λ(R0^(1/α) − 1) and R0 = (1 + r/λ)^α.
  • The coefficient of variation is 1/√α, with α = μ^2/σ^2 and λ = μ/σ^2 from the mean and variance.
  • The Euler–Lotka equation can calculate exponential growth from a Gamma generation-time distribution when R0 is known, or R0 when growth rate is known.
  • Replacing the true generation-time distribution with the backward distribution changes the Gamma rate parameter before solving the Euler–Lotka equation.
  • Using a Gamma distribution with larger variance but the same mean overestimates exponential growth and underestimates R0 when the other parameter is held at its true value.

5 Modelling multiple exposures and estimating the incubation period

The paper models repeated exposures and uses likelihood or moment equations to estimate infection probability and incubation characteristics, finding that moment estimation can be more stable under misspecification.

  • Maximum likelihood and moment equations can estimate p, contact rate μ, and the mean and variance of incubation time from observed contact and symptom timing data.
  • The contact process assumes Poisson contacts, independent infection probability p per contact, and an incubation time T with estimable mean and variance.
  • The infecting contact index is Geom(p), while contacts after infection and before symptoms are Poisson conditional on the incubation time.
  • 1000 replicates used 500 individuals, p = 1/2, incubation mean 11.4 days, standard deviation 8.1 days, and contact intensity 0.0725.
  • Maximum likelihood works well under the correct distribution but less well under misspecification, whereas moment estimation appears comparatively stable.

6 The simulation model for generation and serial intervals

The simulation distinguishes generation times from serial intervals and shows how backward observation changes the apparent generation-time distribution and variance.

  • Generation and serial intervals have the same mean, but their variances differ through a covariance involving symptom, latent, and infectious-period components.
  • The model gives G ∼ Γ(3, 1/5) with variance 75, while simulation estimates are approximately 52.4 for generation-time variance and 54.6 for serial-interval variance.
  • Backward observation predicts a change from Γ(3, 0.2) to Γ(3, 0.2387), yielding variance 52.7, consistent with the simulated generation-time estimate.
  • In the simulation, the serial interval’s coefficient of variation is 1.026 times the generation-time coefficient of variation.

7 Estimating the growth rate

The paper compares several estimators of exponential growth and tests their use for forecasting epidemic size, finding broadly similar performance with slight overestimation and a small advantage for logarithmic cumulative regression.

  • The study compares logarithmic regression on cumulative cases, logarithmic regression on daily cases, daily cumulative ratios, a branching-process estimator, and a renewal-equation approach.
  • The simulations use 1000 epidemic trajectories to estimate growth before 4500 notified cases and predict epidemic size six weeks later.
  • The true growth rate is 0.03870, with exp(r) = 1.03946 and six-week growth factor exp(42r) = 5.07973.
  • All four direct growth-rate estimators produce reasonably concentrated values around the true rate 0.0387.
  • All estimators slightly overestimate growth by less than 1% on average, and estimator (a) has mean six-week prediction factor 5.14729 with a [−16%, +21%] prediction interval.
  • Using the estimated growth factor with the last cumulative datum leaves about 1% overestimation, while estimator (a)'s interval narrows to [−13%, +18%].
Loading 1803.01688v1…