Source-linked AI summary

A Modified SIR Model for the COVID-19 Contagion in Italy

Giuseppe C. Calafiore, Carlo Novara, Corrado Possieri

arXiv:2003.14391v1physics.soc-phcs.CEcs.SIeess.SY

TL;DR

The paper addresses COVID-19 modeling when detected positives do not equal actual infections and key initial conditions are unknown. It develops a modified SIR framework, estimates its parameters using grid search and weighted least squares, and applies it to Italian data through March 30, 2020. The estimated detection-to-actual-infection factor is about 63, implying a lower mortality rate when calculated over actual infections.

  • Problem

    COVID-19 models must account for underdetected infections and unknown initial susceptible populations rather than equating detected positives with all infections.

  • Method

    The paper combines a modified SIR model with grid search over nonlinear parameters, weighted least-squares estimation of remaining parameters, and weighted multi-step prediction.

  • Results

    The estimated proportionality factor is about 63; mortality is 1.18% using detected positives versus 0.019% using estimated actual infections.

  • Takeaways & Limitations

    Accounting for undetected infections materially changes estimated mortality and supports prediction of detected infections, recoveries, and deaths.

Abstract

from arXiv · show

The purpose of this work is to give a contribution to the understanding of the COVID-19 contagion in Italy. To this end, we developed a modified Susceptible-Infected-Recovered (SIR) model for the contagion, and we used official data of the pandemic up to March 30th, 2020 for identifying the parameters of this model. The non standard part of our approach resides in the fact that we considered as model parameters also the initial number of susceptible individuals, as well as the proportionality factor relating the detected number of positives with the actual (and unknown) number of infected individuals. Identifying the contagion, recovery and death rates as well as the mentioned parameters amounts to a non-convex identification problem that we solved by means of a two-dimensional grid search in the outer loop, with a standard weighted least-squares optimization problem as the inner step.

I. INTRODUCTION

The paper develops a modified collective SIR framework for COVID-19 in Italy that addresses unobserved infections and difficult parameter identification. It combines grid search, convex optimization, and weighted multi-step prediction for practical epidemic modeling.

  • Motivation: Mathematical models support epidemic prediction, transmission-parameter estimation, scenario simulation, and evaluation of control strategies.Their potential applications include containment measures, lockdowns, and vaccination campaigns.
  • Modeling context: Collective models use few parameters and collective variables, making them simpler than networked models but less detailed.The paper focuses on collective models because they can remain useful under data scarcity and for nonexpert operators.
  • Identification challenge: Observed positive cases may substantially undercount actual infections, while epidemic-model identification often involves non-convex optimization.Treating detected cases as the true infected population can produce misleading epidemiological interpretations.
  • Contribution: The proposed variant models actual infections and identifies nonlinear parameters through a grid search with convex optimization for the remaining parameters.The approach is designed for collective models with relatively few nonlinear parameters.
  • Prediction method: Weighted averaging of multi-step predictions from all available initial conditions is introduced to reduce noise and improve long-term prediction accuracy.The method addresses the limitation that one-step-error minimization does not directly minimize multi-step prediction errors.
  • Case study: A real-data case study applies the framework to the COVID-19 epidemic in Italy.

II. SIRD MODEL FOR COVID-19 CONTAGION

The paper formulates a discrete-time SIRD model for an isolated region, distinguishing susceptible, active infected, recovered, and deceased populations. It modifies the model to account for undetected infections and estimates initial susceptible population alongside epidemiological rates.

  • State variables: The model tracks susceptible, active infected, recovered, and deceased individuals through discrete-time epidemic equations.Recovered and deceased populations accumulate through recovery and mortality processes.
  • Model structure: The SIRD formulation uses infection, recovery, and mortality rates with positive initial conditions for the population compartments.Time is represented in days, and the model neglects deaths unrelated to the disease under study.
  • Assumptions: The model assumes recovered individuals are no longer susceptible and treats each geographical region as isolated from others.Isolation may be reasonable when containment measures are enforced.
  • Observed infections: Because observations detect only a fraction of infections, the model distinguishes actual infected individuals from detected infected individuals.Undetected cases may include asymptomatic individuals.
  • Modified formulation: The modified equations incorporate detected recovered and susceptible quantities through the proportionality factor and associated weighted variables.The resulting equations are referred to as the SIRD model.
  • Parameter estimation: Model behavior depends strongly on initial conditions, which affect both the amplitude and timing of the infection peak.The paper therefore estimates S(t0), α, β, γ, and ν from available data.

III. MODEL IDENTIFICATION

The model parameters are identified from official Italian COVID-19 data using a weighted least-squares regression nested inside a grid search over nonlinear parameters, including the susceptible-population fraction.

  • Data and setup: Official Italian data from February 24 to March 30, 2020 provide the time series used for parameter identification.The data comprise detected infected, recovered, and deceased individuals, with days represented on a discrete time scale.
  • Model parameters: ω represents the fraction of the regional population initially susceptible, with S(t0) = ωP.This parameter addresses the unavailability of the initial susceptible population.
  • Inner optimization: For fixed α and ω, β, γ, and ν̃ are estimated by solving a convex weighted mean-square optimization problem.The weighting parameter ρ ∈ (0, 1) gives greater relevance to recent observations.
  • Outer search: The overall identification problem is non-convex because MSE(α, ω) need not be convex.The method grids α and ω, solves the inner optimization for each grid point, and selects the parameter set minimizing the objective.
  • Algorithm: Algorithm 1 returns ω⋆, α⋆, β⋆, γ⋆, and ν̃⋆ after evaluating the candidate grid.Its inputs include the observed time series, a maximum α, the weighting parameter ρ, and the total population P.

IV. MODEL PREDICTIONS

The prediction procedure uses the identified SIRD model to generate forward trajectories from available observations and combines these forecasts through a weighted average.

  • Prediction procedure: The prediction algorithm estimates future detected infected, recovered, and deceased individuals from the identified model and available data.It produces predictions of ˜I(t), ˜R(t), and D(t) over the prediction horizon.
  • Forward simulation: For every available observation time, the algorithm projects the model states forward across the prediction horizon.The projected states include susceptible, infected, recovered, and deceased individuals.
  • Forecast aggregation: For T > Θ, the returned prediction is a weighted average of forward simulations initialized at different available times.This combines forecasts generated from the available initial conditions rather than relying on a single starting point.
  • Results display: Figure 2 compares the weighted-average forecast with individual forward predictions and one-step projections using the parameters in Table I.The figure concerns Italy and uses the model parameters obtained for the prediction task.

V. DISCUSSION

For aggregated Italian data, the model estimates a large factor linking detected positives to actual infections and correspondingly lowers the mortality estimate; its March 30 calibration is limited by non-stationary lockdown effects and uncertain data.

  • Estimated infections: α = 63 is the estimated proportionality factor between detected positives and actual infected individuals in aggregated Italian data.Multiplying 101739 detected positives by α = 63 gives about 6.4 million infected individuals.
  • Mortality estimate: ν̃ = 1.18% based on detected positives decreases to ν = 0.019% when mortality is referenced to actual infected individuals.The decrease follows the paper’s distinction between detected positives and the estimated total infected population.
  • Uncertainty: The parameter and prediction estimates are sensitive to input data, while no Monte Carlo analysis was performed to derive reliability intervals.The paper also expects substantial variability because of uncertainty in data-collection procedures.
  • Prediction scope: The model’s predictions are expected to be pessimistic because lockdown measures changed the contagion dynamics after the March 30 calibration period.The paper describes the underlying process as non-stationary and notes a roughly two-week delay before the lockdown effect appeared in the data.
Loading 2003.14391v1…