Source-linked AI summary
Estimating the COVID-19 Infection Rate: Anatomy of an Inference Problem
Charles F. Manski, Francesca Molinari
TL;DR
The paper addresses how to bound COVID-19 infection rates when testing data are incomplete and test accuracy is imperfect. It combines partial-identification analysis with testing data and assumptions about untested persons and test accuracy, finding that credible monotonicity assumptions restrict the infection-rate bounds to about width 0.5 in the COVID-19 context. Narrower bounds require stronger assumptions with considerable identifying power.
Problem
Missing evidence about infection among untested persons and uncertainty about test accuracy impede credible, informative bounds on COVID-19 infection and severe-illness rates.
Method
The paper applies partial-identification methods, combining testing data with assumptions that bound infection among untested persons and account for test accuracy.
Results
Credible monotonicity assumptions restrict the population infection-rate bounds to about width 0.5 in the current COVID-19 context.
Takeaways & Limitations
The findings quantify how uncertainty about untested populations and test negative predictive value determines what can be learned about infection and severe-illness rates.
Takeaways & Limitations
Narrower bounds are logically possible only by imposing stronger assumptions with considerable identifying power.
Abstract
from arXiv · showhide
As a consequence of missing data on tests for infection and imperfect accuracy of tests, reported rates of population infection by the SARS CoV-2 virus are lower than actual rates of infection. Hence, reported rates of severe illness conditional on infection are higher than actual rates. Understanding the time path of the COVID-19 pandemic has been hampered by the absence of bounds on infection rates that are credible and informative. This paper explains the logical problem of bounding these rates and reports illustrative findings, using data from Illinois, New York, and Italy. We combine the data with assumptions on the infection rate in the untested population and on the accuracy of the tests that appear credible in the current context. We find that the infection rate might be substantially higher than reported. We also find that the infection fatality rate in Italy is substantially lower than reported.
2. Methods
The paper frames COVID-19 infection-rate estimation as a partial-identification problem caused by missing infection data, untested populations, and imperfect test accuracy. It derives bounds by combining observed testing data with assumptions about untested persons and test accuracy, then extends the bounds to severe illness and asymptomatic infections.
- Scope: The analysis bounds the population infection rate and then uses those bounds to study severe illness conditional on infection and rates conditional on observed patient characteristics.The paper first treats the infection-rate problem abstractly, derives bounds under credible assumptions, and extends the analysis to related rates.
- Target quantity: P(Cd = 1) denotes the fraction of the population infected by date d, but this quantity is not directly observable.Population surveillance instead provides testing rates and positive-result rates among those tested.
- Identification strategy: The infection rate is decomposed by positive results, negative results among tested persons, and untested persons using the Law of Total Probability.Equations (1)–(4) express the target in terms of observed testing outcomes and infection probabilities for tested and untested groups.
- Test accuracy: Test accuracy enters through positive predictive value and one minus negative predictive value, while sensitivity and specificity require knowledge of infection prevalence among tested persons.The analysis assumes positive predictive value equals one but acknowledges substantial uncertainty about negative predictive value.
- Bound width: The width of the infection-rate bound combines uncertainty about infection among untested and test-negative persons, weighted by the fractions in those groups.With no information about infection among the untested, the bound width cannot be smaller than the untested fraction, even with precise test-accuracy information.
- Additional restrictions: With monotonicity assumptions, the bound depends on the upper bound for infection among untested persons and modestly improves lower bounds in the data used.The illustrative results show that conclusions about infection prevalence depend on assumptions about P(Cd = 1|Td = 0).
3. Data
The empirical analysis combines daily testing and positive-result data from Illinois, New York, and Italy with population counts from official sources. Italy additionally provides cumulative severe-outcome data, while comparable Illinois and New York severe-outcome time series were unavailable.
- Study populations: The study analyzes Illinois, New York, and Italy using population counts from national or state statistical sources.Illinois and New York population counts come from the U.S. Census Bureau; Italy’s comes from Istituto Nazionale di Statistica.
- Testing data: Daily data report cumulative numbers tested and positive test results, beginning February 24 for Italy, March 10 for Illinois, and March 2 for New York.The analysis begins on March 16, after all three locations had at least 100 confirmed cases.
- Severe outcomes: Italy’s data include cumulative severe outcomes—hospitalization, intensive-care-unit admission, and death—but official Illinois and New York severe-outcome time series were unavailable.This limits comparable severe-outcome analysis across all three locations.
- Estimation: The testing-rate quantities used in the bounds are estimated as simple frequencies, such as tested individuals divided by population size.The data sources are the Illinois Department of Public Health, New York State Department of Health, and Italian Protezione Civile.
4. Results
Using data from Illinois, New York, and Italy, the paper reports infection-rate bounds under monotonicity and test-related assumptions. The bounds remain wide because few people were tested, but they contain information and imply substantially lower Italian fatality than among confirmed infections.
- Observed testing and outcomes: 0.005, 0.017, and 0.012 were the April 6 tested fractions in Illinois, New York, and Italy, respectively.Testing increased from 0 in Illinois, 0.001 in New York, and 0.002 in Italy on March 16.
- Infection-rate bounds: 0.455 to 0.516, 0.48 to 0.637, and about 0.51 were the bound widths for Illinois, New York, and Italy, respectively.The substantial widths reflect the very small fraction of tested individuals.
- Infection-rate bounds: The bounds retain informational content despite being much wider than bounds available with credible information about untested persons.The paper contrasts these results with extremely wide bounds obtained without such information.
- Infection-rate bounds: 0.001 to 0.517, 0.008 to 0.645, and 0.003 to 0.510 were the April 6 infection-rate bounds for Illinois, New York, and Italy.These bounds incorporate the paper’s monotonicity assumptions.
- Additional assumptions: Expert opinion slightly raised the lower infection-rate bounds by a factor of (0.75)^-1 without lowering the upper bounds.For April 6, the updated bounds were [0.002, 0.517], [0.011, 0.645], and [0.004, 0.510].
- Severe outcomes: On April 6, Italy’s fatality probability conditional on infection was bounded by [0.001, 0.086], below the confirmed-infected fatality rate of 0.125.The hospitalization bound was [0.001, 0.172], while the intensive-care bound was [0, 0.02].
5. Discussion
The paper applies partial-identification methods to bound COVID-19 infection and severe-illness rates under credible assumptions, while emphasizing that current data leave substantial uncertainty. It concludes that stronger bounds require stronger assumptions or better data, especially random testing and improved knowledge of test accuracy.
- The analysis uses partial-identification methods to study COVID-19 infection and severe-illness rates under available data and maintained assumptions.The authors treat the data and assumptions as determining the inferences that can logically be drawn.
- The authors illustrate the analysis using data from two American states and Italy, with monotonicity assumptions and a conjectured bound on asymptomatic infection.The latter assumption is presented as an example of a less firmly grounded assumption that may be used if considered credible.
- With very low testing coverage, little can be concluded about the population infection rate without assumptions about infection among untested people.The paper identifies uncertainty about the untested population as the dominant concern.
- The monotonicity assumptions restrict the population infection rate to bounds with about 0.5 width in the current COVID context.Narrower bounds would require additional assumptions with considerable identifying power.
- The authors do not report narrower bounds because they do not see a credible basis for adding assumptions that would justify them.Readers who can motivate stronger assumptions may adapt the analysis to study their implications.
- Data from other locations help tighten inference only when combined with assumptions that support credible extrapolation across locations.The paper points to intersection bounds as a formal way to combine such information.
- The analysis abstracts from several possible uncertainties, including immunity after recovery, retesting, diagnostic accuracy, and death-cause coding.The authors state that these assumptions may be inaccurate and that the analysis can be extended to incorporate further uncertainty.
- The empirical results are presented as logical inferences rather than statistical estimates because the data are exact population counts and no reasonable sampling process is specified.Accordingly, the paper does not provide measures of statistical precision.
6. Conclusion
The paper argues that better data are needed to improve knowledge of COVID-19 infection and severity rates. It particularly emphasizes random population testing and better understanding of the negative predictive value of the tests in use.
- Better data are the more satisfactory way to increase knowledge of the COVID-19 infection rate.The authors caution that narrowing bounds through stronger assumptions is less satisfactory when those assumptions lack credible support.
- Random testing of populations would contribute enormously to understanding the infection rate.
- Improving knowledge of the negative predictive value of tests is also important for interpreting infection and severity rates.