Source-linked AI summary

The geographic spread of COVID-19 correlates with the structure of social networks as measured by Facebook

Theresa Kuchler, Dominic Russel, Johannes Stroebel

arXiv:2004.03055v3physics.soc-phcs.SIq-bio.PE

TL;DR

The paper asks whether geographically structured social connections can help forecast communicable-disease spread. Using aggregated Facebook friendship data and regional COVID-19 outcomes, it finds that social connectedness predicts outbreaks beyond physical distance and demographics, with broad availability as a practical advantage.

  • Problem

    Forecasting outbreak risk requires information about cross-region physical interactions, but the geographic structure of social connections is difficult to measure at national or global scale.

  • Method

    The paper uses aggregated Facebook friendship data to measure social connectedness and tests its relationship with COVID-19 cases and predictive models across U.S. counties and Italian provinces.

  • Results

    Social connectedness predicts COVID-19 prevalence and subsequent case growth beyond physical proximity and demographics, while improving out-of-sample predictions across U.S. counties.

  • Takeaways & Limitations

    Stable, broadly available social-connectedness data can provide predictive power over other available measures for epidemiological work.

  • Takeaways & Limitations

    The methodology is not intended to be a state-of-the-art epidemiological model, and its additional predictive value is small when analysis is limited to counties with LEX and Google data.

Abstract

from arXiv · show

We use aggregated data from Facebook to show that COVID-19 is more likely to spread between regions with stronger social network connections. Areas with more social ties to two early COVID-19 "hotspots" (Westchester County, NY, in the U.S. and Lodi province in Italy) generally had more confirmed COVID-19 cases by the end of March. These relationships hold after controlling for geographic distance to the hotspots as well as the population density and demographics of the regions. As the pandemic progressed in the U.S., a county's social proximity to recent COVID-19 cases and deaths predicts future outbreaks over and above physical proximity and demographics. In part due to its broad coverage, social connectedness data provides additional predictive power to measures based on smartphone location or online search data. These results suggest that data from online social networks can be useful to epidemiologists and others hoping to forecast the spread of communicable diseases such as COVID-19.

1 Data Description

The paper measures regional social connectedness using aggregated Facebook friendship data, covering U.S. counties and Italian provinces. The Social Connectedness Index normalizes cross-region friendship links by the active Facebook user populations of both regions.

  • Data sources: Facebook’s March 2020 snapshot covers active users and friendship networks across broad geographic regions.Locations are assigned using Facebook profile, device, and connection information.
  • Data sources: Facebook friendships require mutual consent and are capped at 5,000 connections per person, making them more likely to represent real-world acquaintances than many platform links.
  • Measure: The Social Connectedness Index measures the probability that Facebook users in two regions are friends with one another.It is computed from cross-region friendship links divided by the product of the two regions’ active Facebook user counts.
  • Validation and related work: Prior research links the measure to trade, patent citations, investment decisions, travel patterns, and other social interactions.
  • Outcome data: COVID-19 county and province case data come from Johns Hopkins University and Italy’s Dipartimento della Protezione Civile.The study also uses COVID-19 deaths because regional testing differences may bias case-based results.

2 Early Hotspot Analysis

Across U.S. counties and Italian provinces, stronger social ties to early COVID-19 hotspots aligned with higher case prevalence. The relationships remained informative after accounting for geographic distance and regional characteristics, while hotspot-specific connections provided the strongest forecasting signal.

  • U.S. hotspot: Westchester-connected U.S. counties generally had higher COVID-19 case prevalence on March 30, 2020, especially in coastal regions and urban centers.Figure 1 compares social connectedness to Westchester with confirmed cases per 10,000 residents.
  • U.S. hotspot: 0.88 COVID-19 cases per 10,000 residents is the increase associated with doubling social connectedness to Westchester, with R-squared 0.093.The relationship excludes counties within 50 miles of Westchester.
  • Controls and robustness: The hotspot relationships remain after controlling for geographic distance, population density, income, and urban-rural classifications.Placebo regressions across other large U.S. counties contextualize the specificity of Westchester’s predictive relationship.
  • U.S. hotspot: Westchester’s incremental R-squared is 0.037, second only to New York City, with nine of the top ten counties in the New York-Newark CSA when excluding counties within 150 miles.The results indicate that connections to other demographically similar counties without early outbreaks were less useful for forecasting spread.
  • Interpretation and scope: The analysis is intended to demonstrate predictive power from online social-network data beyond variables already readily available, rather than establish a complete epidemiological model.The paper benchmarks this signal against smartphone location and Google symptom-search measures in later analysis.
  • Italy hotspot: Italian provinces with stronger ties to Lodi and higher case densities were concentrated in Lombardy and nearby Piemonte and Veneto, with Rimini also showing both patterns.Southern provinces had social ties to Lombardy through worker and student mobility.

3 Time Series Analysis

The analysis tests whether time-varying social exposure to COVID-19 cases predicts subsequent county case growth beyond physical proximity and demographic controls. Out-of-sample results show that social proximity improves predictions across U.S. counties, although its incremental value is smaller where smartphone and Google data are available.

  • Time-varying exposure measures: Social Proximity to Cases measures county exposure to COVID-19 through social connections, while Physical Proximity to Cases measures exposure through geographic distance.The analysis also compares these measures with smartphone-based LEX exposure and Google symptom-search changes.
  • Regression analysis: Lagged social proximity to new cases is positively associated with subsequent county case growth after controlling for physical proximity and regional characteristics.The specification uses two-week periods, two lags of own case growth, and additional controls in stricter models.
  • Regression analysis: The relationship remains similar when COVID-19 deaths replace cases, suggesting the case-growth result is not driven by differential testing across counties.The death analysis uses four-week periods with exposure measures lagged by four and eight weeks.
  • Comparison with other predictors: Social proximity remains a significant predictor alongside Google symptom searches and lagged LEX proximity to cases in the in-sample analysis.This comparison motivates the subsequent out-of-sample forecasting exercise.
  • Out-of-sample prediction: 0.652 log new cases per 10,000 residents is the final RMSE for the model including lagged social proximity, versus 0.667 without it.Including social proximity lowers RMSE in every prediction-period row, indicating improved out-of-sample fit.
  • Out-of-sample prediction: Social proximity adds only small predictive value among 1,976 counties with both LEX and Google data, but consistently improves predictions across the full U.S. county set.Its broader availability reflects Facebook’s large user base and the relative stability of social connectedness over time.

4 Conclusion

Forecasting outbreak risk is important for determining public health responses, and information about the geography of social connections is relevant because social connections shape physical interactions. The authors do not present their methods as a state-of-the-art epidemiological model, but suggest social connectedness may add predictive power because of its broad geographic coverage and fine granularity.

  • Forecasting outbreak risk helps regions determine optimal public health responses to communicable diseases.The risk is partly determined by physical interactions between residents of different areas, especially areas with severe outbreaks.
  • Geographic information about social connections is crucial because those connections shape patterns of physical interaction between regions.
  • The methodologies are not intended to constitute a state-of-the-art epidemiological model.
  • Social connectedness may provide predictive power beyond other available measures because it offers broad geographic coverage and high granularity.

A Out-of-Sample Hotspot Analysis

This hotspot exercise trains models on Italian provincial data and tests them on U.S. counties, finding that hotspot social connectedness improves predictions of later COVID-19 spread.

  • The exercise trains models on an Italian hotspot and tests them using a subsequent U.S. hotspot.
  • Italian models predict U.S. county cases using density, income, distance, and optionally social connectedness to the hotspot.The Italian training outcome is March 10 cases per 10,000 people, while U.S. predictions target March 30.
  • Adding hotspot social connectedness lowers linear-regression RMSE from 0.99 to 0.97 and random-forest RMSE from 1.04 to 1.01.The RMSEs are measured in standard deviations from mean cases per 10,000 residents.
  • Including hotspot social connectedness increases rank-rank correlations between predicted and true county rankings for both models.

B Additional Time Series Regressions

Time-series analyses show that prior social proximity to COVID-19 cases predicts subsequent county case growth, including after accounting for physical proximity and other predictive measures.

  • A doubling in social proximity to cases corresponds to a 10.9%–66.4% increase in actual cases in the next two-week period.This relationship remains after controlling for physical proximity, prior case growth, and regional characteristics.
  • Lagged social proximity to cases significantly predicts subsequent case growth in every two-week period from March 30 through November 2.
  • Google symptom-search changes are strongly correlated with case-growth changes, with weaker persistence for one-period-lagged searches.
  • Social proximity remains a significant predictor of future case growth after adding the other predictive measures.

C Out-Of-Sample Prediction: COVID-19 Deaths

An out-of-sample exercise for COVID-19 deaths compares models with and without social proximity to deaths, finding lower prediction error when that measure is included.

  • The exercise evaluates out-of-sample predictions of COVID-19 deaths rather than cases.
  • Including lagged social proximity to COVID-19 deaths lowers RMSE in every period.
  • The models use random forests trained on data from periods preceding the period being predicted.

D Additional Details on Google Symptom Search Data

The paper constructs county-level COVID-19 symptom-search measures from Google data, combining daily and weekly series while accounting for unavailable observations and aggregating changes biweekly.

  • Google provides county-by-week normalized probabilities that users make searches related to COVID-19 symptoms.Data may be available at daily or weekly frequency depending on Google quality and privacy thresholds.
  • When daily measures are available, they are averaged by week to create county-week time series; weekly measures fill periods without qualifying daily data.Measures unavailable at both frequencies are omitted.
  • The analysis uses searches for three common COVID-19 symptoms: fever, cough, and fatigue.
  • For biweekly analysis, symptom-search changes are defined as the percent change between the second week of the current period and the second week of the previous period.The prediction exercise uses a one-period lagged version of this measure.
Loading 2004.03055v3…