Source-linked AI summary

Assessing Vaccination Sentiments with Online Social Media: Implications for Infectious Disease Dynamics and Control

Marcel Salathé, Shashank Khandelwal

arXiv:1105.4502v2cs.SIphysics.soc-phq-bio.PE

TL;DR

Measuring health behaviors across populations is resource-intensive, motivating the use of online social media as an alternative data source. The paper analyzes vaccination-related social-media data to measure sentiment over time and space, network opinion structure, and outbreak implications. It finds strong regional agreement with CDC-estimated vaccination rates, same-sentiment information flow, polarized communities, and higher outbreak likelihood when susceptibility is clustered.

  • Problem

    Measuring population health behaviors over time and space requires substantial resources.

  • Method

    The study used publicly available social-media data to measure vaccination sentiment, analyze opinionated-user networks, and simulate infectious-disease transmission.

  • Results

    The study found strong correlation with CDC-estimated vaccination rates, same-sentiment information flow, predominantly positive or negative communities, and increased outbreak likelihood under clustered susceptibility.

  • Takeaways & Limitations

    Online social media can provide inexpensive, efficient access for identifying intervention target areas and evaluating intervention effectiveness.

  • Takeaways & Limitations

    Because the study is observational, the data cannot determine whether information-flow links cause shared vaccination sentiments and cannot exclude confounding.

Abstract

from arXiv · show

There is great interest in the dynamics of health behaviors in social networks and how they affect collective public health outcomes, but measuring population health behaviors over time and space requires substantial resources. Here, we use publicly available data from 101,853 users of online social media collected over a time period of almost six months to measure the spatio-temporal sentiment towards a new vaccine. We validated our approach by identifying a strong correlation between sentiments expressed online and CDC- estimated vaccination rates by region. Analysis of the network of opinionated users showed that information flows more often between users who share the same sentiments - and less often between users who do not share the same sentiments - than expected by chance alone. We also found that most communities are dominated by either positive or negative sentiments towards the novel vaccine. Simulations of infectious disease transmission show that if clusters of negative vaccine sentiments lead to clusters of unprotected individuals, the likelihood of disease outbreaks are greatly increased. Online social media provide unprecedented access to data allowing for inexpensive and efficient tools to identify target areas for intervention efforts and to evaluate their effectiveness.

AUTHOR SUMMARY

The study uses publicly available online social-media data to measure vaccination sentiments over time and space. It finds that sentiment clustering could increase disease-outbreak risks if it corresponds to clustered vaccination.

  • Measuring vaccination sentiments and their population distribution is difficult and resource-intensive.
  • Publicly available Twitter data were used to measure the evolution and distribution of sentiment toward the novel influenza A(H1N1) vaccine during fall 2009.
  • Positive and negative vaccination opinions were clustered in the online social network.
  • A similarly clustered distribution of vaccination would strongly increase disease-outbreak risks.

INTRODUCTION

The introduction frames online social media as a survey-free, potentially real-time way to measure health behavior across populations. The paper applies this approach to vaccination sentiment, network structure, and outbreak risk during the 2009 H1N1 pandemic.

  • Traditional survey methods for measuring health behaviors across time and space are labor-intensive and expensive.
  • Online social media create new possibilities for measuring health behavior without survey responses, often in real time.
  • The study collected publicly available English-language vaccination tweets and user-location information from Twitter in the United States from August 2009 to January 2010.
  • A manually labeled tweet subset trained a machine-learning classifier to predict positive, negative, or neutral vaccination sentiment in the remaining messages.
  • The classified data produced localized temporal sentiment scores and an information-flow network for studying opinion distribution.
  • The study extrapolated these findings to empirical infectious-disease contact networks to examine how non-random vaccination distributions affect outbreak likelihood.

RESULTS

Online vaccination sentiment tracked CDC-estimated vaccination coverage and showed strong homophily and community polarization. Simulations indicated that clustered susceptibility substantially raises the probability of large outbreaks.

  • 318,379 of 477,768 collected tweets were classified as relevant to influenza A(H1N1) vaccination.
  • Temporal sentiment: The vaccination sentiment score began negative in late summer 2009 and its 14-day moving average became positive in mid-October as the vaccine became available.
  • Validation: The online sentiment score correlated strongly with estimated vaccination coverage at both HHS-region and state levels.Weighted Pearson correlations were r = 0.78 at the HHS-region level and r = 0.52 at the state level.
  • Network structure: The opinionated network of 39,284 users had assortativity r = 0.144, exceeding randomized-network maxima of 0.0056.
  • Network structure: Users received significantly more information from same-sentiment users than expected under randomized opinions.The original mean same-sentiment incoming-edge fraction was 0.601 versus a randomized mean of means of 0.531, with p < 10^-95 for all tests.
  • Community structure: Nearly every sufficiently large community was significantly more positive or negative than the giant-component average, ranging from p(-) = 0.764 to p(-) = 0.266.

DISCUSSION

The study shows that online vaccination sentiments form clustered, opinion-homogeneous communities, while cautioning that the observed links may reflect confounding factors. Simulations indicate that corresponding clustering in vaccination status can increase outbreak risk, especially near herd-immunity thresholds.

  • Network structure: Almost 40,000 opinionated users showed more information flow between users sharing the same vaccination sentiments than expected by chance.This assortative structure is consistent with online social media acting as an echo chamber for medical opinions.
  • Network structure: Most communities were dominated by either positive or negative sentiments toward the novel vaccine.The study therefore identifies substantial opinion clustering rather than a uniformly mixed sentiment distribution.
  • Disease dynamics: If sentiment clusters produce similarly clustered vaccination status, the probability of large disease outbreaks is greatly increased.The simulations link non-random vaccination distributions to outbreak likelihood, without assuming that online and real-world contact networks strongly overlap.
  • Disease dynamics: This outbreak-risk effect is strongest when vaccination coverage approaches the herd-immunity level expected under a random distribution.Communities with very low vaccination rates may lack herd-immunity protection even when overall coverage is high.
  • Implications: Communication strategies that reduce assortative vaccination distributions may complement efforts to increase overall vaccination rates.The authors suggest that identifying high-risk communities would benefit from empirical data on assortativity in vaccination status.
  • Implications: Public online social-media data can support inexpensive, efficient identification of regional areas for intensified vaccine communication.Twitter-like network data also expose information flows and the social structures in which opinions and behaviors spread.

Data Collection

The study collected English-language tweets about vaccination in real time, together with tweet timing, user locations, and social-network identifiers, from August 25, 2009 through January 19, 2010.

  • Tweets were collected in English when they contained vaccination-related search terms, including vaccine, vaccination, immunization, or immunized.
  • The dataset included tweet text, publication date and time, user location when available, user IDs, follower IDs, and friend IDs.
  • Information flow was defined from a user to that user’s followers, who receive the user’s messages.
  • Tweets were collected daily in real time until January 19, 2010, aside from occasional short-term technical disruptions.

Sentiment Analysis

Tweets were assigned positive, negative, neutral, or irrelevant vaccination sentiment using manually rated data to train and evaluate machine-learning classifiers.

  • Each tweet was classified as positive, negative, neutral, or irrelevant with respect to influenza A(H1N1) vaccination.
  • 64 students submitted 88,237 ratings to create the manually labeled sentiment dataset.
  • The high-confidence test set contained 630 tweets after filtering for sufficient ratings and agreement criteria.
  • The ensemble combined Naive Bayes for positive and negative tweets with Maximum Entropy for neutral and irrelevant tweets.
  • 84.29% was the accuracy of the ensemble classifier.
  • Short tweets, non-standard language, and limited context made perfect automated classification unrealistic.

Geocoding

User-provided locations were resolved to state and country information where possible, with ambiguous or unresolved locations handled through additional review or exclusion.

  • Twitter profile locations were free-form, could omit a location, and sometimes changed over time.
  • The study aimed to resolve user locations within the United States to the state level.
  • Yahoo! PlaceFinder converted recognized location strings into state- and country-level information.
  • Users reporting locations in multiple states were excluded after locations were resolved.

Network creation

The network represented opinionated Twitter users as nodes and information-flow relationships as directed edges, producing a large network with a dominant giant component.

  • Each user with at least one positive, negative, or neutral tweet was represented by a network node.
  • A directed edge connected users when a follower relationship indicated information flow between them.
  • The network was treated as static rather than dynamic, despite follower and friend counts generally increasing over time.
  • The collection cutoff at 5,000 friends or followers was addressed by searching both follower and friend lists.
  • Users without positive or negative influenza A(H1N1) vaccine sentiment scores were removed from the network.
  • The overall vaccine sentiment score was calculated as (n+ - n-) / (n+ + n- + n0).
  • 39,284 opinionated users formed 685,719 edges in the resulting network.
  • 34,025 nodes and 685,390 edges belonged to the giant component, containing 99.95% of network edges.

Disease simulations

The study uses an SEIR model on an empirical high-school contact network to examine how vaccination coverage and assortative clustering affect influenza-like disease outbreaks.

  • Model framework: The simulations use an SEIR model parameterized with influenza outbreak data, with transmission occurring exclusively along measured contacts lasting at least 30 minutes.Individuals occupy susceptible, exposed, infectious, or recovered classes.
  • Model framework: Vaccinated individuals are assigned to the recovered class, while all other individuals initially begin susceptible and one random susceptible becomes the index case.The simulation ends when no exposed or infectious individuals remain.
  • Model assumptions: Transmission is restricted to daytime weekdays because the high-school network excludes contacts outside school.The authors explicitly note that this assumption does not hold in reality but permits analysis of spread from a single infected case.
  • Vaccination distributions: Vaccination coverage is fixed at 0.624, matching the proportion of positive Twitter sentiments, while vaccination assortativity is varied through status redistribution.The redistribution algorithm preserves coverage while increasing the assortativity index r.
  • Simulation experiment: The experiment generated 100,000 vaccination redistributions for each minimum r from 0 to 0.145 in increments of 0.005, producing 3,000,000 simulation runs.All runs used constant network structure and vaccination coverage but different assortativity values.

FIGURE LEGENDS

The figure legends describe vaccine sentiment over time and space, sentiment composition within network communities, and how assortative vaccination patterns relate to outbreak risk.

  • Figure 1: Figure 1C compares estimated vaccination rates with sentiment scores across HHS regions and states using separate regression lines.Black dots represent HHS regions, gray dots represent states, and region numbers follow the U.S. Department of Health and Human Services definition.
  • Figure 2: Most network communities have negative or positive sentiment proportions significantly different from the overall opinionated-network proportions, except community E.Figure 2A uses p(-), the proportion of negative sentiments, with a dashed line for the overall network proportion.
  • Figure 2: Figure 2C compares relative outbreak-risk increases across outbreak sizes for assortativity values r=0.075 and r=0.145, the latter matching the opinionated Twitter network.The comparison is against the random vaccination distribution with r~0.
Loading 1105.4502v2…