Source-linked AI summary
Social Contagion Theory: Examining Dynamic Social Networks and Human Behavior
Nicholas A. Christakis, James H. Fowler
TL;DR
Determining interpersonal effects in social networks remains difficult because available data and methods rely on imperfect observations and assumptions. This paper reviews longitudinal network analyses and finds that associations can extend beyond one degree but often fade within about three degrees.
Problem
Determining interpersonal effects with confidence is difficult because network data and methods are imperfect and subject to assumptions or biases.
Method
The paper examines longitudinal social-network data using systematic tie ascertainment and simulations of fully versus partially observed network paths.
Results
Associations between behaviors or attributes often extend to three degrees of separation, including links between individuals and friends’ friends’ friends.
Takeaways & Limitations
The evidence supports interpersonal associations spreading beyond one degree while fading within a few degrees across phenomena and datasets.
Takeaways & Limitations
The authors acknowledge that interpersonal effects are difficult to discern confidently because network data and methods remain imperfect and assumption-dependent.
Abstract
from arXiv · showhide
Here, we review the research we have done on social contagion. We describe the methods we have employed (and the assumptions they have entailed) in order to examine several datasets with complementary strengths and weaknesses, including the Framingham Heart Study, the National Longitudinal Study of Adolescent Health, and other observational and experimental datasets that we and others have collected. We describe the regularities that led us to propose that human social networks may exhibit a "three degrees of influence" property, and we review statistical approaches we have used to characterize inter-personal influence with respect to phenomena as diverse as obesity, smoking, cooperation, and happiness. We do not claim that this work is the final word, but we do believe that it provides some novel, informative, and stimulating evidence regarding social contagion in longitudinally followed networks. Along with other scholars, we are working to develop new methods for identifying causal effects using social network data, and we believe that this area is ripe for statistical development as current methods have known and often unavoidable limitations.
(1) The FHS-Net Data and Its Pertinent Features
The FHS-Net combines longitudinal follow-up with systematically collected family and social ties, enabling observation of relationships and outcomes across participants and nonparticipants. Its network coverage includes substantial friendship, coworker, and neighborhood connections, with limited observed differences between sampled and unsampled contacts on several behaviors and outcomes.
- Cohorts and follow-up: The Framingham Heart Study enrolled 5,209 people in the Original Cohort in 1948 and 5,124 people in the Offspring Cohort in 1971, with almost no Offspring loss to follow-up.Only 10 Offspring Cohort cases dropped out, apart from deaths.
- Network construction: FHS name generators systematically recorded first-order relatives and at least one unrelated close friend, providing broad social-tie ascertainment.The network also captured changing family and social contacts over follow-up as relationships, residences, employment, and friendships changed.
- Network coverage: 53,228 observed familial and social ties connected 5,124 subjects observed from 1971 to 2009, including ties to people both inside and outside the sample.Attributes could be observed longitudinally for connected individuals who also participated in the FHS.
- Network coverage: 45% of the 5,124 subjects had an FHS friendship tie, 39% had at least one coworker captured, and 10% had an immediate non-relative residential neighbor present.The dataset contained 3,542 friendships, averaging 0.7 friendship ties per subject.
- Sample representativeness: Egos whose named contacts participated in the FHS did not significantly differ from those whose contacts did not participate in weight, smoking, alcohol consumption, happiness, loneliness, or depression.The authors report that sampled alter types and counts were generally not far from unrestricted national-sample data.
(2) Basic Analyses and Findings: Clustering
The analyses test whether traits cluster among connected individuals beyond chance using permutation-based comparisons, finding associations that often extend to three degrees of separation. These clustering results support broad empirical regularities but do not by themselves distinguish influence from homophily or shared context.
- Clustering method: Permutation tests preserve network topology and trait prevalence while comparing observed clustering with randomly assigned traits to assess departures from chance.The resulting confidence intervals test whether observed-minus-permuted clustering differs from zero.
- Three-degree pattern: In many empirical cases, trait correlations between egos and alters remain statistically significant and substantively meaningful up to three degrees of separation.Other datasets with more complete tie ascertainment also often show clustering to three degrees.
- Interpretation and limitations: Observed clustering establishes prediction beyond chance, but it cannot determine whether the association reflects influence, homophily, or shared contextual factors.The null model rejects only the simplest random-assignment explanation and may omit more complex network assumptions.
- Interpretation and limitations: The authors describe the three-degree pattern as evidence that diverse phenomena spread beyond one degree while associations fade within a few degrees, not as a deterministic rule.They acknowledge that the evidence directly concerns clustering rather than influence.
- Variation across phenomena: Detectable clustering varies by phenomenon: regression analyses find friend correlations for obesity but not neighbors, while health screening and sexual orientation show no spread across observed ties.Thus, different traits may spread in different ways and to different extents.
- Variation across phenomena: Permutation results also raise the possibility that traits can skip over an intermediary in a social chain, rather than spreading through every successive tie.This possibility depends on assumptions about how change occurs along paths.
(3) Partial Observation of FHS-Net Ties
Partial observation of FHS-Net ties does not necessarily overestimate the distance over which influence spreads: sampled shortest paths can be shorter or longer than actual paths, depending on network structure and path multiplicity. Similar three-degree clustering also appears in datasets with more fully observed ties, including Add Health.
- Partial observation: Partial observation does not necessarily overestimate influence-path length because actual, shortest, and sampled paths are distinct.The authors distinguish the unobservable stochastic transmission path from shortest paths in fully observed and sampled networks.
- Partial observation: 3.9 million cell phone users showed that shortest paths in sampled networks may be shorter than actual paths.Sampling can produce path lengths that are either shorter or longer than the paths taken by diffusion.
- Partial observation: A sampled shortest path cannot be shorter than the original shortest path, but it may be equal or longer when portions of that path disappear.If the original shortest path has length three, the sampled path may remain three when alternate paths exist or become longer.
- Partial observation: Transmission probability depends on the number of alternative paths, so a longer sampled shortest path need not represent the most likely transmission distance.Multiple paths of length four can collectively make four-step transmission more probable than transmission through a single three-step path.
- Partial observation: Three-degree clustering also appears in networks with almost fully observed ties, including Add Health, where subjects could name up to 10 friends.Ninety percent of Add Health subjects named fewer than the maximum number of friends.
(5) An Identification Strategy Involving Directional Ties
The section presents friendship directionality as an identification strategy for assessing peer-effect causality, while emphasizing that its interpretation depends on statistical assumptions and testing choices. Evidence across behaviors and datasets supports an ordered pattern of effects, but overlapping confidence intervals and methodological limitations temper the conclusions.
- Identification strategy: The proposed strategy uses differences across tie directions to shed light on confounding and causal ordering.It analogizes directional ties to the directional nature of time, which researchers use to establish a sequence consistent with causal ordering.
- Identification strategy: Directional friendship ties distinguish mutual, ego-perceived, and alter-perceived relationships for studying asymmetric peer effects.Sociocentric network studies record both nominations within each friendship, making these three directional categories observable.
- Evidence: The pattern mutual tie > ego-perceived tie > alter-perceived tie generally appears across many behaviors and affective states in two datasets.This directional pattern is presented as evidence against covariance being explained by unobserved contemporaneous exposures shared by both friends.
- Limitations: Confidence intervals for the three friendship types often overlap, so the reported ordering does not by itself establish pairwise statistical significance.The papers reported confidence intervals and comparison tests derived from a single model with an interaction term.
- Limitations: Whether the directional pattern can be stated confidently depends on the null hypothesis being tested.Testing whether all three relationships share one distribution or whether their effects follow the specified order differs from testing a two-group comparison.
- Related methodological work: Subsequent work has examined the strengths and limitations of the network directionality test and identified additional assumptions that may be necessary or implicit.The section notes contributions from computer scientists, econometricians, statisticians, and other researchers.
(6) Using Geographic Information to Address Certain Types of Confounding
Geographic information helps distinguish social influence from shared physical context. Physical distance did not alter correlations in obesity and health behaviors, whereas affective-state associations were positive only among nearby friends and siblings.
- Geographic confounding: Geographic distance between participants varied substantially, with the most distant sextile of friends’ residences averaging nearly 500 miles.Residence data, including changes over time, enabled comparison of geographic and social distance as potential explanations for correlated outcomes.
- Health behaviors: For obesity, smoking, and drinking, geographic distance had no discernable role in outcome correlations.The interaction between geographic distance and the alter’s outcome at time t+1 had a coefficient near 0 and was insignificant.
- Health behaviors: A friend living hundreds of miles away appeared to have a similar effect as one living next door, suggesting social distance mattered more than physical distance.This finding was reported for obesity and other health-behavior follow-up studies.
- Affective states: For happiness, loneliness, and depression, positive associations occurred only among friends and siblings living within a few miles.The authors suggested physical proximity may be required for affective contagion, while noting that this interpretation is not definitive.
(7) Availability of Data and Code
The authors have publicly released substantial network data and code and shared additional materials on request. FHS-Net access remained limited by clinical-record origins and study rules, requiring secure collaboration and confidentiality-preserving data changes that affect replication.
- The authors placed substantial network data and code in the public domain and promptly shared code and supplementary results with requesters.Examples include Facebook, biological, experimental, and political network datasets; Add Health is publicly available.
- FHS-Net data could not be fully released because of clinical-record origins and FHS rules, limiting outside researchers’ ability to replicate FHS-based results.Data were shared with collaborators through secure servers, and administrators posted a secure version in 2009.
- The posted FHS data were modified to protect confidentiality, including monthly rather than daily dates, 9,000 rather than 12,000 cases, exclusions of nonconsenting individuals, and available covariates.The changes also removed certain non-genetically related relative ties, including adopted siblings and step-children.
(8) Social Influence and Social Networks
The section argues that social-network data provide evidence of interpersonal influence across diverse behaviors, while emphasizing that causal interpretation remains difficult because observational methods rely on imperfect assumptions. It therefore supports continued methodological development, including experiments and instrumental-variable approaches.
- Evidence for social influence: Obesity may spread through social networks in a quantifiable pattern shaped more by social ties than geographic distance.The cited discussion notes that common exposures and shared events could otherwise explain simultaneous weight changes.
- Evidence for social influence: Immediate neighbors’ weight gain did not affect egos, and geographic distance did not alter effects for friends or siblings, weakening local-environment explanations.Models also controlled for egos’ previous weight status to address time-stable confounding.
- Related work: Other investigators have used different datasets and approaches to confirm the authors’ findings, often reproducing the magnitude of the observed effects.The section places this corroboration within a broader tradition of research on peer effects.
- Limitations and causal inference: Interpersonal effects are difficult to discern confidently because network data and methods are imperfect and subject to assumptions or biases.The authors emphasize transparency about these limitations and avoid strong mechanistic claims in scientific papers.
- Limitations and causal inference: Experiments and gene-based instrumental-variable methods are proposed as complementary approaches that may provide different confidence in causal inference.The authors argue that scientifically accurate observation is possible as network data become increasingly available.
- Evidence for social influence: 2 to 4 degrees of separation showed significant associations across 14 behaviors and affective states in four observational and experimental datasets.Network permutation tests compared the probability that an ego had a trait when an alter had it versus when the alter did not.