Source-linked AI summary
The Spread of Evidence-Poor Medicine via Flawed Social-Network Analysis
Russell Lyons
TL;DR
The paper examines whether highly publicized Christakis and Fowler studies provide evidence that personal characteristics spread through social networks. By critiquing their models, estimation methods, statistical tests, and interpretations, it finds that the studies do not establish transmission or distinguish it from alternative explanations. The authors use the episode to argue for more explicit assumptions, stronger review, and improved statistics education.
Problem
Professional publications, including top medical journals, can unwittingly misuse statistics, making the evidential basis of influential social-network contagion claims important to examine.
Method
The paper presents detailed statistical critiques of Christakis and Fowler’s observational Framingham network analyses, including their models, estimation procedures, significance tests, and interpretation of directional associations.
Results
The studies do not support their claims of social transmission or three degrees of influence; their reported differences are not statistically significant and remain consistent with homophily, shared environment, and induction.
Takeaways & Limitations
The episode supports stating models and assumptions explicitly, reporting estimation methods clearly, and placing greater emphasis on critical statistics education and scientific review.
Takeaways & Limitations
The Framingham friendship data are unavailable to others and sparse, preventing basic replication and potentially making the assembled network misleading.
Abstract
from arXiv · showhide
The chronic widespread misuse of statistics is usually inadvertent, not intentional. We find cautionary examples in a series of recent papers by Christakis and Fowler that advance statistical arguments for the transmission via social networks of various personal characteristics, including obesity, smoking cessation, happiness, and loneliness. Those papers also assert that such influence extends to three degrees of separation in social networks. We shall show that these conclusions do not follow from Christakis and Fowler's statistical analyses. In fact, their studies even provide some evidence against the existence of such transmission. The errors that we expose arose, in part, because the assumptions behind the statistical procedures used were insufficiently examined, not only by the authors, but also by the reviewers. Our examples are instructive because the practitioners are highly reputed, their results have received enormous popular attention, and the journals that published their studies are among the most respected in the world. An educational bonus emerges from the difficulty we report in getting our critique published. We discuss the relevance of this episode to understanding statistical literacy and the role of scientific review, as well as to reforming statistics education.
1. Introduction
The paper critiques widely acclaimed Christakis and Fowler studies claiming that personal characteristics spread through social networks to three degrees of separation. It argues that flawed models, statistical interpretations, and assumptions leave little support for those conclusions and highlight the need for stronger statistical education and review.
- 1. Introduction: The paper uses these cases to expose failures in statistical assumptions, scientific review, and quality control, and to call for more critical statistics education.The authors aim their analysis at educators, practitioners, and readers concerned with the quality of statistical research.
- 1. Introduction: Christakis and Fowler infer social-network transmission of obesity, smoking cessation, happiness, and loneliness, extending their proposed influence to three degrees of separation.Their studies analyzed observational data from the Framingham Heart Study using statistical techniques that produced these two major inferences.
- 1. Introduction: The paper argues that C&F’s evidence does not establish either transmission or the three-degrees-of-influence rule, and may even suggest that transmission is untrue.The critique targets both the evidential basis and the interpretation of the original studies’ results.
- 1. Introduction: C&F analyze obesity associations in the Framingham network using observational data and logistic regression models incorporating participants’ prior obesity and linked participants’ current and prior obesity.The analysis focuses on roughly 5,000 Offspring Cohort participants in relation to all roughly 12,000 study participants.
- 1. Introduction: The critique identifies two central errors: models that contradict the data and conclusions, and incorrect interpretations of statistical models and tests.These errors undermine the arguments offered to substantiate C&F’s claims.
- 1. Introduction: The reported 57% and 13% obesity-risk increases are model-derived and statistically uncertain; closer analysis finds them indistinguishable, while lagged obesity terms may amplify rather than remove homophily.The proposed directional differences also remain consistent with homophily, shared environment, and induction, so the explanations are not distinguished.
- 1. Introduction: The claimed three-degree pattern may partly reflect sparse network data, because unrecorded friendship ties can make the assembled network misleading.The paper therefore treats the observed three-degree rule as constrained by the structure and incompleteness of the data.
2. Directionality
Christakis and Fowler interpret directional friendship differences as evidence of obesity transmission, but the critique argues those differences are neither statistically significant nor diagnostic of causation. The same ordering can arise from homophily or shared environment, and their modeling approach has additional inferential problems.
- Directional-effects argument: C&F’s directional-effects argument rests on comparing mutual, FP→LP, and LP→FP friendship associations to distinguish induction from homophily and shared environment.They interpret stronger mutual-friend and FP→LP associations as evidence that influence is directional.
- Statistical significance: 57% and 13% are not statistically distinguishable because each estimate falls within the other direction’s confidence interval.The reported intervals are 6% to 123% for FP→LP and approximately −28% to 68% for LP→FP.
- Model comparison: Comparing coefficients from different models cannot support an inference about their difference without a valid model containing both coefficients.The unavailable Framingham network data also prevents basic replication and makes errors harder to detect.
- Homophily control: Including the linked participant’s lagged obesity does not establish that the estimated association is net of homophily.The critique notes that the current and lagged coefficients have opposite signs and sum to approximately zero, making their interpretation unstable.
- Alternative explanations: Directional differences are consistent with induction, homophily, and shared environment rather than distinguishing among these explanations.In a nearest-neighbor friendship process, the expected ordering is FP↔LP strongest, followed by FP→LP and LP→FP under both similarity-based selection and geographic proximity.
3. Random Networks
C&F’s random-network procedure shows that associations extend across three degrees in the observed network, but this result describes the assembled network rather than demonstrating real-world contagion. The critique emphasizes that sparse friendship data and the separation between random-network analysis and causal regression weaken the interpretation.
- Random-network procedure: C&F preserve the network while randomly redistributing obesity, retaining the same number of obese individuals, and compare the observed network with randomized networks.They report the same three-degree pattern for obesity, smoking cessation, happiness, and loneliness.
- Data limitations: Only 45% of 5124 focal participants named a friend, leaving about 0.7 observed current friends per focal participant.The critique therefore characterizes the friendship network as thin and incomplete.
- Interpretation: The three-degree pattern appears in the assembled network but has not been demonstrated for the real world because unrecorded friendships distort the network.The network treats some people as nonfriends even when they are friends in reality.
- Causal interpretation: Random-network analysis summarizes observed network structure but is unrelated to the cause of the statistical associations.C&F use regression models, rather than the randomization procedure, to argue for causal conclusions.
- Model versus reality: Calling the three-degree pattern a risk or real-world prediction blurs the distinction between model output and reality.The critique specifically rejects interpreting the reported 52% loneliness figure as a friendship risk or prediction about the real world.
4. Modeling
Christakis and Fowler model obesity and loneliness within social networks using regression equations intended to describe associations and support causal inference. The critique argues that these models contradict the data and conclusions, including through mathematically inconsistent directionality claims.
- Modeling: C&F use logistic or linear regression models because they lack experimental data and sufficient data for multidimensional cross-tabulation.Their analyses therefore depend on assumptions about what the data would look like under an experiment or with much more data.
- Modeling: The obesity model predicts an individual’s current status from the connected person’s current and previous status, prior individual status, and covariates.The covariates include exam timing, age, sex, and education; the lagged connected-person status is intended to control for homophily.
- Modeling: The critique identifies a mismatch between C&F’s Fig. 3 caption and text about whether obesity risk was measured in an ego or an alter.The same interchange of “alter” and “ego” also occurs in Fig. 2 of another C&F paper.
- Modeling: The model’s one-equation-per-tie structure creates too many equations when an individual names multiple connected people.For network ties, dependent variables appear both as outcomes and predictors, producing simultaneous-equation constraints.
- Modeling: The model contradicts the data unless β1 = β2 = 0, implying that no individual affects another at any time.A separate algebraic argument shows β1 = 0 regardless of the data, contradicting C&F’s directionality conclusions.
5. Model Estimation
The critique examines C&F’s use of generalized estimating equations for dependent network observations. It argues that the estimation method does not apply to their model because the model implies β1 = 0 while C&F estimate β1 ≠ 0.
- Model Estimation: C&F estimate 12 coefficients with generalized estimating equations, a method designed for repeated measures or other dependencies.GEE nevertheless assumes independence among subjects or groups of subjects.
- Model Estimation: C&F’s stated clustering appears to group all measurements on each focal participant, but those groups are not independent because the analysis concerns their dependencies.GEE also relies on having a large number of independent groups.
- Model Estimation: The dependent variables appear on both sides of C&F’s network equations, a structure outside the authors’ cited GEE literature.The critique says the method might work if the model holds, but mathematical proof or relevant simulations would be needed.
- Model Estimation: Because the model implies β1 = 0 while C&F estimate β1 ≠ 0, the critique concludes that their estimation method is faulty.This conclusion follows from the model’s internal constraint rather than from a disagreement over a particular estimated effect.
6. The Role of Review
The paper uses the C&F episode to question statistical quality control and the treatment of critiques in leading journals. It reports both widespread concern about statistical errors and substantial barriers to publishing the critique.
- The Role of Review: C&F’s papers appeared in highly respected medical and psychology journals, including the New England Journal of Medicine, BMJ, and the Journal of Personality and Social Psychology.The critique emphasizes that publication in prestigious venues did not prevent fundamental statistical errors.
- The Role of Review: A senior BMJ statistics editor attributed the abundance of statistical errors in published medical articles largely to inadequate statistical review.The paper presents this assessment as evidence relevant to evaluating peer review at top journals.
- The Role of Review: Problems with peer review are longstanding, and the authors report that one proposed remedy has already been shown to fail.They also note that the policies of journals toward critiques had not been systematically studied.
- The Role of Review: The New England Journal of Medicine and BMJ rejected the critique without peer review, while JAMA, Lancet, and PNAS declined because they had not published the criticized studies.The psychology journal involved reportedly required new data for critiques, even of papers it had published.
- The Role of Review: Statistical Science received three referee reports after five months, with two recommending publication after revisions and one questioning the critique’s importance.The editor decided two months after receiving the reports.
- The Role of Review: The authors propose a journal devoted specifically to critiques to improve knowledge about which studies are trustworthy.They argue that methodological cautions and recommendations are often ignored.
7. Conclusions
The paper argues that flawed statistical assumptions and methods can make published analyses misleading, especially when observational data are modeled without adequate scrutiny. It calls for stronger statistical literacy, explicit reporting of models and assumptions, and greater emphasis on critical thinking.
- Problems in C&F’s studies: The authors summarize seven problems in C&F’s studies, including unavailable sparse data, contradictory models, invalid estimation, unsupported significance tests, and failure to distinguish competing explanations.They also report that associations at a distance may be better explained by homophily than induction.
- Broader implications: They recommend that analysts state models, equations, estimation methods, and assumptions explicitly so readers and reviewers can evaluate them.The paper presents this clarity as useful for published research and statistical review.
- Broader implications: The authors argue that statistics education should devote more attention to assumptions and practical evaluation, particularly in advanced courses where assumptions are subtler.They note that students often learn techniques without questioning whether their assumptions hold in real situations.
- Broader implications: The paper links these errors to a broader pattern in which statistical ideas are adopted easily but resisted when they conflict with investigators’ vested interests.The authors identify questioning assumptions as a way to counteract this tendency.
- Broader implications: The authors extend the concern beyond medicine and academia, citing flawed statistical models in economics and urging educators to teach critical thinking.They frame improved statistics education as an overdue response to recurring misuse of statistical models.
Appendix A. Directionality Table
The appendix explains how directional friendship estimates are represented and why reported probabilities and confidence intervals require careful interpretation. Logistic regression’s nonlinear transformation means that population-mean covariate values yield only a vague average of individual results.
- Interpretation of reported probabilities: Logistic regression transforms coefficient values into probabilities, but calculating a probability requires choosing values for every covariate.Because the transformation is nonlinear, changing one coefficient produces uncertainty that depends on the selected values of the other covariates.
- Interpretation of reported probabilities: C&F report probabilities and confidence intervals after assigning covariates their population means, producing a vague average rather than values representing any individual participant.For example, a binary gender covariate becomes 1/2 even though each person has a value of 0 or 1.
- Table 1: Table 1 reports directional differences for friendship ties using FP↔LP, FP→LP, and LP→FP, with 95% confidence intervals and standard errors.FP denotes the focal participant or ego, while LP denotes the linked participant or alter.
Appendix B. Further Lack of Statistical Significance
The appendix identifies further errors in interpreting statistical significance. C&F treat a significant-versus-nonsignificant contrast as evidence of a difference and sometimes treat statistically indistinguishable effects as zero.
- Statistically insignificant comparisons: C&F’s comparison of coresident and non-coresident spouses does not establish that only coresident spouses affect happiness.The non-coresident estimate’s confidence interval, −18% to 31%, overlaps the coresident estimate’s interval, 0.2% to 16%.
- Statistically insignificant conclusions: C&F sometimes report that an effect is zero when their methods show only that it cannot be statistically distinguished from zero.The paper illustrates this with the claim that obesity in an opposite-sex sibling did not affect the other sibling’s obesity risk.
- Statistically insignificant conclusions: Table 3 documents cases where a coefficient is labeled statistically insignificant even though its confidence interval is not close to zero.The authors distinguish lack of statistical significance from evidence that the effect itself equals zero.