Source-linked AI summary

Identification and estimation of treatment and interference effects in observational studies on networks

Laura Forastiere, Edoardo M. Airoldi, Fabrizia Mealli

arXiv:1609.06245v4stat.ME

TL;DR

Interference complicates causal inference in networked observational studies and can make individual treatment effects misleading when neighborhood influence is ignored. The paper defines treatment and spillover effects, derives bias results, and develops propensity-score adjustment methods whose simulations support estimation under the stated model.

  • Problem

    Interference matters both as a source of bias in individual treatment-effect estimation and as a mechanism scientists and policymakers may seek to understand or leverage.

  • Method

    The paper derives bias expressions for naive estimators, extends adjustment to individual and neighborhood covariates, and uses propensity-score methods to estimate main and spillover effects.

  • Results

    The derived bias combines unmeasured-confounding and interference components, with the interference component proportional to interference intensity and the association between individual and neighborhood treatments; simulations assess the proposed estimators.

  • Takeaways & Limitations

    Adjustment for the joint individual and neighborhood propensity score is needed in the simulated settings to estimate treatment and spillover effects while addressing interference.

  • Takeaways & Limitations

    The estimands condition on a fully known, fixed observed network, while residual subclass imbalance and neighborhood-propensity model misspecification may bias real-study estimates.

Abstract

from arXiv · show

Causal inference on a population of units connected through a network often presents technical challenges, including how to account for interference. In the presence of local interference, for instance, potential outcomes of a unit depend on its treatment as well as on the treatments of other local units, such as its neighbors according to the network. In observational studies, a further complication is that the typical unconfoundedness assumption must be extended - say, to include the treatment of neighbors, and indi- vidual and neighborhood covariates - to guarantee identification and valid inference. Here, we propose new estimands that define treatment and interference effects. We then derive analytical expressions for the bias of a naive estimator that wrongly assumes away interference. The bias depends on the level of interference but also on the degree of association between individual and neighborhood treatments. We propose an extended unconfoundedness assumption that accounts for interference, and we develop new covariate-adjustment methods that lead to valid estimates of treatment and interference effects in observational studies on networks. Estimation is based on a generalized propensity score that balances individual and neighborhood covariates across units under different levels of individual treatment and of exposure to neighbors' treatment. We carry out simulations, calibrated using friendship networks and covariates in a nationally representative longitudinal study of adolescents in grades 7-12, in the United States, to explore finite-sample performance in different realistic settings.

1 Introduction

The paper develops causal inference methods for estimating individual and spillover effects in observational network studies with interference. It formalizes network interference, derives bias from ignoring it, and extends propensity-score adjustment to account for individual and neighborhood treatments and covariates.

  • Motivation: Interference can be a nuisance for individual-effect estimation but is also substantively relevant when spillovers may be leveraged or prevented.
  • Motivation: The paper targets estimation of individual treatment and spillover effects when units are connected by a known network and treatment assignment is observational.
  • Contributions: The authors formalize network interference under potential outcomes and define causal estimands based on individual treatment and summarized treatment among immediate neighbors.
  • Contributions: They derive bias formulas for naive estimators that assume no interference, identifying interference and residual individual-neighborhood treatment association as key drivers.
  • Contributions: They define a generalized propensity score whose balancing properties support covariate adjustment across joint individual and neighborhood treatment levels.
  • Evaluation: Simulations use friendship-network and covariate data to assess proposed estimators, bias from neglecting interference, bias formulas, and adjustment for the joint propensity score.

2 Interference based on the exposure to neighborhood treatment

The paper models local network interference by representing each unit’s outcome through its own treatment and a summary of immediate neighbors’ treatments. It defines main and spillover effects under fixed or observed-distribution neighborhood exposure and establishes identification results.

  • Notation: The network is an undirected graph with units as nodes and edges representing links; each unit’s neighborhood contains its connected neighbors.
  • Notation: Covariates are decomposed into individual or contextual characteristics and neighborhood features such as degree, topology, centrality, and aggregated neighbor characteristics.
  • Interference assumption: Neighborhood interference rules out effects from neighbors’ neighbors while allowing a unit’s outcome to depend on neighbors’ treatments through a specified summary function.
  • Interference assumption: Under this framework, each unit receives an individual treatment Zi and a neighborhood treatment Gi, forming a joint treatment.
  • Causal estimands: The main effect compares potential outcomes under individual treatment levels while holding neighborhood treatment fixed, then averages across the observed neighborhood-treatment distribution for the overall effect.
  • Causal estimands: The spillover effect compares neighborhood treatment level g with 0 under fixed individual treatment z, while the overall spillover effect averages these contrasts over neighborhood-treatment exposure.
  • Causal estimands: Unlike estimands based on hypothetical interventions over the whole treatment vector, the overall estimands fix a unit’s treatment and retain neighbors’ observed treatments.
  • Identification: Theorem 1 states identification of the average dose-response function under no multiple versions of treatment, neighborhood interference, and unconfoundedness.

3 Bias when SUTVA is wrongly assumed

The paper shows that covariate-adjusted estimators assuming SUTVA can be biased when neighborhood interference is present. The bias depends on interference and on residual association between individual and neighborhood treatments, with additional unmeasured-confounding bias when unconfoundedness fails.

  • Naive estimation under interference: Under interference, estimators comparing units by individual treatment alone generally do not estimate the intended main or spillover effects.They compare potential outcomes with different neighborhood-treatment values while ignoring neighborhood exposure.
  • Conditional independence: If individual and neighborhood treatments are conditionally independent given X⋆, SUTVA-based covariate adjustment can remain unbiased for the overall main effect.This result holds even when interference is present.
  • Conditional dependence: Residual association between individual and neighborhood treatments after conditioning on X⋆ produces bias in the naive estimator.The bias depends on both the level of interference and the residual treatment association; omitted neighborhood covariates, peer influence, and homophily can contribute to that association.
  • Unmeasured confounding: When unconfoundedness fails conditional on X⋆, the naive estimator combines bias from interference with bias from unmeasured confounders.Theorem 2.B attributes the additional component to Ui = Xi\X⋆.
  • Special cases: If treatments are conditionally independent or interference is absent, the remaining bias is attributable only to unmeasured confounders.Under SUTVA, this agrees with the corresponding unmeasured-confounding expression.

4 Definition and Properties of Generalized Propensity Score Under Neighborhood Interference

The paper generalizes the propensity score to jointly represent individual treatment and neighborhood treatment under network interference. This joint score balances covariates and supports adjustment methods that can handle the two treatments jointly or separately.

  • Balancing property: Under neighborhood interference, the generalized score has balancing properties analogous to those of the ordinary propensity score under SUTVA.Units with the same score have the same covariate distribution across joint-treatment arms.
  • Joint propensity score: The joint propensity score is the probability of receiving individual treatment z and neighborhood treatment g conditional on observed covariates.It is defined as ψ(z; g; x) = P(Zi = z, Gi = g|Xi = x).
  • Conditional unconfoundedness: Conditional unconfoundedness given covariates implies conditional independence of potential outcomes from individual and neighborhood treatments given the joint propensity score.This permits imputing missing potential outcomes across arms for units with the same exposure propensity.
  • Separate propensity scores: The joint propensity score can be factorized into an individual propensity score and a neighborhood propensity score.The individual score models individual treatment, while the neighborhood score models neighborhood treatment conditional on individual treatment and covariates.
  • Separate propensity scores: Separate adjustment can use covariate subsets affecting individual treatment and neighborhood treatment, which need not be identical.Individual characteristics, neighborhood characteristics, and network structure may enter the respective assignment mechanisms.

5 Propensity Score-Based Estimator for Main Effects and Spillover Effects

The paper estimates main and spillover effects under network interference using separate individual and neighborhood propensity-score adjustments. Its semi-parametric approach factorizes the joint propensity score and combines individual-treatment subclassification with generalized propensity-score modeling.

  • The proposed estimator targets the average dose-response function and derives marginal main effects τ(g) and spillover effects δ(z; g).
  • Identification can use a joint propensity score balancing individual and neighborhood covariates across levels of the bivariate treatment.The joint score is difficult to stratify or match directly because neighborhood treatment may have many values.
  • The method factorizes the joint propensity score into individual and neighborhood propensity scores for separate adjustment.
  • Adjustment requirements differ by estimand: spillover estimation requires neighborhood-score adjustment, while main-effect estimation reverses the relevant adjustment logic.
  • The strategy is particularly appropriate when neighborhood treatment has a domain contained in R, and inference uses bootstrap resampling tailored to the sampling scheme.
  • Individual propensity-score subclassification is followed by neighborhood generalized propensity-score modeling within each subclass.The procedure estimates neighborhood propensity scores and uses them to predict potential outcomes before averaging across units.

6 Realistic Simulation Study leveraging Add Health data

Simulations calibrated to friendship networks assess bias from ignoring interference and performance of the proposed estimators. Across realistic treatment-assignment scenarios, jointly adjusting individual and neighborhood propensity scores yields unbiased estimates under the stated assumptions.

  • The study uses friendship networks from 29 Add Health schools containing 16,410 students, with simulated health-insurance treatment and illness-related school absences.The setting represents possible flu-health spillovers among students’ friends.
  • Four scenarios vary dependence between individual treatment and neighborhood treatment, using friends’ insurance proportions or treated-friend counts as neighborhood exposure.
  • Naive estimators that ignore interference have bias proportional to interference and association between individual and neighborhood treatments.
  • Main-effect simulations: In Scenario 1, subclassification-based estimators are essentially unbiased across interference levels when individual and neighborhood treatments are conditionally independent.
  • Main-effect simulations: Regression estimators show model-misspecification bias ranging from −3.50 with individual covariates to −3.25 with individual and neighborhood covariates.
  • Main-effect simulations: The proposed subclassification-and-GPS estimator shows no bias across scenarios and interference levels, while joint covariate adjustment reduces interference bias.
  • Spillover simulations: For spillover effects, estimators using the correct neighborhood-treatment function are unbiased under SUTNVA when confounding covariates are properly adjusted.
  • Spillover simulations: The estimated dose-response results indicate that more insured friends reduce average illness-related missed school days, with larger spillovers for uninsured students.The individual insurance effect decreases as friends’ insurance coverage increases.

7 Concluding Remarks

The paper formalizes treatment and spillover effects under network interference, derives bias formulas for naive estimators, and proposes propensity-score adjustment for observational network studies. It also identifies scope and inference limitations, including dependence on a known fixed network and model specification.

  • Framework: SUTNVA limits interference to immediate neighbors and represents each potential outcome using individual and neighborhood treatments, (Zi, Gi).The neighborhood treatment summarizes neighbors’ treatment assignments.
  • Bias analysis: Naive treatment-effect estimators can combine bias from unadjusted neighborhood confounders with interference bias proportional to interference and treatment association.The interference component can vanish when individual and neighborhood treatments are conditionally independent after covariate adjustment.
  • Estimation: The paper extends propensity-score balancing to joint, individual, and neighborhood treatments for estimating treatment and spillover effects.The proposed semi-parametric estimator targets both effects under an extended unconfoundedness assumption.
  • Simulations: Simulations evaluate main-effect and spillover estimation, naive-estimator bias, bias formulas, and the need for joint-propensity-score adjustment.The studies use multiple scenarios to assess finite-sample performance and different bias sources.
  • Limitations: The estimator assumes a fully known, fixed social network and may suffer residual-imbalance or model-misspecification bias in real studies.The authors are working to make the semi-parametric estimator less model-dependent.
  • Limitations: Network statistical inference remains difficult because correlated observations complicate uncertainty quantification across sampling schemes.The paper proposes unit- or cluster-level bootstrapping but notes that alternative schemes require further investigation.
  • Estimands: Conditional main effects compare individual treatments among units sharing neighborhood treatment g, while conditional spillover effects compare neighborhood treatments among controls.These estimands restrict comparisons to observed treatment and neighborhood-treatment subsets.

A.2.1 Conditional Propensity Scores of the Neighborhood Treatment

The appendix develops a conditional neighborhood propensity score for estimating spillover effects and notes that subclassification can exploit its balancing properties.

  • Propensity score: The conditional neighborhood propensity score is defined to support covariate adjustment for conditional spillover effects.It is presented with balancing and conditional-unconfoundedness properties.
  • Estimation: A subclassification-based estimator can use the conditional neighborhood propensity score.The construction follows the same general propensity-score strategy used for binary treatments on restricted subsets.

B Data Generating Model for Simulations

The simulations use the Add Health friendship network and student race and grade as covariates, while generating treatments and outcomes under four scenarios.

  • Simulation data: The simulation network comes from Add Health, with race and grade recorded for students.Race is binary in the simulation, while grade ranges from 6 through 12.
  • Simulation design: Treatment and outcome variables are generated rather than directly taken from the observed dataset.The passage states that four data-generating scenarios are used.

B.1 Correlation Structure of the Joint Treatment and Covariate Balance

The simulation scenarios vary how individual and neighborhood treatments depend on individual covariates, neighborhood covariates, degree, and one another. Covariate-balance analyses examine these induced dependence structures.

  • Scenario 1: Scenario 1 generates individual treatment from individual race and grade, while neighborhood treatment depends on neighbors’ treatment.The scenario is designed to create a treatment structure involving individual covariates and network exposure.
  • Scenario 2: Scenario 2 makes individual treatment depend on individual and friends’ race and grade, so individual and neighborhood treatments remain dependent given only individual covariates.The stated logistic model includes gradei, racei, friends.gradei, and friends.racei.
  • Scenario 3: Scenario 3 associates individual and neighborhood treatments through student degree, with Gi counting treated friends and differing across individual-treatment arms.Degree enters individual-treatment generation and also affects the number of treated friends.
  • Balance diagnostics: The simulations report covariate balance across individual and dichotomized neighborhood treatment arms using tables and weighted logistic regressions.Additional figures compare neighborhood propensity scores across race, grade, and friends’ characteristics.
  • Scenario 4: Scenario 4 directly correlates individual and neighborhood treatments through an iterative assignment procedure.Neighborhood treatment is computed as the proportion of treated friends among the first five best friends.
  • Scenario 4: Residual dependence between Zi and Gi persists after conditioning on all covariate types in Scenario 4.The passage reports this dependence in the covariate-balance results.

B.3 True Main and Spillover Effects

The section describes simulation scenarios through treatment-assignment association and true main and spillover effects, then outlines propensity-score estimation of dose-response functions.

  • Simulation scenarios: Simulation scenarios vary the correlation between individual treatment Zi and neighborhood treatment Gi, alongside true overall main and spillover effects.True effects are computed as means of simulation-specific estimates from the stated outcome models.
  • Neighborhood propensity score: Within each subclass, the method models the neighborhood propensity score conditional on neighborhood covariates and uses it to predict potential outcomes for each joint treatment level.The neighborhood-treatment model depends on whether Gi is defined as the proportion or number of treated neighbors.
  • Dose-response estimation: The average dose-response function is obtained by averaging predicted potential outcomes across units for each individual and neighborhood treatment combination.This construction supports estimation of marginal treatment and spillover quantities.
  • Individual propensity score: The estimation procedure first models individual treatment assignment and forms subclasses with similar individual propensity scores and sufficient treatment-group balance.The simulation uses five subclasses based on quintiles.
  • Outcome model: The simulation outcome model relates potential outcomes to individual treatment, neighborhood treatment, their interaction, and the estimated neighborhood propensity score.The authors note that alternative models, including a cubic polynomial, could be used in more realistic settings.

C.2 Statistical Inference and Interval Estimation

The section uses unit-level bootstrap resampling to estimate uncertainty for main and spillover effects, while identifying sampling-mechanism dependence as an important boundary of validity.

  • Bootstrap inference: The proposed inference procedure independently resamples units with replacement after computing neighborhood treatments and neighborhood covariates.These network-related quantities are carried as node attributes during resampling.
  • Simulation assessment: The reported standard errors are consistent with Monte Carlo root mean squared errors from Tables 2, 3, and 4.
  • Sampling assumptions: The bootstrap variance estimator is valid for an egocentric sampling mechanism, where sampled units carry their neighborhood treatment and covariates regardless of neighbor resampling.Its properties depend on matching the resampling mechanism to the mechanism that generated the observed sample.
  • Sampling assumptions: When the sampling mechanism is not independent, the proposed solution is not ideal; clustered bootstrap methods may be appropriate for clustered samples.For networked clusters, the authors suggest that smaller correlated clusters could be defined using community detection.

D.4 Proof of Corollary 2.A.2

The proof continues the bias derivation without assuming conditional independence between individual and neighborhood treatments, then simplifies it when treatment interaction is absent.

  • The derivation removes the assumption that Zi and Gi are conditionally independent given X⋆.
  • Without interaction between Zi and Gi, the conditional mean terms do not depend on g, yielding a simplified bias expression.

D.5 Proof of Theorem 2.B

The proof develops the bias formula by iterated expectations and treatment–unobserved-variable decompositions, then reduces it under no individual-by-neighborhood treatment interaction.

  • The proof uses iterated equations and the definition of Vg as part of the derivation.
  • It decomposes conditional outcome means using the distributions of an unobserved variable Ui under different individual and neighborhood treatment levels.
  • Adding and subtracting conditional mean terms that do not depend on g or u produces the intermediate bias decomposition.
  • When Zi and Gi have no interaction, the bias formula reduces to a simpler expression.

D.6 Proof of Corollary 2.B.1

The proof derives simplified bias expressions under no interaction between individual treatment and the relevant covariate, then establishes balancing and unconfoundedness properties for the joint propensity score. It also extends these arguments to a factorized propensity-score representation involving individual and neighborhood covariates.

  • The proof uses conditional unconfoundedness, consistency, iterated equations, and marginalization over Gi to derive the bias expression.
  • Under no interaction between individual treatment Zi and Ui, the bias formula simplifies.
  • The joint propensity score ψ(z; g; Xi) balances the joint treatment and neighborhood exposure conditional on covariates.The proof shows both conditional probabilities equal ψ(z; g; Xi), using that the score is a function of Xi and the score definition.
  • Under Assumption 3, conditioning on the potential outcome Yi(z, g) does not alter the joint treatment probability once ψ(z; g; Xi) is given.The argument applies iterated expectations and the unconfoundedness assumption to show equality with the joint propensity score.
  • The same proof strategy establishes the corresponding property for the factorized score λ(z; g; Xg_i), using its relationship to Xi and ψ(z; g; Xi).The factorization ψ(z; g; Xi) = λ(z; g; Xg_i)φ(z; Xz_i) is used in the argument.
Loading 1609.06245v4…