Source-linked AI summary

Design and analysis of experiments in networks: Reducing bias from interference

Dean Eckles, Brian Karrer, Johan Ugander

arXiv:1404.7530v2stat.MEcs.SIphysics.soc-ph

TL;DR

Interference in connected networks makes standard causal estimators biased and complicates direct comparisons of global treatment assignments. This paper studies network-correlated assignment and neighborhood-based analysis, finding that graph cluster randomization can substantially reduce bias and error under clustered networks and strong direct and peer effects.

  • Problem

    Interference makes outcomes depend on other units’ treatments and behaviors, while common experimental designs and analyses assume no interactions and can produce biased causal-effect estimators.

  • Method

    The paper formalizes experimentation in networks, proves sufficient conditions for bias reduction through design and analysis, and evaluates methods using simulations with varied network structures and social decision processes.

  • Results

    Graph cluster randomization substantially reduces bias with comparatively small variance increases when networks are substantially clustered and treatments have substantial direct and peer effects.

  • Takeaways & Limitations

    Bias reduction is generally larger through design than analysis, while neighborhood-based analysis can further reduce bias but may substantially reduce precision.

  • Takeaways & Limitations

    Combining graph cluster randomization with the FTNR estimand is not necessarily bias reducing under monotonic responses without modifying the estimand to count matching clusters.

Abstract

from arXiv · show

Estimating the effects of interventions in networks is complicated when the units are interacting, such that the outcomes for one unit may depend on the treatment assignment and behavior of many or all other units (i.e., there is interference). When most or all units are in a single connected component, it is impossible to directly experimentally compare outcomes under two or more global treatment assignments since the network can only be observed under a single assignment. Familiar formalism, experimental designs, and analysis methods assume the absence of these interactions, and result in biased estimators of causal effects of interest. While some assumptions can lead to unbiased estimators, these assumptions are generally unrealistic, and we focus this work on realistic assumptions. Thus, in this work, we evaluate methods for designing and analyzing randomized experiments that aim to reduce this bias and thereby reduce overall error. In design, we consider the ability to perform random assignment to treatments that is correlated in the network, such as through graph cluster randomization. In analysis, we consider incorporating information about the treatment assignment of network neighbors. We prove sufficient conditions for bias reduction through both design and analysis in the presence of potentially global interference. Through simulations of the entire process of experimentation in networks, we measure the performance of these methods under varied network structure and varied social behaviors, finding substantial bias and error reductions. These improvements are largest for networks with more clustering and data generating processes with both stronger direct effects of the treatment and stronger interactions between units.

1. Introduction.

Interference makes network experiments unable to directly compare global treatment assignments and can bias standard causal estimators. The paper evaluates realistic design and analysis changes, proving sufficient bias-reduction conditions and testing them through simulations.

  • Motivation: Network interference means one unit’s outcome may depend on other units’ treatments and behaviors, potentially across the entire network.This complicates estimating effects of global treatment assignments.
  • Motivation: The ATE compares outcomes when all units receive treatment with outcomes when all receive control, but each potential outcome depends on the global assignment vector.Additional assumptions are therefore required for identification.
  • Motivation: SUTVA and no-interference assumptions can identify the ATE under random assignment, but are implausible when units interact.The paper avoids replacing interference with other strong assumptions.
  • Approach: The paper formalizes network experimentation, proves sufficient conditions for bias reduction, and evaluates design and analysis methods through extensive simulations.The process treats design as assigning conditions and analysis as combining responses into causal estimators.
  • Findings: Graph cluster randomization substantially reduces bias with limited variance increases, especially in clustered networks with strong social interactions.Neighborhood-based analysis can reduce bias further but often sacrifices precision, making simple estimators preferable in error.
  • Findings: No design-analysis combination works uniformly across different network and behavioral settings, but simulations provide practical guidance for experiments with peer effects.The paper also identifies conditions under which design and analysis changes will not increase bias.

2. Model of experiments in networks.

The paper models a network experiment as a sequence of initialization, treatment assignment, outcome generation, and estimation. Each complete pass through these phases represents one experimental instance.

  • 2. Model of experiments in networks: Network experimentation consists of initialization, treatment assignment, outcome generation, and estimation.Treatment assignment represents design, while estimation represents analysis.
  • 2. Model of experiments in networks: Simulations implement the same four phases repeatedly to evaluate experimental methods across generated instances.This links the formal model directly to the simulation process.

2.1. Initialization.

Initialization covers everything before treatment assignment, including network formation and the generation of vertex characteristics and prior behaviors. It may be treated as random or conditioned on as fixed.

  • 2.1. Initialization: Initialization includes network formation and processes producing vertex characteristics and prior behaviors.The paper later generates networks from small-world models and degree-corrected blockmodels.
  • 2.1. Initialization: Researchers may average design and analysis performance over a distribution of networks or condition on one observed network and its vertex characteristics.The choice depends on whether initialization is regarded as random or fixed.
  • 2.1. Initialization: After initialization, the experiment has a fixed graph G = (V, E), adjacency matrix A, and vertex characteristics X that may affect outcome generation.The characteristics need not be related to graph structure.

2.2. Design: Treatment assignment.

The design phase maps network vertices to treatment conditions, focusing here on binary treatment assignment and graph cluster randomization. Clustered assignments increase network autocorrelation but can create positivity and variance trade-offs.

  • 2.2. Design: Treatment assignment: The treatment assignment phase maps vertices to treatment or control, with independent assignment modeled through Bernoulli treatment indicators.The treatment probability is q.
  • 2.2. Design: Treatment assignment: Graph cluster randomization partitions vertices into clusters and assigns each cluster a shared treatment.Vertex treatments inherit the assignments of their clusters.
  • 2.2. Design: Treatment assignment: Assigning entire clusters to one treatment can make some effective treatments impossible for certain vertices, violating positive assignment probabilities.A modified assignment allowing some vertices to differ from their cluster can address this issue.
  • 2.2. Design: Treatment assignment: The paper uses ϵ-net clustering in simulations, while other possible cluster mappings include community detection, observed membership, and geography.Global community detection can produce clusters that are too large, increasing variance excessively.
  • 2.2. Design: Treatment assignment: Independent random assignment is a special case of clustered assignment in which every vertex forms its own cluster.This provides a direct baseline for comparing graph-clustered designs.

2.3. Outcome generation and observation.

The paper models network outcomes as functions of global treatment assignments and stochastic processes, then studies dynamic peer effects that expand effective treatment dependence across the graph.

  • Outcome generation: Observed outcomes are generated from a global treatment assignment Z and an independent stochastic component U through Y = f(Z, U).The vertex-level response functions are obtained by decomposing this outcome-generation function.
  • Treatment response assumptions: Treatment-response assumptions define when different global assignments are equivalent for a vertex through effective-treatment functions g_i(·).ITR restricts outcomes to depend only on own treatment, while more general CTR assumptions map global assignments into equivalence classes.
  • Dynamic interference: Peer-driven dynamics can violate tractable neighborhood assumptions because behavior propagates from neighbors to neighbors over successive time steps.The paper therefore evaluates methods under outcome-generating processes consistent with such interference theories.
  • Dynamic interference: At time step t, effective treatment depends on assignments within distance t −1, quickly expanding toward the full graph.This makes limited-scope exposure models increasingly difficult to justify as the dynamic process unfolds.
  • Simulation outcome model: Simulations use a stochastic mean-neighbor-behavior model in which α is baseline, β controls direct treatment effects, and γ controls peer effects.The process is run through a maximum time T, with binary behavior implemented using a probit specification.
  • Simulation outcome model: The simulated process can be interpreted as a noisy best-response model and, for γ > 0, as a graphical game with strategic complements.This connects the outcome model to settings where vertices anticipate neighbors’ prior behavior.

2.4. Analysis and estimation.

The analysis defines estimands under network interference and establishes conditions under which graph-clustered designs or alternative effective-treatment definitions reduce bias, while highlighting variance and specification trade-offs.

  • Estimands and bias: The target ATE compares mean outcomes under global treatment assignments, but individualistic treatment-response estimands can be biased when other units’ assignments affect outcomes.Each unit’s contribution to bias depends on the difference between its observed-design expectation and its outcome under the global assignment of interest.
  • Bias reduction through design: Under a linear outcome model with monotonic responses and nonnegative interaction coefficients, some graph clusterings yield no greater absolute bias than independent assignment.The result holds for a fixed treatment probability p and an appropriate mapping of vertices to clusters.
  • Bias reduction through design: Bias reduction depends on whether clusters capture the network’s dependence structure, so clusters formed from closer or more connected vertices can be more useful than arbitrary groupings.The coefficient matrix B specifies the relevant dependence structure, and poor clustering can leave graph-clustered estimators with substantial bias.
  • Bias reduction through analysis: Alternative effective-treatment definitions compare estimands through monotonicity conditions, allowing analysis choices to reduce absolute bias under suitable response behavior.Theorem 2.2 gives a sufficient condition for one independent-assignment estimand to have no greater absolute bias than another.
  • Bias reduction through analysis: Neighborhood-based estimators can be unbiased when their effective-treatment assumptions hold, but misspecifying the functions g_i(·) creates estimand bias.The effective treatment assumption is what permits generalizing from one sampled assignment to the global treatment assignments of interest.
  • Estimator trade-offs: Combining graph cluster randomization with neighborhood-based estimands is not necessarily bias reducing unless the estimand is modified to count matching clusters.The unmodified FTNR estimand lacks this guarantee even under monotonic responses.
  • Estimator trade-offs: Estimators requiring all neighbors to share a condition can have substantially increased variance because few vertices qualify and Hájek weights become highly imbalanced.Fractional relaxation or additional modeling can borrow information from other vertices to address this variance problem.

3. Simulations.

Simulations evaluate graph-clustered assignment and neighborhood-based analysis across network structures and social-interaction strengths. Both methods can reduce bias, but design generally yields larger reductions and may improve RMSE when bias reduction is substantial.

  • Simulation setup: Simulations vary rewiring probability, direct treatment effect, peer-effect strength, and network structure to assess bias and error reduction.Small-world simulations use N = 1,000 vertices, initial degree k = 10, five rewiring probabilities, and β and γ values from 0.0 to 1.0.
  • Design: Graph cluster randomization assigns entire 3-net clusters to treatment or control and is compared with independent random assignment.The design aims to place vertices near neighbors receiving the same treatment, making observed outcomes closer to global-treatment counterfactuals.
  • Small-world results: Graph cluster randomization reduces bias most when peer and direct effects are strong relative to the baseline and the network has substantial local clustering.In small-world networks, substantial clustering corresponds to small rewiring probability prw.
  • Small-world results: Approximately 40% RMSE reduction occurs with prw = 0.01 and γ = 0.5, although clustered assignment can also increase RMSE when variance costs outweigh bias reduction.When both direct and peer effects are strong, RMSE reductions are observed across rewiring structures.
  • Analysis: Neighborhood-based exposure models further reduce bias but use fewer vertices and increase variance, making their MSE worse for many simulated parameter values.The added bias reduction is reported for both independent and clustered assignment, while the variance cost is reflected in RMSE and MSE results.
  • Stochastic blockmodel results: Degree-corrected blockmodel networks show smaller bias and error reductions than small-world networks, attributed to higher-degree vertices and less local clustering.With pcomm = 0.8, clustering remains comparable to small-world networks with prw = 0.5, and reductions are likewise comparable.

4. Discussion.

The paper evaluates design and analysis strategies for reducing bias from interference without relying on strong, often unrealistic interference assumptions. Theory and simulations show that graph cluster randomization and neighborhood-based estimators can reduce bias, with important scope limitations.

  • Prior work often assumes restricted interference patterns to make network-treatment analysis tractable and estimators unbiased or consistent.
  • The paper formalizes experimentation in networks, proves sufficient bias-reduction conditions, and evaluates methods through simulations with dynamic peer effects.
  • Graph cluster randomization substantially reduces bias with comparatively small variance increases when networks are clustered and direct and indirect treatment effects are substantial.
  • Specific estimators can provide additional bias reduction even when their effective-treatment assumptions are incorrect.
  • The simulations use monotonic outcomes without vertex characteristics beyond degree or prior behaviors, motivating evaluation on other networks and data-generating processes.

A.1. Modified graph cluster randomization: hole punching.

Hole punching modifies graph cluster randomization by adding vertex-level randomness, leaving clusters predominantly assigned to one treatment while isolating a small fraction of vertices.

  • Hole punching adds vertex-level randomness so some vertex assignments differ from their cluster assignments.
  • Independent switching variables set each vertex to its cluster treatment with probability η and otherwise flip the assignment.
  • The resulting isolated treatment positions may help estimate differences between direct treatment effects and peer effects.

A.2. Bias reduction from design: balanced linear case.

In the balanced linear case, the theory shows that suitable graph clustering can reduce absolute bias relative to balanced independent assignment while preserving treatment-control balance.

  • Balanced graph cluster randomization assigns equal numbers of clusters to treatment and control when the number of clusters is even.
  • Under the linear outcome model and monotonicity with nonnegative interaction terms, some cluster mappings guarantee no greater absolute bias than independent assignment.
  • The bias-reduction comparison is established for a fixed treatment probability p under graph cluster randomization.
  • The balanced analysis derives the true ATE and corresponding expressions under balanced graph cluster randomization.
  • For sufficiently large numbers of clusters, relative bias is identical with or without balance-preserving sampling, so reduction depends on clustering relationships to the B_ij values.

A.3. Bias reduction from analysis.

The analysis results show that more restrictive exposure conditions can reduce bias under independent assignment and, with a cluster-level reformulation, under graph cluster randomization. However, the benefit is not universal for the original vertex-level formulation.

  • The analysis compares estimators based on exposure functions, including ITR, NTR, and FNTR conditions defined by matching treatment assignments within selected neighborhoods.
  • Under independent assignment, more restrictive exposure conditions have no greater absolute bias when responses are monotonically increasing or decreasing in treatment.
  • The proof uses monotonicity to show that expected potential outcomes increase with the number of neighborhood assignments matching the global treatment vector.
  • Under graph cluster randomization, the independent-assignment proposition does not generally apply, and a counterexample shows that a more restrictive exposure condition can increase bias.
  • A cluster-level redefinition of exposure conditions yields an analogous sufficient condition covering graph cluster randomization and independent assignment as a special case.
  • The corollary specifically covers comparing FNTR and ITR under graph cluster randomization because both can be expressed as cluster-level exposure conditions.
Loading 1404.7530v2…