Source-linked AI summary
Reasoning about Interference Between Units
Jake Bowers, Mark Fredrickson, Costas Panagopoulos
TL;DR
Interference makes causal hypotheses difficult when treatment affects both treated and control units. The paper formalizes and tests theory-driven counterfactual models of spillover using a Fisherian framework. Simulations show that the method controls rejection of true hypotheses and can reject false ones across varied situations, while spillover can make effects harder to distinguish.
Problem
When treatment given to one unit changes another unit’s potential outcomes, average treatment effects may not be identified or meaningful under interference.
Method
The paper proposes Fisherian testing of specific interference models rather than requiring a no-interference assumption or relying on estimation.
Results
The method is not overly likely to reject true null hypotheses and, across many situations, has power to reject false hypotheses.
Takeaways & Limitations
Interference can be modeled and tested as a theoretically meaningful causal phenomenon rather than treated as an untestable assumption.
Takeaways & Limitations
When spillover occurs, it becomes increasingly difficult to discriminate effects, and some results may be more difficult for other datasets, networks, and models.
Abstract
from arXiv · showhide
If an experimental treatment is experienced by both treated and control group units, tests of hypotheses about causal effects may be difficult to conceptualize let alone execute. In this paper, we show how counterfactual causal models may be written and tested when theories suggest spillover or other network-based interference among experimental units. We show that the "no interference" assumption need not constrain scholars who have interesting questions about interference. We offer researchers the ability to model theories about how treatment given to some units may come to influence outcomes for other units. We further show how to test hypotheses about these causal effects, and we provide tools to enable researchers to assess the operating characteristics of their tests given their own models, designs, test statistics, and data. The conceptual and methodological framework we develop here is particularly applicable to social networks, but may be usefully deployed whenever a researcher wonders about interference between units. Interference between units need not be an untestable assumption; instead, interference is an opportunity to ask meaningful questions about theoretically interesting phenomena.
1 Introduction
The paper develops a theory-driven framework for modeling and testing causal effects when treatment spills over between experimental units. It uses Fisherian testing to examine direct and indirect effects without requiring average treatment effects or a no-interference assumption.
- The framework formalizes explicit spillover between units and connects theoretically driven models to potential outcomes, hypotheses, and tests.Researchers can represent how subjects react to treatment received by their neighbors.
- Simulation results indicate that the method does not reject true hypotheses or fail to reject false hypotheses too often, for hypotheses with or without spillover.The simulations use treatment spillover from treated to control units through a known network.
- The approach models causal-effect flows across networks and tests model parameters without requiring probability models of observed outcomes.
- The paper treats no interference as an implication of using simple averages, not a fundamental assumption, and frames interference as directly testable.The framework supports substantively relevant, theoretically motivated hypotheses about interference and causal effects.
- Fisherian testing focuses on hypotheses rather than estimation, including theory-driven statements about direct and indirect effects.The framework does not require conceptualizing causal effects as averages.
- Design choices involving the model, data distribution, sample size, network density, and treated-unit share can affect experimental efficiency, sometimes counter-intuitively.
2 Setting
The paper studies randomized treatment assignments in pre-existing networks, where outcomes may depend on other units’ treatment assignments. It formalizes interference through network structure and counterfactual potential outcomes rather than treating no interference as mandatory.
- Network setting: The setting is a randomized experiment in a pre-existing network, with treatment allocated to units connected by observed or hypothesized relationships.The network may be fixed, like a road network, or represent theorized possible connections.
- Simulated experiment: The simulated testbed contains 256 subjects connected by 512 edges, with treatment randomly assigned to 128 units.The example uses equal-sized treated and control groups.
- Network representation: An adjacency matrix S records network relationships, while graphical representations show the same links among treated and control units.For undirected networks, a link from i to j implies a reciprocal link from j to i.
- Potential outcomes: Potential outcomes are indexed by the full treatment-assignment vector, allowing unit i’s outcome to depend on treatments assigned to other units.Interference occurs when one unit’s potential outcomes depend on another unit’s treatment assignment.
- Modeling interference: The paper argues that no interference is an implication of a modeling choice, not a requirement of the potential-outcomes framework.Writing outcomes only as functions of unit i’s own assignment excludes other units’ treatment statuses by construction.
3 Method: Hypotheses and Models
The method defines causal models that map potential outcomes under one treatment assignment to outcomes under another, including sharp null models and interference. It then tests these models with randomization distributions and a Kolmogorov-Smirnov statistic.
- Causal models: A causal model H transforms potential outcomes under one assignment into potential outcomes under another, with parameters encoding the model’s causal effects.The framework permits hypotheses about outcomes generated by complex interference patterns.
- Uniformity trial: The uniformity trial y0 represents the potential outcomes when every unit receives control, providing a baseline for sharp-null testing.Depending on the application, it can represent a world without an experiment or one using an established treatment.
- Sharp null: Under the sharp null, treated and control outcomes should be random samples from a common distribution, so their observed distributions should differ only through sampling noise.Systematic distributional differences indicate that the no-effects model does not fit the observed experiment.
- Test statistic: The Kolmogorov-Smirnov statistic compares treated and control empirical distributions and is sensitive to differences in center, spread, skew, and other distributional features.The statistic is small for similar distributions and large for dissimilar ones.
- Randomization inference: The Fisherian procedure evaluates the test statistic over every possible treatment assignment to form its null randomization distribution.The causal model links observed outcomes to the uniformity trial, while the design determines the assignment distribution.
- Randomization test: 6.83 × 10^-14: the simulated experiment’s sharp-null p-value, providing very strong evidence against the no-effects model.The p-value is obtained from the test-statistic distribution under the model H(yz, 0) = yz ≡ y0.
4 Models of interference
The paper formulates interference as a theoretically driven causal model that transforms baseline potential outcomes into outcomes under treatment. It then tests hypotheses about direct effects and spillover parameters using randomization-based inference.
- 4 Models of interference: The framework writes testable causal models in which treatment can affect treated units directly and control units through network spillovers.The model represents network influence through the number of treated neighbors and includes direct-effect and spillover parameters.
- 4 Models of interference: The network is represented by an adjacency matrix, with treated-neighbor exposure summarized as the scalar z^TS.This reduction makes the spillover component depend on the number of treated neighbors for each unit.
- 4 Models of interference: The model multiplies treated units’ outcomes by β, while control-unit spillovers increase toward β as treated-neighbor exposure grows at a rate governed by τ.The growth curve remains between 1 and β, with τ controlling how quickly spillover increases.
- 4 Models of interference: Special parameter values nest simpler models: τ = 0 removes spillovers, while β = 1 yields the sharp null of no effects regardless of τ.These nested cases connect the interference model to familiar no-spillover and no-effect hypotheses.
- 4 Models of interference: The approach treats spillover parameters as substantively meaningful quantities that can be tested jointly with causal-effect parameters rather than merely as nuisance parameters.Joint p-value regions can provide evidence about both causal effects and structural features of interference.
- 4 Models of interference: For hypothesized β and τ, the method adjusts observed outcomes back to a uniformity trial and evaluates a test statistic over random treatment assignments.The resulting p-value measures how consistent the adjusted data are with the proposed causal model.
5 Operating Characteristics of Tests for Models of Interference
The paper evaluates whether its interference tests control false rejections and detect false hypotheses under researcher-specified models, designs, and data. Simulations show controlled size but complex power patterns shaped by sample size, network structure, and treatment allocation.
- 5 Operating Characteristics of Tests for Models of Interference: The tests should rarely reject true hypotheses while producing small p-values and high rejection rates for false hypotheses.These are the paper’s minimum criteria for controlled error and useful power.
- 5 Operating Characteristics of Tests for Models of Interference: The simulation framework lets researchers assess operating characteristics for their own models, experimental designs, test statistics, and data.The paper presents simulations as an applied complement to analytic asymptotic results whose relevance may vary across settings.
- 5 Operating Characteristics of Tests for Models of Interference: The simulations show that the testing method maintains its promised error rate when evaluating true hypotheses.The assessment compares test size with the prespecified level and finds that the method keeps its promises.
- 5 Operating Characteristics of Tests for Models of Interference: Increasing sample size increases power, but network density and the percentage treated can have non-monotonic relationships with power.These patterns arise from interactions between the model and the fixed network.
5.1 Simulation setup
The simulations vary network size, density, treatment assignment, baseline outcomes, model parameters, and test statistics to evaluate inference under interference. Across the reported settings, size remains controlled, while power depends on sample size, network structure, and the parameter being tested.
- 5.1 Simulation setup: The canonical experiment contains 256 subjects and 512 network edges, with additional networks varying sample size and edge count.The simulated networks use a fixed unit order and grid locations, placing edges between closest pairs first.
- 5.1 Simulation setup: All three test statistics maintain appropriate size, rejecting false nulls no more than 5% of the time when the null is true.The comparison includes mean difference, Mann-Whitney U, and Kolmogorov-Smirnov statistics.
- 5.1 Simulation setup: The three statistics have similar power as β and τ move away from their true values.The statistics represent mean differences, rank-based location comparisons, and general distributional divergence.
- 5.1 Simulation setup: With 32 units, roughly a fourfold change in β is needed for rejection at α = 0.05, while almost no tested τ value is rejected at that level.Larger samples enable rejection of more false spillover hypotheses.
- 5.1 Simulation setup: Power increases with sample size, but the direct-effect parameter can require less additional sampling than the spillover parameter.For the reported setup, a sample one quarter as large provides nearly as much power for β, whereas larger samples can matter more for τ.
- 5.1 Simulation setup: For the reported network and model, the most powerful design randomly assigns half of the units to treatment, while network density can diminish β power and increase τ power.The simulations identify a minimum density around 1 for maintaining power against reasonable spillover hypotheses in that dataset.
- 5.1 Simulation setup: Large spillover can mask direct effects, making hypotheses with simultaneously large β and τ observationally similar to the truth.The joint rejection surface therefore contains regions where false hypotheses are difficult to distinguish.
5.2 Comparing Models
The simulations assess whether the proposed tests protect true hypotheses, reject false ones, and distinguish competing interference models. They also show that spillover and misspecified functional forms can make different models observationally similar.
- Comparing models: Spillover makes competing models harder to discriminate because different models can imply similar adjustments and treated-control distributions.The low-power hypothesis makes the treated and control groups appear similar, while spillover increases the difficulty of distinguishing models.
- Simulation properties: The method maintains test size and power when the tested model is true, protecting the sharp null and rejecting false hypotheses.Under the sharp null, hypotheses including β = 1 were rejected in 3.3% of simulations at α = 0.05.
- No spillover: When β = 2 and τ = 0, the method rarely rejects the true no-spillover model and seldom suggests substantial spillover.The simulations indicate researchers would be unlikely to infer large spillover when none exists.
- Misspecified models: 71.4% of additive-model simulations rejected the least-rejected multiplicative hypothesis, indicating strong rejection of incorrectly specified models.The least-rejected hypothesis was β = 2 and τ = 0.395, rejected in 71.4% of simulations.
- Misspecified models: 10.8% of spillover-model simulations rejected the least-rejected additive hypothesis, showing that incorrect functional forms are not always rejected.The closest additive approximation was α = 48.718.
5.3 Summary of Simulation Studies
The simulation studies indicate that the proposed procedure satisfies standard requirements for statistical tests across varied designs and information sources. They also show why graphical diagnostics and study-specific simulations remain useful.
- Simulation conclusions: The method controls test size and has power against false parameter sets across the simulation studies.Power depends on relevant information such as sample size, network density, baseline outcomes, network relationships, and treatment proportion.
- Model assessment: Data may remain consistent with more than one model even when power increases, so graphical methods help compare model-implied adjustments.The studies also inform experimental design and interactions among models, designs, and data.
6 Discussion
The discussion frames interference as a setting for theory-driven causal questions rather than an assumption that must be imposed away. The proposed framework specifies and tests unit-level models while identifying computational, inferential, and diagnostic boundaries.
- Contribution: The framework directly specifies and tests theorized forms of interference instead of limiting analysis to no-effects hypotheses.It focuses on unit-level causal models and uses hypothesis tests to evaluate them.
- Computation: For small samples, exact p-values can be obtained by enumeration, while large-sample approximations can speed computation when appropriate.Simulation-based assessment can reveal problems with approximations, such as failure to control Type I error.
- Theory and inference: The approach allows substantive theory to guide model selection before statistical technique, including theories involving spillover between units.Researchers translate theoretical implications into models for experimental subjects.
- Relation to effect estimation: The framework complements average-effect analyses by offering unit-level tests of the processes that may generate direct and indirect effects.The authors suggest workflows combining average-effect reporting, weak-null tests, and unit-level model assessment.
- Open problems: The framework remains limited because models can imply similar data adjustments, and describing relationships among multi-parameter models remains an open problem.Additional work is needed to categorize, describe, diagnose, and summarize such models.
Appendix A Code Examples
The appendix illustrates how the simulations were implemented and how researchers can use code to explore designs and models.
- Implementation: The simulations use RItools and demonstrate how applied researchers can simulate outcomes while varying sample sizes, design elements, and models.The complete simulation generation process is available in the paper’s source code.
Appendix A.1 Model
The model is parameterized by a network adjacency matrix S and returns a UniformityModel with forward and inverse mappings between observed outcomes and uniformity-trial data. Its growthCurve function specifies how spillover depends on β, τ, and network exposure.
- Model construction: The model maker takes adjacency matrix S and returns a UniformityModel parameterized by β and τ.The object combines mappings from observed data to uniformity-trial data and back to model-generated observations.
- Model construction: The first function maps observed outcomes to uniformity-trial data, while the second generates observed outcomes from that trial.Here, y0 denotes the uniformity trial, y the observed outcome, and z the treatment assignment.
- Spillover function: The spillover mechanism is represented by growthCurve(β, τ, x) = β + (1 − β) * exp(−τ^2 * x).Network exposure is computed from treatment assignment z and adjacency matrix S as zS.
- Outcome generation: The implementation uses treatment assignment, network exposure zS, and model parameters to transform uniformity-trial outcomes into observed outcomes.The supplied code includes separate treatment and non-treatment components in this transformation.
Appendix A.2 Sample Size Simulation
The sample-size simulation evaluates inference under increasing network-study sample sizes by generating data under a known interference model and repeatedly testing parameter values. Its outputs support assessment of type I error and power across the tested parameter space.
- Simulation design: For each sample size, the procedure creates uniformity-trial data, draws 1000 design-consistent treatment assignments, and generates 1000 observed datasets under the true model.The model is instantiated for the simulated network before outcomes are generated.
- Inference procedure: RItest searches parameter values around the truth and reports a p-value for each tested value.The implementation can search parameters separately or evaluate a two-dimensional grid of β and τ values.
- Simulation outputs: The simulation summarizes how often true values are rejected as type I error and how often false values are rejected as power.Separate result and power objects are produced for τ and β searches.
- Simulation design: The simulation evaluates sample sizes 32, 256, and 1024 while fixing the number of network edges through a density-based rule.For each sample size, baseline data and an adjacency matrix are generated before inference is run.
Appendix A.3 Testing Hypotheses from a Model
The appendix illustrates testing joint hypotheses generated by a specified interference model using a canonical network dataset. It constructs model-consistent outcomes and applies RItest across a parameter search space.
- Canonical example: The example creates a canonical dataset with 256 units and 512 edges, then draws one treatment assignment from its sampler.The canonical model is built from the dataset’s adjacency matrix.
- Canonical example: Observed outcomes are generated by applying the canonical model to uniformity-trial outcomes using the true β and τ values.This produces data consistent with the specified interference model.
- Hypothesis testing: RItest evaluates joint hypotheses from the model by using the canonical outcomes, treatment assignment, KS test statistic, and parameter search space.The test is run with the model supplied as the margin-of-error function and uses the asymptotic option.