Source-linked AI summary
Spillover Effects under Network Interference When Neighbours' Treatment Effects Are Heterogeneous
Faezeh Dehghan Tarzjani, Bhaskar Krishnamachari
TL;DR
The paper asks how budget-constrained network interventions can predict spillovers when neighbours differ in their own treatment responsiveness. It proves that separable summaries and covariate-only estimates are insufficient, then introduces SpilloverNet and a pilot-based per-unit response measurement. The pilot escapes the covariate-only error floor and improves budgeted targeting, while the empirical evidence remains model-generated.
Problem
Network spillover depends on individual neighbours’ treatment responsiveness, but standard network data do not observe that personal variation.
Method
The paper proves non-separability and an irreducible covariate-only error floor, then introduces SpilloverNet with per-neighbour gating and pilot-based response estimates.
Results
The CATE plug-in reaches 51.9% error, while direct-response measurements improve targeting by 3 to 14 points over degree-only policies.
Takeaways & Limitations
A small pilot measuring an intermediate response before spillover arrives can provide information unavailable from standard network data and improve budgeted targeting.
Takeaways & Limitations
All outcomes, including real-topology experiments, are generated by the authors’ model, and observed responsiveness on a network remains unavailable.
Abstract
from arXiv · showhide
Optimizing budget-constrained network interventions requires evaluating not just who is connected, but predicting how strongly individual recipients will propagate the treatment's benefits. Existing models predict spillover from neighbors' treatments and attributes. We argue that spillover also depends on how strongly each neighbor responded to its own treatment. We prove that models summarizing neighbor treatments and attributes independently cannot capture this interaction, and we introduce SpilloverNet, a graph neural network designed to preserve neighbor-level response dynamics. Because a neighbor's response is unobserved, a natural approach is to estimate it from covariates and plug it in. However, we prove that any predictor relying solely on standard network data faces an irreducible error floor set by unobserved personal responsiveness. Empirically, the plug-in's error climbs to 51.9% as heterogeneity grows - worse than using no responsiveness estimate at all - while its overall correlation still looks acceptable. A per-unit estimate from a direct-response measurement, collected in a small pilot before spillover arrives, escapes this bound and recovers up to 14 percentage points of oracle-optimal welfare in a budgeted targeting problem. On two real social graphs, SpilloverNet reaches 7.7-8.4% error, outperforming standard GNNs as well as specialized causal-representation baselines.
1 The problem
Network interventions must predict how much benefit treated neighbours transmit, not merely which neighbours receive treatment. Because individual responsiveness is unobserved, the paper proposes measuring it in a pilot before spillover arrives.
- Motivation: Spillover prediction guides scarce-treatment allocation because planners optimize direct effects plus benefits received from treated neighbours.Prediction error can create policy regret under a budget constraint.
- Motivation: Strongly responsive treated neighbours can transmit benefits, whereas unresponsive neighbours become dead ends despite identical treatment placement.Counting treated neighbours can therefore rank transmitting and non-transmitting hubs alike.
- Motivation: Individual sender response is unobserved when allocation is made, and estimating it from covariates does not work.The paper instead uses a pilot to record an intermediate outcome before spillover arrives.
- Contributions: SpilloverNet accepts per-unit response estimates and targeting on the pilot estimate recovers up to 14 points of oracle-optimal welfare over a degree rule.The model uses a per-neighbour gate to preserve sender-level response information.
2 Why existing models cannot express this
Existing models summarize neighbour treatments and attributes separately, but heterogeneous sender responsiveness makes spillover depend on their interaction. The paper formalizes this mismatch and proves that such separable models cannot represent the mechanism.
- Setup: Responsiveness is a fixed trait revealed only under treatment, while treatment is randomized given the network and covariates may explain only part of it.The outcome model includes direct treatment response, spillover, and noise.
- The gap: Existing models predict spillover from separate summaries of neighbours’ treatments and attributes, with each summary invariant to neighbour relabelling.The attribute summary includes covariates and standardized responsiveness.
- The gap: The model above is not separable for any choice of treatment summary, attribute summary, or combining function.This is the paper’s non-separability proposition.
- Proof: A separable model gives the same prediction when the stronger or weaker of two neighbours is treated first, although the true spillovers differ.Because every reviewed model is separable, the limitation is structural rather than estimator-specific.
3 Any covariate-based predictor has irreducible error
Standard network data cannot reveal unobserved personal responsiveness, creating an irreducible spillover-prediction error floor for covariate-based predictors. The floor worsens with unexplained responsiveness variation and treated degree.
- Information limit: Any predictor using only (X, Z, G) has irreducible spillover error when covariates leave individual responsiveness unexplained.A CATE plug-in is subject to the same restriction because it estimates E(τ | X) and removes the unexplained component.
- Evaluation: Table 1 reports the Floor as conditional-median NMAE and No-τ as SpilloverNet with the response input zeroed.The table uses Monte Carlo resampling and reports correlation r between estimated and true τ.
- Information limit: Substituting any estimate of E(τ | X) for τ leaves spillover error at least V⋆, regardless of sample size or estimator.The bound applies to covariate-based responsiveness estimates generally.
- Implications: The error floor grows with the unexplained responsiveness share and treated degree, concentrating difficulty in units whose covariates predict responsiveness worst.The paper computes V⋆ by resampling unexplained responsiveness while holding standard network information fixed.
4 SpilloverNet: a graph neural network with a normalised gate
SpilloverNet preserves neighbour-level response dynamics by separating each node’s own state from incoming messages and assigning neighbours learned, normalized weights. This targets the heterogeneous-neighbour mechanism that standard aggregators cannot express.
- Architecture: After L message-passing layers, each node summarizes its L-hop neighbourhood; two layers suffice for the paper’s two-hop spillover function, although four fit better empirically.No fixed-point condition is needed because the target is a fixed two-hop function.
- Architecture: SpilloverNet separates a node’s own state from aggregated neighbour messages and gives each neighbour its own learned weight.This decoupling keeps self-information from blending into incoming spillover messages.
- Normalized gate: The normalized gate converts neighbour scores into weights and constrains their total contribution to less than one at every node.The learned scalar ρℓ∈(0, 1) supplies the normalization constraint.
- Ablation: Decoupling the self path improves NMAE by 14–17 points over graph attention, while ρℓ< 1 adds 3.1 points at Flickr’s mean degree of 63.The model is trained on spillover with node inputs including covariates, degree, treatment, and a feasible response estimate.
5 Experiments
Experiments show that covariate-only responsiveness estimates face an irreducible error floor, whereas direct-response measurements improve prediction and budget-constrained targeting. On real topologies, comparisons use generated outcomes, with SpilloverNet outperforming causal-representation baselines under random treatment.
- Experimental protocol: The synthetic protocol uses 500 graphs across three random-graph families, 10% seeding, three diffusion rounds, and ση values from 0.1 to 1.5.Across the sweep, ση/στ ranges from 0.30 to 0.98.
- Responsiveness estimation: At ση = 1.5, the CATE plug-in reaches 51.9% error, worse than using no responsiveness estimate at all.The covariate-only bound is 41.1%, while SpilloverNet with the response input zeroed reaches 45.6%.
- Responsiveness estimation: A direct-response measurement beats the CATE plug-in by 4 to 33 points and reaches 19.1% versus the 41.1% floor for ση ≥0.6.It is recorded after treatment but before spillover, preserving individual responsiveness without containing spillover.
- Budgeted targeting: Direct-response targeting holds regret near 30%, while degree-only regret rises from 33% to 47%, yielding 3 to 14 points of improvement.At ση = 1.5, direct-response targeting selects units with mean true τ of 1.49 without losing reach.
- Real-graph evaluation: On Flickr and BlogCatalog, the comparison uses generated outcomes on fixed real topologies over 25 runs per model.The real-graph ordering is the same on both, and causal-representation baselines trail a graph-free MLP because treatment is random.
- Limitations: The authors identify generated outcomes as the main limitation because no network dataset with observed responsiveness is available.A variant using neighbours’ observed outcomes is excluded because it fails a network-form reflection-problem scrambling test.
A Generative model: full specification
The generative model combines covariate-driven outcomes, heterogeneous individual responsiveness, neighbour spillover, and two-hop effects on a network. A separability proof shows that independent summaries of treatments and attributes cannot represent sender-level response interactions, while the resulting spillover is exactly two-hop localized.
- Generative specification: The edge coefficient combines baseline, spatial distance, clustering, and common-neighbour terms using fixed coefficients.The specified coefficients are (cb, cs, cc, cn) = (0.02, 0.50, 0.20, 0.35).
- Generative specification: The model uses a treated-neighbour threshold factor, a two-hop spillover term, and diffusion probability proportional to treated-neighbour share.The threshold factor is 1.3 when the treated fraction exceeds one half and 1 otherwise; diffusion probability is 0.30 times treated-neighbour share.
- Non-separability proof: Independent permutation-invariant summaries of neighbour treatments and attributes assign identical predictions to configurations that swap which neighbour is treated.With heterogeneous responses, the true spillovers differ, yielding a contradiction for separable models.
- Locality: If assignments agree on node i’s two-hop ball, the model gives identical Yi, so a two-layer message-passing network is correctly specified for conditional spillover prediction.The proof traces spillover dependence to treated one-hop neighbours, their responses, and two-hop influencers; τiZi supplies the remaining assignment dependence.
- Contraction property: A weight-tied message-passing iteration has a unique fixed point with geometric convergence when κ = maxℓρℓ∥Wℓ∥ < 1.With a self-term Uℓh, the contraction modulus becomes ρℓ∥Wℓ∥ + ∥Uℓ∥.
B.4 Theorem B.4 (co-estimation contraction) and proof
The co-estimation update is a contraction when its Lipschitz modulus κτ is below one, yielding geometric convergence to a unique fixed point. With saturation, this guarantee requires excluding cases where a unit has no treated neighbour.
- Contraction condition: The update map Φ is Lipschitz in the sup norm with modulus κτ determined by treatment-effect sensitivity, network influence, and treated-neighbour counts.The modulus scales with (ατ/στ)θmax and the maximum treated-neighbour fraction d_j^(1)/(d_j+1).
- Contraction condition: If κτ < 1, repeated co-estimation converges geometrically to a unique fixed point.This follows by applying the Banach fixed-point theorem to the contraction map.
- Residual update: The residual update inherits the neighbour-effect bound after pooling over d_j+1 units, with the inner influence sum restricted to treated neighbours.The pooling contributes a factor 1/(d_j+1).
- Saturation caveat: Under saturation, ψ′(R_i)=1/(2|R_i|^1/2) becomes unbounded when R_i approaches zero, which occurs when unit i has no treated neighbour.A finite modulus can be recovered only on the restricted event min_i|R_i|≥ς0.
B.5 Proposition B.5 (error propagation) and proof
The appendix analyzes error propagation and representational scope for alternative iterative and message-passing constructions. It distinguishes the fixed-point route from the supervised SpilloverNet model and records empirical stability-certificate results.
- Model scope: The supervised two-hop-ball model does not require a fixed-point condition at any depth, unlike the iterative co-estimation alternative.The appendix states that the iterative theorems price an alternative route rather than describe the shipped model.
- Error propagation: A covariate-based responsiveness predictor remains measurable from standard network information and therefore cannot recover the unobserved individual component.The proof identifies E(τj | Xj) as I-measurable and notes that it carries none of ηj.
- Representation: A two-layer injective message-passing network can uniformly approximate spillover when the target is a permutation-invariant function of the two-hop neighborhood.The construction relies on injective aggregation and universal update maps.
- Empirical check: On the constant-responsiveness generator, SpilloverNet and Mean were indistinguishable, with NMAE 1.33% versus 1.85% and p = 0.41.The passage attributes this null comparison to the absence of variation for a neighbour-specific gate to exploit.
- Stability: κ < 1 certifies geometric convergence to a unique fixed point for the weight-tied iteration, but trained checkpoints in the reported full and reduced-scale runs were not certified.The full-grid checkpoints had κ in [1.61, 1.90], while the reduced-scale run had κ in [1.39, 1.56].
H Experimental details
The experiments use synthetic and real social graphs, controlled diffusion, supervised training, and held-out evaluations of error floors and targeting. Ablations compare response information, gating, and several GNN and causal-representation baselines.
- Synthetic graphs: 500 synthetic graphs of 100 nodes span Erdős–Rényi, Barabási–Albert, and Watts–Strogatz families with 10% random seeding and three diffusion rounds.Covariates include two node features and degree; untreated adoption probability is 0.30 times treated-neighbour share.
- Training and baselines: SpilloverNet, GraphSAGE, GATv2 variants, NetDeconf, and NetEst are trained with AdamW for 200 epochs and selected by validation NMAE.SpilloverNet and comparison GNNs use four hidden layers of width 128; GATv2 uses four heads.
- Response estimators: CATE plug-in estimates come from an R-learner, while direct response uses a mid-period outcome measurement and the pre-period method uses change relative to untreated units.Theorem 1 establishes the CATE error floor as method-independent; the direct-response construction uses independently drawn noise.
- Real graphs: Real-graph evaluations use Flickr and BlogCatalog topologies with principal-component node attributes, standardised log-degree, and clustering coefficient.Outcomes use the supplementary model with 10% seeding and the same diffusion process, across five realisations and five training seeds per model.
- Evaluation: Error floors are evaluated on 50 held-out graphs with 300 fixed-(X, Z, G) resamples, while targeting uses 50 graphs and a 10% seeded pilot.All runs used a single Google Colab GPU, with results and checkpoints persisted to Google Drive.
- Ablations: Ablation results report SpilloverNet NMAE with true τ, shuffled τ, and the gate removed, using means and standard deviations over five seeds.The seed-variability reporting for Table 1 likewise uses means and standard deviations over five seeds.
I Predictions fixed before the runs
The experiments test preregistered directional predictions about direct-response estimates, model comparisons, and targeting regret. The predictions are largely held, with specific exceptions at low heterogeneity and on selected graph settings.
- Response estimates: Direct-response correlation rises with ση from 0.24 to 0.39, and exceeds pre-period correlation at every ση.Direct-response NMAE is no higher than pre-period NMAE at every ση, with ties within seed standard deviation at 0.1 and 0.3.
- Model comparisons: GATv2+self stays within two points of GraphSAGE on synthetic graphs.This prediction is marked held in the supplied results.
- Model comparisons: GATv2+self trails SpilloverNet on Flickr when ρℓ<1 matters at high degree.The supplied prediction table marks this comparison as exceeded rather than held.
- Prediction record: The reported targeting and model-comparison predictions include a 3.1-point result in the supplied prediction record.The record labels this result as held and associates it with the relevant supplementary comparison.
- Targeting: CATE targeting regret is at least degree-only regret for every ση, held for ση ≥0.3 and reversed at 0.1.SpilloverNet is approximately equal to Mean on Zhang et al.’s constant-τ generator.