Source-linked AI summary

Overlap in Observational Studies with High-Dimensional Covariates

Alexander D'Amour, Peng Ding, Avi Feller, Lihua Lei, Jasjeet Sekhon

arXiv:1711.02582v4math.ST

TL;DR

The paper studies how overlap constrains causal inference with high-dimensional covariates, where adding covariates may aid unconfoundedness but make overlap harder. It uses an information-theoretic likelihood-ratio framework to derive bounds, showing that strict overlap imposes fixed information and population-level restrictions as dimension grows. These restrictions remain relevant alongside regularization and motivate careful consideration of overlap when adjusting for rich covariates.

  • Problem

    High-dimensional causal analyses face a tension between making unconfoundedness more plausible with richer covariates and satisfying overlap.

  • Method

    The paper reframes strict overlap as a likelihood-ratio bound and applies an analytical framework for covariate sequences to derive imbalance restrictions.

  • Results

    Strict overlap keeps the information distinguishing treated and control covariate distributions fixed as dimension grows, imposing binding population-level restrictions.

  • Takeaways & Limitations

    Overlap assumptions should be carefully considered when adjusting for rich covariates, because regularization does not avoid the resulting restrictions.

Abstract

from arXiv · show

Estimating causal effects under exogeneity hinges on two key assumptions: unconfoundedness and overlap. Researchers often argue that unconfoundedness is more plausible when more covariates are included in the analysis. Less discussed is the fact that covariate overlap is more difficult to satisfy in this setting. In this paper, we explore the implications of overlap in observational studies with high-dimensional covariates and formalize curse-of-dimensionality argument, suggesting that these assumptions are stronger than investigators likely realize. Our key innovation is to explore how strict overlap restricts global discrepancies between the covariate distributions in the treated and control populations. Exploiting results from information theory, we derive explicit bounds on the average imbalance in covariate means under strict overlap and show that these bounds become more restrictive as the dimension grows large. We discuss how these implications interact with assumptions and procedures commonly deployed in observational causal inference, including sparsity and trimming.

1 Introduction

High-dimensional covariates may make unconfoundedness more plausible while making overlap harder to satisfy. The paper formalizes this tension and shows that strict overlap increasingly restricts treated-control covariate differences as dimension grows.

  • 1 Introduction: Richer covariate sets may measure previously unmeasured confounding, but can also more closely predict treatment assignment for some subgroups.This creates opposing implications for unconfoundedness and overlap.
  • 1 Introduction: Strict overlap requires the propensity score to remain bounded away from zero and one with probability one.The paper treats this as a local assignment-probability constraint with broader distributional implications.
  • 1 Introduction: Strict overlap imposes global restrictions on discrepancies between treated and control covariate distributions by bounding a likelihood ratio.The analysis reframes overlap using a likelihood-ratio problem studied in information theory.
  • 1 Introduction: Explicit bounds on covariate imbalance become more restrictive as covariate dimension grows.The bounds concern various imbalance measures, including average discrepancies in covariate means.
  • 1 Introduction: As dimension grows, strict overlap implies that covariates must be highly correlated or their average mean differences must become arbitrarily small.This is the paper's central high-dimensional overlap implication.
  • 1 Introduction: The paper examines how these implications interact with modeling assumptions and trimming, while noting that machine-learning adjustment methods are highly sensitive to poor overlap.It situates the results within observational causal inference methods using rich covariates.

2 Preliminaries

The paper frames causal-effect estimation around unconfoundedness and overlap, emphasizing that strict overlap is needed for stronger identification and uniform inference results.

  • Researchers commonly condition on large covariate sets to make unconfoundedness plausible.
  • Population overlap requires propensity scores strictly between zero and one almost surely.The propensity score is e(X1:p) = P(T = 1 | X1:p).
  • Strict overlap strengthens population overlap by imposing a common bound η ≤ e(X1:p) ≤ 1 − η.
  • Strict overlap is necessary for uniformly n^1/2-consistent regular semiparametric estimators of the average treatment effect in a nonparametric model family.This necessity may fail under additional outcome restrictions, although relaxing strict overlap leads to non-standard asymptotics and difficult uniform inference.
  • Without strict overlap, uniform inference can be difficult or impossible, and standard bootstrap or pivotal inference may be asymptotically invalid.The appropriate inferential procedure can depend on the unknown data-generating truth.

3 Implications of Strict Overlap

The paper reframes strict overlap as a bound on treated-control likelihood ratios and derives restrictions on distributional discrepancies that become stronger as covariate dimension grows.

  • Strict overlap restricts the overall discrepancy between treated and control covariate distributions.
  • Strict overlap is equivalent to bounding the likelihood ratio between the treated and control covariate measures.Information-theoretic results then bound divergences measuring distributional discrepancy.
  • 3.1 Framework: For one-dimensional Gaussian covariates, strict overlap implies identical treated and control covariate distributions.If the distributions differ, the log-density ratio diverges in the tails, making propensity scores arbitrarily close to zero or one.
  • 3.2 Strict Overlap Implies Bounded Mean Discrepancy: When the smaller covariance operator norm grows more slowly than p, strict overlap forces average covariate-mean discrepancy toward zero.If the bound remains non-zero, both operator norms must grow at the same rate as p, implying strong covariance concentration in a finite-dimensional subspace.
  • 3.2 Strict Overlap Implies Bounded Mean Discrepancy: Strict overlap also bounds general square-integrable functional discrepancies between treated and control covariate distributions.
  • 3.3 Strict Overlap Restricts General Distinguishability: In the large-p limit, strict overlap implies no consistent classifier can distinguish the treated and control covariate distributions.The result follows from an upper bound on arbitrary classifier accuracy.
  • 3.3 Strict Overlap Restricts General Distinguishability: As p grows, strict overlap implies that average unique discriminating information in each covariate converges to zero.The conditional distributions of each covariate given preceding covariates become arbitrarily close to balance on average.

4 Strict Overlap and Modeling Assumptions

The paper examines how modeling assumptions and trimming can weaken or manage the implications of strict overlap, while showing that high-dimensional trimming may discard much of the sample and alter the estimand.

  • Treatment models: Sparse or latent-variable propensity-score models trade stronger assignment-model assumptions for weaker implications of strict overlap.
  • Treatment models: Sufficient summaries of covariates can make overlap in the summary sufficient for overlap in the full covariate set.The paper studies balancing-score specifications, including sparse and latent-variable propensity-score models.
  • Treatment models: If the propensity score is a function of a lower-dimensional covariate subset, strict overlap in that subset implies strict overlap in the full set.
  • Outcome models: Constant treatment effects can reduce the overlap requirement for estimating the average treatment effect to strict overlap with positive probability.This structural assumption also supports extrapolation and some trimming strategies.
  • Trimming: The paper’s trimming discussion connects retained-sample proportions to the accuracy of the Bayes optimal treatment classifier.
  • Trimming: Trimming changes the estimand unless additional structure, such as a constant treatment effect, is imposed.
  • Trimming: In high dimensions, trimming may need to discard a large proportion of units to achieve desirable overlap in the target population.This is especially relevant when small imbalances accumulate across many dimensions.

5 Discussion

The paper shows that strict overlap imposes strong, population-level restrictions in high-dimensional settings, with information distinguishing treated and control covariates remaining fixed as dimension grows. These restrictions matter for adjusting for rich covariates and are not removed by regularization.

  • As covariate dimension grows, the information distinguishing treated and control covariate distributions must remain fixed.
  • Strict overlap imposes binding population-level restrictions on the data-generating process in high-dimensional covariate settings.
  • Regularization does not avoid the restrictions implied by strict overlap, although it is often necessary for high-dimensional estimation.
  • The results suggest that overlap assumptions should be carefully considered when adjusting for rich covariates.
  • For any fixed overlap bound, finite-sample exact tests can be constructed, and the paper suggests empirical validation should become standard practice.
  • When unconfoundedness is violated, overlap may play a key role in bias amplification from adjusting for highly treatment-predictive covariates that do not predict outcomes.

A Strict Overlap Implies Bounded f-Divergences

This section reframes strict overlap as a likelihood-ratio bound and uses an information-theoretic theorem to bound f-divergences between treated and control covariate distributions.

  • An information-theoretic theorem due to Rukhin is adapted to derive general implications of strict overlap.
  • Strict overlap is reframed as a likelihood-ratio bound between the treated and control covariate distributions.
  • The theorem implies upper bounds on f-divergences between P0 and P1 when the defining convex function is bowl-shaped with a minimum at 1.
  • f-divergences are discrepancy measures between probability distributions defined through a convex function f.
  • Examples include Kullback–Leibler divergence with f(t) = t log t and χ2- or Pearson divergence with f(t) = (t−1)2.

B.1 Strict Overlap Implies Bounded Functional Discrepancy

This section applies the f-divergence result to derive bounds on functional discrepancies under strict overlap, including a χ2-divergence bound that supports the paper’s main theorem.

  • Strict overlap implies an upper bound on functional discrepancies for measurable functions under both covariate distributions.
  • The functional-discrepancy result is used in the proof of Theorem 1 and is stated to be general enough for independent interest.
  • The bound is established by applying the general theorem to the special case of the χ2-divergence.
  • Corollary 4 states that strict overlap implies a bound on the χ2-divergence.
  • With balanced treatment assignment, π = 0.5, the two directional χ2 bounds have a simple form.
  • Corollary 5 extends the result to functional discrepancies, including cases where both variances are infinite.

B.2 Proof of Theorem 1

The proof of Theorem 1 specializes the functional-discrepancy bound to a linear function aligned with the difference in covariate means, then controls variances using operator norms.

  • Theorem 1 follows by choosing a linear function aligned with the difference between treated and control covariate means.
  • The chosen vector has unit length and is normalized by the norm of the covariate-mean difference.
  • The variances of the resulting linear projections are upper-bounded by the operator norms of the corresponding covariance matrices.

C Other implications of strict overlap

The paper derives additional bounds on functional discrepancies using Hölder’s inequality and χ_α-divergences. These bounds tighten the dependence on strict-overlap parameter η but require higher-order moments.

  • Hölder’s inequality combined with χ_α-divergence bounds yields additional upper bounds on mean discrepancy in g.
  • The resulting tighter bounds in η depend on higher-order moments of g(X1:p).
  • χ_α-divergences generalize the χ2-divergence and can be defined in the opposite direction by switching P0 and P1.
  • Under strict overlap with bound η, Theorem 2.1 of Rukhin (1997) supplies the relevant inequality for the derivation.
  • For small η, the new bound scales as η^-1/α, compared with η^-1/2 for (B.3).

D Operator Norm

The operator norm captures how covariance structure affects the bounds on mean imbalance. When covariates are not too correlated, strict overlap forces average mean discrepancies toward zero as dimension grows.

  • The operator norm serves as a proxy for the degree of correlation among covariates under the corresponding probability measure.
  • In independent and stationary covariance cases, the operator norm remains O(1) as dimension grows.
  • With restricted rank s_p=s and component-wise variances bounded away from 0 and ∞, the operator norm is O(p).
  • More generally, when s_p is non-decreasing in p, the operator norm grows as O(p/s_p).
  • If covariates are not too correlated so that ∥Σ1:p∥op=o(p), strict overlap makes mean discrepancies vanish on average as p grows.
Loading 1711.02582v4…