Source-linked AI summary

Semiparametric proximal causal inference

Yifan Cui, Hongming Pu, Xu Shi, Wang Miao, Eric Tchetgen Tchetgen

arXiv:2011.08411v4stat.MEmath.ST

TL;DR

Observational causal inference can fail when measured covariates do not capture latent confounding, because they may be only imperfect proxies. The paper develops semiparametric proximal methods using treatment- and outcome-inducing proxies, establishing identification and doubly robust, locally efficient estimation for the ATE and analogous results for the ATT. The proposed estimators are consistent if either of two working models is correct and achieve the semiparametric efficiency bound when both are correctly specified.

  • Problem

    Exchangeability may be untenable when measured covariates are imperfect proxies for latent confounding mechanisms.

  • Method

    The paper develops semiparametric proximal causal inference using treatment- and outcome-inducing proxies, with alternative identification conditions and efficiency theory for the ATE and ATT.

  • Results

    The proposed proximal estimators are consistent when either of two working models is correct and attain the semiparametric efficiency bound when both are correctly specified.

  • Takeaways & Limitations

    Proximal inference supports causal effect estimation in settings where exchangeability based on measured covariates fails.

  • Takeaways & Limitations

    The framework requires correctly selecting treatment- and outcome-inducing proxies relative to a latent factor sufficient to account for confounding.

Abstract

from arXiv · show

Skepticism about the assumption of no unmeasured confounding, also known as exchangeability, is often warranted in making causal inferences from observational data; because exchangeability hinges on an investigator's ability to accurately measure covariates that capture all potential sources of confounding. In practice, the most one can hope for is that covariate measurements are at best proxies of the true underlying confounding mechanism operating in a given observational study. In this paper, we consider the framework of proximal causal inference introduced by Miao et al. (2018); Tchetgen Tchetgen et al. (2020), which while explicitly acknowledging covariate measurements as imperfect proxies of confounding mechanisms, offers an opportunity to learn about causal effects in settings where exchangeability on the basis of measured covariates fails. We make a number of contributions to proximal inference including (i) an alternative set of conditions for nonparametric proximal identification of the average treatment effect; (ii) general semiparametric theory for proximal estimation of the average treatment effect including efficiency bounds for key semiparametric models of interest; (iii) a characterization of proximal doubly robust and locally efficient estimators of the average treatment effect. Moreover, we provide analogous identification and efficiency results for the average treatment effect on the treated. Our approach is illustrated via simulation studies and a data application on evaluating the effectiveness of right heart catheterization in the intensive care unit of critically ill patients.

1 Introduction

The paper develops a semiparametric proximal framework for causal inference when measured covariates are imperfect proxies and exchangeability fails. It provides identification, efficiency, and doubly robust estimation results for the ATE and ATT using treatment- and outcome-inducing proxies.

  • Motivation: Measured covariates may be imperfect proxies for latent confounding mechanisms, making exchangeability based on measured covariates questionable in observational studies.The framework addresses settings where confounding mechanisms cannot be learned with certainty from measured covariates.
  • Proximal framework: Proximal causal inference requires classifying measured covariates as common causes, treatment-inducing proxies, or outcome-inducing proxies relative to latent confounding factors.Treatment-inducing proxies relate to treatment through an unmeasured common cause, while outcome-inducing proxies relate to outcomes through such a cause.
  • Related work: The proximal framework unifies identification and inference using proxy variables developed across related negative-control and instrumental-variable literatures.The paper positions proximal causal inference as a framework leveraging multiple proxy types from prior work.
  • Contributions: The paper develops semiparametric proximal inference for the ATE and ATT while allowing many observed covariates and using treatment- and outcome-inducing proxies.The framework is intended for point-treatment settings where exchangeability on measured covariates fails.
  • Contributions: The authors establish an alternative condition for nonparametric proximal identification of the ATE and derive semiparametric efficiency bounds under multiple observed-data models.The efficiency analysis includes restricted semiparametric models and a nonparametric model relaxing both sets of restrictions.
  • Contributions: The proposed proximal estimators are doubly robust and locally efficient, remaining consistent when either of two working models is correct and attaining the efficiency bound when both are correct.The relevant intersection submodel is the setting in which all working models are correctly specified.

2 Nonparametric proximal identification of the av-

The paper develops proximal identification of average treatment effects when exchangeability fails because measured covariates may be imperfect proxies for unmeasured confounding. Under proxy, independence, positivity, and completeness conditions, the ATE is identified without directly measuring or modeling the latent confounder.

  • 2.1 Background: Proximal causal inference can preserve identification of the average treatment effect despite unmeasured confounding and failure of exchangeability.The approach uses treatment- and outcome-inducing confounding proxies rather than requiring direct measurement of the latent confounder.
  • 2.2 Proximal identification: Measured covariates are partitioned as X, Z, and W, where Z and W serve as treatment- and outcome-inducing confounding proxies, respectively.The proxy structure is formalized through conditional independence assumptions involving the latent confounder U.
  • 2.2 Proximal identification: Under consistency, conditional independence, conditional randomization, positivity, and completeness, the counterfactual mean and ATE are nonparametrically identified by proximal g-formulas.The ATE is represented as the contrast of identified outcome-bridge functions integrated over W and X.
  • 2.2 Proximal identification: The identification strategy accounts for the latent confounder without measuring U directly or estimating its distribution.Solutions to the relevant integral equations yield a unique ATE value and the same identifying proximal g-formula.
  • 2.3 A new proximal identification result: Theorem 2.2 extends proximal identification beyond the ATE to the marginal counterfactual distribution and smooth functionals defined through moment equations.The result establishes identification of Pr(Y(a)|X) and corresponding functionals under the stated conditions.
  • 2.3 A new proximal identification result: The proposed proximal doubly robust estimator remains valid when at least one of the low-dimensional models for h and q is correctly specified.It is consistent and asymptotically normal over a semiparametric union model, and locally efficient in the semiparametric model.

5 Numerical experiments

The simulations compare proximal and standard doubly robust estimators across correctly specified and misspecified confounding-bridge models. Proximal estimators generally retain small bias or nominal coverage when at least one bridge is correctly specified, whereas the standard estimator is severely biased under unmeasured confounding.

  • Simulation results: Proximal doubly robust estimation has small bias in the first three scenarios, while misspecifying both bridge functions produces bias comparable to proximal IPW and proximal OR.Scenario 4 reflects simultaneous misspecification of both confounding bridge functions.
  • Simulation results: The standard doubly robust estimator is severely biased in all four scenarios because of unmeasured confounding.It is valid in principle only under exchangeability.
  • Simulation results: Table 1 reports simulation absolute bias and MSE for the compared estimators.The simulation results use absolute bias and MSE scaled by 10^-2.
  • Simulation results: Proximal OR yields the narrowest confidence intervals with nominal coverage when h is correctly specified, whereas proximal DR has nominal coverage in the first three scenarios.Proximal IPW confidence intervals have correct coverage in Scenarios 1 and 2.

6 Data analysis

The data application uses proximal causal inference to evaluate right heart catheterization when treatment selection may depend on unrecorded information. Using physiological measurements as treatment- and outcome-confounding proxies, the proximal estimators consistently suggest a more harmful effect on 30-day survival than the standard doubly robust estimator.

  • 6 Data analysis: The study analyzed RHC effects on 30-day survival among 5,735 individuals, including 2,184 treated patients and 3,551 controls.There were 3,817 survivors and 1,918 deaths within 30 days.
  • 6 Data analysis: The selected proxy sets showed significant pairwise partial correlations between Z and W conditional on X and A at the 0.05 level.The outcome and treatment confounding bridge functions were specified using the paper's stated models, including Z-A and X-A interactions for the treatment bridge.
  • 6 Data analysis: The proximal doubly robust, IPW, and outcome-regression estimates were all much larger than the standard doubly robust estimate, with concordance suggesting a more harmful RHC effect on 30-day survival.The analysis reports confidence intervals for the treatment-effect estimates and interprets agreement among the three proximal estimators as support for the modeling assumptions.
  • 6 Data analysis: Sensitivity analyses removing one proxy from Z or W suggested that the proxies were not equally relevant, while remaining broadly aligned with the main estimates.This indicates that proxy selection remains consequential for the application.
  • 7 Discussion: The paper contributes an alternative nonparametric identification condition, semiparametric efficiency results, and doubly robust locally efficient estimators for the ATE and related treatment effects.The framework generalizes standard doubly robust estimation to settings with potential unmeasured confounding by leveraging proxy variables.

A Proof of Theorem 2.2

The proof derives the bridge-function identities needed for proximal identification and uses completeness to connect observed conditional expectations to causal quantities. It then characterizes solution existence through operator-based regularity conditions.

  • The proof uses the treatment bridge equation and completeness to recover E[Y(a)|U, X = x] from observed conditional expectations.The argument factors conditional expectations involving q(Z, a, X) and invokes completeness to establish the required identity.
  • The proof applies singular value decomposition and compact-operator arguments to characterize when the bridge equations admit solutions.The operator formulation uses square-integrable function spaces and range conditions for the relevant operators.
  • The treatment bridge equation is shown to have a solution under stated completeness and regularity conditions.Existence is linked to conditions involving the relevant conditional distributions and bridge-equation assumptions.

C Theorem C.1 and its proof

Theorem C.1 establishes existence and uniqueness of the treatment confounding bridge function under completeness conditions. Its proof shows that any two solutions must agree almost surely.

  • The theorem assumes the existence of a bridge function and invokes the corresponding completeness condition to establish identification.The stated theorem begins with assumptions and existence of a bridge function before proving uniqueness.
  • The proof verifies that a candidate bridge function satisfies the required conditional-expectation identity involving U, A, and X.The argument integrates the candidate over the conditional distribution of U given W, A, and X, then uses the bridge equation.
  • Completeness implies that the treatment bridge equation has a unique solution q.If two candidate bridge functions satisfy the equation, their difference has conditional expectation zero and must therefore vanish almost surely.

D Proof of Theorem 3.1

The proof establishes uniqueness of the outcome bridge function and derives an influence function for the causal-effect parameter. It also places the influence function within the relevant semiparametric tangent-space characterization.

  • Completeness implies that the outcome bridge equation has a unique solution h almost surely.The proof compares two solutions and uses the completeness condition to show that their difference must vanish.
  • The derived influence function combines a weighted residual based on q with the contrast h(W, 1, X) − h(W, 0, X) − ψ.The final expression is identified as an influence function for ψ.
  • The proof characterizes the influence-function space using the range of an operator and its closure, together with an orthogonal-complement component.The decomposition separates terms associated with the bridge-function restrictions from components orthogonal to the relevant conditional structure.
  • The double-robustness proof is deferred to Section F.

E Theorems E.1 and E.2 and their proofs

Theorems E.1 and E.2 characterize efficient influence functions for finite-dimensional parameters in semiparametric models for the outcome and treatment bridge functions. Their proofs derive the corresponding classes of regular asymptotically linear influence functions and impose efficiency conditions.

  • Theorems E.1 and E.2 give semiparametric efficient influence functions for bridge-function parameters in models M1 and M2.The results cover the parameter b for the outcome bridge and the parameter t for the treatment bridge.
  • All regular asymptotically linear estimators of b in M1 have influence functions generated by residuals Y − h(W, A, X; b) multiplied by functions of Z, A, and X.The proof derives this class before selecting the efficient influence function.
  • The efficient influence function is obtained by solving the orthogonality condition against all admissible functions n(W, A, X).The proof uses conditional-mean decompositions and orthogonality to treatment scores in deriving the efficient form.

F Proof of Theorem 3.2

The proof establishes double robustness and local efficiency for the proximal estimator under stated regularity conditions. It derives the influence function through nuisance-parameter expansions and verifies the required mean-zero derivative condition.

  • The proof imposes boundedness and neighborhood regularity conditions for the estimating-function system and its derivatives.
  • The proof establishes double robustness when either the bridge function h or treatment model q is correctly specified.
  • Under regularity conditions, nuisance estimators converge to their target functions, supporting the subsequent asymptotic analysis.
  • A Taylor expansion of the efficient influence function around the parameter and nuisance vector yields the basis for asymptotic normality and local efficiency.
  • The derivative of the efficient influence function with respect to nuisance parameters has expectation zero under the intersection model.

G Average treatment effect on the treated

This section develops proximal identification and semiparametric efficiency results for the average treatment effect on the treated. The ATT requires only a control-group confounding bridge, and its efficient influence function is derived under operator and regularity conditions.

  • Identifying the ATT requires only the confounding bridge function for the control group, a weaker requirement than for the ATE.
  • Theorem G.1 gives the efficient influence function for the treated mean µ when the bridge equation holds and h and q are unique.
  • The ATT influence function combines treated outcomes with a control-group bridge residual weighted by q(Z, X).
  • The efficient influence function characterizes the semiparametric efficiency bound for the ATT model, with estimation and inference proceeding analogously to the ATE case.
  • The derivation assumes surjectivity of the bridge operators and regularity conditions supporting pathwise differentiation and influence-function calculations.

H Choices of the parameters for data generating

The simulation design specifies data-generating mechanisms for proxies, treatment, outcomes, and covariates, then compares proximal estimators with a standard doubly robust estimator. When working models are correctly specified, the proposed estimators perform well.

  • The simulation generates W independently of A and Z conditional on U and X, while its conditional mean depends on U, A, and X.
  • The outcome model includes treatment, covariates, the latent confounder through E(W|U, X), and a residual component involving W.
  • The standard doubly robust estimator uses logistic regression for treatment and linear regression for the conditional outcome mean.
  • The doubly robust and proximal estimators perform well when the working models are correctly specified.
  • Tables 4 and 5 report absolute bias, MSE, confidence-interval coverage, and average length for the simulation estimators.

I.2 Sensitivity analysis on violation of Assumptions 4 and 5

The sensitivity analysis examines proximal estimation when key identifying assumptions are violated and when dependence between the proxy variables is weakened. Proximal estimators become invalid under violated assumptions but retain favorable bias and coverage relative to the standard doubly robust estimator in the weaker-dependence setting.

  • The violated-assumption simulation specifies Z, treatment, W, and potential outcomes as functions of X, U, and independent errors.
  • Tables 6–8 report simulation performance using absolute bias and MSE, with Table 8 covering the weaker-dependence analysis.
  • Under violation of Assumptions 4 and 5, proximal estimators are invalid because their identifying assumptions no longer hold.
  • When dependence between Z and W is weaker, proximal estimators are comparable to or slightly worse than in the preceding setting.
  • In the weaker-dependence setting, proximal estimators outperform the standard doubly robust estimator in bias and coverage.

I.4 Sensitivity analysis on real data application

The sensitivity analysis removes variables from Z and W and compares treatment-effect estimates across four proxy specifications. Results change little when Z includes pafi1, whereas using paco21 alone produces smaller estimates, a positive but nonsignificant proximal IPW estimate, and possible proxy or model problems.

  • Sensitivity results: Scenarios 1 and 2, which retain pafi1 as the sole Z variable, show little change in the treatment-effect results.These scenarios use W=ph1 or W=hema1 with Z=pafi1.
  • Sensitivity results: The proximal IPW estimate becomes positive under the paco21-only specifications, although it is not statistically significant.This differs from the other estimates reported for the sensitivity analysis.
  • Interpretation: Using paco21 alone may not provide a sufficiently relevant treatment-confounding proxy to completely account for confounding.The paper presents this as an interpretation of the sensitivity results rather than a definitive conclusion.
  • Interpretation: Discrepancies among the proximal estimators suggest potential model misspecification in this application.Table 10 reports the corresponding point estimates and 95% confidence intervals for the four scenarios.
Loading 2011.08411v4…