Source-linked AI summary

An Introduction to Proximal Causal Learning

Eric J Tchetgen Tchetgen, Andrew Ying, Yifan Cui, Xu Shi, Wang Miao

arXiv:2009.10982v1stat.ME

TL;DR

Observational causal inference can fail when measured covariates are only imperfect proxies for unmeasured confounding, leaving standard exchangeability unsupported. The paper introduces proximal causal learning, derives identification through a proximal g-formula, and develops proximal g-computation for point and time-varying treatments. It also illustrates the approach with an application in which the proximal estimate differs from standard OLS.

  • Problem

    Measured covariates may be imperfect proxies for confounding mechanisms, making exchangeability-based causal inference unreliable and leaving causal learning from proxies as an unresolved inverse problem.

  • Method

    The paper develops a potential-outcome framework using treatment-inducing and outcome-inducing proxies, with proximal g-formula identification and proximal g-computation estimation.

  • Results

    The framework provides sufficient conditions under which causal effects can sometimes be identified despite unmeasured confounding, and supports point and time-varying treatment analyses.

  • Takeaways & Limitations

    Proximal causal learning offers a formal way to analyze causal effects when measured covariates do not establish exchangeability, using interpretable proxy-based assumptions.

  • Takeaways & Limitations

    The proximal g-formula may require numerically solving an empirically ill-posed equation, which can be computationally intensive and unstable; estimation also relies on correctly specified outcome confounding bridge functions.

Abstract

from arXiv · show

A standard assumption for causal inference from observational data is that one has measured a sufficiently rich set of covariates to ensure that within covariate strata, subjects are exchangeable across observed treatment values. Skepticism about the exchangeability assumption in observational studies is often warranted because it hinges on investigators' ability to accurately measure covariates capturing all potential sources of confounding. Realistically, confounding mechanisms can rarely if ever, be learned with certainty from measured covariates. One can therefore only ever hope that covariate measurements are at best proxies of true underlying confounding mechanisms operating in an observational study, thus invalidating causal claims made on basis of standard exchangeability conditions. Causal learning from proxies is a challenging inverse problem which has to date remained unresolved. In this paper, we introduce a formal potential outcome framework for proximal causal learning, which while explicitly acknowledging covariate measurements as imperfect proxies of confounding mechanisms, offers an opportunity to learn about causal effects in settings where exchangeability on the basis of measured covariates fails. Sufficient conditions for nonparametric identification are given, leading to the proximal g-formula and corresponding proximal g-computation algorithm for estimation. These may be viewed as generalizations of Robins' foundational g-formula and g-computation algorithm, which account explicitly for bias due to unmeasured confounding. Both point treatment and time-varying treatment settings are considered, and an application of proximal g-computation of causal effects is given for illustration.

1 Introduction

The paper motivates proximal causal learning as a framework for estimating causal effects when measured covariates are imperfect proxies and exchangeability fails. It introduces proxy-based identification and estimation methods, including proximal g-computation for point and time-varying treatments.

  • Motivation: Exchangeability requires sufficiently rich measured covariates, but accurately capturing all relevant confounding mechanisms is difficult and the assumption is empirically untestable.Residual confounding may remain even after adjusting for measured covariates.
  • Proximal causal learning: Proximal causal learning uses imperfect covariate measurements as proxies to potentially learn causal effects when exchangeability based on measured covariates does not hold.The framework explicitly acknowledges that measured covariates may not capture underlying confounding mechanisms.
  • Proxy structure: Treatment-inducing and outcome-inducing proxies are selected alongside other observed covariates to represent distinct relationships with unmeasured confounding.The paper illustrates outcome-inducing proxies using baseline measurements of cognitive impairment and treatment adherence-related factors.
  • Linear illustration: A proximal control variable based on E(W|A, Z, X) can yield a consistent treatment-effect slope, whereas omitting it or substituting W or (W, Z) generally remains biased.This result is presented first in a linear structural model, while later arguments relax linearity and interaction restrictions under suitable conditions.
  • Framework and estimation: The paper develops nonparametric identification conditions, the proximal g-formula, and proximal g-computation for point and time-varying treatment settings.The proximal g-formula generalizes Robins’ g-formula to account for confounding bias due to unmeasured factors.

2 Notation and definitions

The notation section defines potential outcomes and the target population causal effect, then reviews exchangeability-based identification and proxy configurations represented with directed acyclic graphs. It sets up the distinction between standard measured-confounding adjustment and settings involving proxy types.

  • Notation: Observed data consist of i.i.d. samples on treatment A, covariates L, and outcome Y, with Ya denoting the potential outcome under treatment level a.The framework also invokes the standard consistency assumption linking observed and potential outcomes.
  • Causal estimand: The target is the population average causal effect, expressed through contrasts of counterfactual means such as E(Y1) − E(Y0) for binary treatment.Causal effects may be defined on different scales for binary, polytomous, or continuous treatments.
  • Standard identification: Standard identification uses exchangeability or no unmeasured confounding conditional on measured covariates, together with positivity, yielding the g-formula.Exchangeability is often interpreted as requiring measured covariates to include all common causes of treatment and outcome.
  • Proxy configurations: Measured covariates can be organized as common causes, treatment-inducing proxies, and outcome-inducing proxies in alternative directed acyclic graph configurations.The configurations include settings where proxies share unmeasured common causes with the opposite variable.
  • Graphical conditions: The reviewed graphical settings include cases where exchangeability holds and a case where it fails because an unmeasured common cause of A and Y remains.One exchangeability configuration additionally requires conditional independence of specified unmeasured variables.

3 Proximal identification in point exposure studies

Proximal identification replaces exchangeability based on measured covariates with proxy-based conditions that can identify causal effects despite unmeasured confounding. For point exposures, the approach uses treatment- and outcome-inducing proxies, completeness conditions, and a proximal g-formula based on an outcome confounding bridge function.

  • Figure 2 represents point-exposure settings where measured covariates fail to ensure exchangeability because an unmeasured common cause affects treatment and outcome.
  • The method classifies Z as a treatment-inducing confounding proxy and W as an outcome-inducing confounding proxy under proxy-validity assumptions.Neither proxy must directly cause the treatment or outcome, provided both are relevant to the unmeasured confounder and satisfy the required assumptions.
  • Completeness requires proxies to have sufficient variability relative to the unmeasured confounder; for categorical variables, Z and W must each have at least as many categories as U.Together with a matrix rank condition, this can suffice for the relevant integral equation to admit a solution.
  • The proximal g-formula identifies the counterfactual mean by integrating an outcome confounding bridge function h(a, x, w) against the distribution of W and X.The bridge function solves an integral equation linking observed conditional outcome distributions to the latent-confounder adjustment target.
  • Solving the bridge equation is an inverse problem that may be computationally intensive, unstable, and empirically ill-posed, motivating regularization or model-based stable solutions.Small estimation uncertainty in the equation’s left-hand side can induce excessive uncertainty in the solution.
  • Under the structural linear model, proximal g-computation recovers the causal parameter βa as ηa by applying the proximal g-formula.The resulting effect equals the expected difference between h(A = 1, X, W; η) and h(A = 0, X, W; η).

4 Proximal identification in complex longitudinal studies

The longitudinal framework replaces sequential exchangeability with proxy-based conditions that identify causal effects despite unmeasured confounding. Under bridge equations and completeness, the counterfactual mean is identified through outcome- and treatment-history proxy information.

  • Longitudinal setup: The framework considers two follow-up treatment and covariate measurements, with the outcome observed at the end of follow-up.The setup uses time-varying treatment and covariate data at visits j = 0, 1 and outcome Y at j = 2.
  • Motivation: Longitudinal proximal identification targets causal effects when sequential exchangeability cannot be guaranteed in observational studies.Measured covariates are incorporated as proxies for underlying confounding mechanisms.
  • Identification conditions: Identification requires longitudinal proxy conditions, completeness, and bridge functions satisfying equations (23) and (24).The first bridge equation links the observed outcome conditional mean to H1, while the second links H1 to H0 using baseline proxies.
  • Identification result: Under assumptions (19)–(24), the causal parameter is identified as β(a) = E{H0(a)} = E{h0(W(0), a, X(0))}.The bridge-function representation identifies the marginal counterfactual mean without requiring unique identification of H1(a) and H0(a).

5 Proximal g-computation

Proximal g-computation estimates causal effects by modeling outcome bridge functions and evaluating their implied counterfactual means. The section develops parametric, least-squares, and more flexible implementations, with consistency under the stated model conditions.

  • Core approach: Proximal g-computation specifies parametric models for outcome bridge functions and estimates causal effects from their fitted counterfactual means.The procedure is presented first for point treatment and then extended to time-varying treatment.
  • Core approach: Directly modeling the outcome confounding bridge function avoids solving complicated ill-posed integral equations and acts as regularization.The approach is appropriate when the outcome mean admits the specified bridge-function representation.
  • Inference: Under correct model specification and the identification result, the estimated counterfactual mean is consistent and approximately normally distributed.The proposed longitudinal estimator is obtained through recursively fitted regression models.
  • Implementation: The proximal recursive least squares algorithm recursively fits multivariate and ordinary least-squares regressions to estimate bridge-function components.The steps include fitting proxy regressions, substituting fitted values, and estimating H0.
  • Robustness: Under linearity of H0 and H1 in W, the recursive least-squares estimator remains consistent even when specified linear models for W and W(0) are misspecified.In the point-treatment setting, consistency can likewise persist when the model for the proxy regression is misspecified, provided h0 and h1 are correctly specified.
  • Efficiency and flexibility: When the model for f is correctly specified, proximal g-computation can generally be more efficient than proximal recursive least squares and recursive generalized method of moments.More flexible semiparametric and nonparametric models for h are also described as options for reducing specification concerns.

6 Data applications

The applications illustrate proximal estimation in point-treatment and time-varying-treatment settings, showing stronger estimated effects than standard analyses in both examples.

  • 6.1 Point treatment application: The SUPPORT application evaluates right heart catheterization during initial ICU care on survival through 30 days using 73 patient covariates.The study included 2,184 RHC-treated and 3,551 untreated patients.
  • 6.1 Point treatment application: P2SLS estimated a more harmful RHC effect than OLS, changing the estimate from −1.25 (SE 0.28) to −1.80 (SE 0.43).The outcome was survival time up to 30 days.
  • 6.1 Point treatment application: The analysis found moderate empirical evidence that unmeasured confounding biased the OLS estimate through the outcome-inducing proxy and confounding bridge parameter.The reported bridge parameter was η̂w = −16.92 with standard error 8.8.
  • 6.2 Time-varying treatment application: The rheumatoid arthritis application analyzes joint causal effects of methotrexate on the average number of tender joints over follow-up.Methotrexate exposure was classified as ever-treated or never-treated, providing a conservative efficacy estimate.
  • 6.2 Time-varying treatment application: Proximal recursive least squares estimated a stronger protective methotrexate effect than IPW least squares, with estimates of −0.37 and −0.23, respectively.The proximal estimate had a 95% confidence interval of (−0.67, −0.13), while the IPW interval was (−0.43, −0.02).

7 Discussion

The discussion presents proximal causal learning as a framework for observational analyses with potentially incomplete confounding measurements, while identifying model specification and finite-sample evaluation as continuing concerns.

  • 7 Discussion: The framework treats measured covariates as proxy measurements of underlying confounding factors and provides a potential-outcome basis for identifying causal effects from proxies.It includes proximal g-formula and proximal g-computation algorithms for point-treatment and time-varying-treatment settings.
  • 7 Discussion: Proximal causal learning is closely related to negative control methods for detecting and sometimes estimating effects of point-treatment interventions.The discussion directs readers to a review of the negative control literature.
  • 7 Discussion: The proposed g-computation, proximal two-stage least squares, and recursive least-squares methods rely on correct specification of outcome confounding bridge functions.The authors report ongoing development of treatment-bridge-based estimators and doubly robust estimators.
  • 7 Discussion: Doubly robust estimators are described as remaining unbiased in large samples when at least one confounding bridge function model is correct.The discussion states that finite-sample performance will be evaluated in future work.

Appendix

The appendix provides recursive-estimation details, closed-form and likelihood expressions, consistency arguments, proxy-assumption graphs, and empirical-application result tables.

  • Closed-form expressions: For binary outcomes and a continuous scalar proxy, the appendix gives a closed-form expression using a known link function and a bridge-error distribution.The displayed formulation specifies g as a known link function and integrates over fεW.
  • Likelihood estimation: The likelihood-based formulation estimates parameters by maximizing a log-pseudo-likelihood combining the binary-outcome contribution with the bridge-error density.Probit and logit links use different bridge-distribution specifications, including Gaussian and logistic forms.
  • Likelihood estimation: For multivariate proxies with continuous and discrete components, the appendix factorizes the conditional proxy density and uses a multivariate Gaussian specification for continuous components.The covariance structure is represented by Σ under a probit-link formulation.
  • Recursive estimation: The recursive least-squares appendix repeatedly applies least-squares regressions to fitted proxy functions and recursively updated outcome quantities.The procedure begins with fitted values for ĉw,j and sets ĤJ = Y before backward recursion.
  • Consistency results: The appendix states that the recursive estimator is consistent for the causal parameter under the specified conditions, including when a relevant model is misspecified.The proof uses conditional mean-zero relations and least-squares projection properties.
  • Proxy-assumption graphs: Table A.1 combines graph pieces for treatment-inducing and outcome-inducing proxies, marking graphs invalid when key proxy assumptions are violated.The examples encode relationships among Z, A, U and W, Y, U.
  • Empirical results: Tables A.2 and A.3 report results for the right-heart-catheterization and methotrexate empirical applications, respectively.The appendix tables correspond to the two application analyses described in the paper.
  • Empirical results: The appendix notes conventional significance thresholds of p<0.1, p<0.05, and p<0.01.These thresholds are marked with one, two, and three asterisks, respectively.
Loading 2009.10982v1…