Source-linked AI summary
Conformal Inference of Counterfactuals and Individual Treatment Effects
Lihua Lei, Emmanuel J. Candès
TL;DR
Existing causal-inference methods often focus on CATE while uncertainty quantification for individual treatment effects remains difficult, despite its importance for treatment decisions. The paper develops conformal intervals for counterfactuals and ITEs, obtaining finite-sample coverage in randomized settings and a doubly robust property more broadly. Numerical studies report desired coverage with reasonably short intervals, unlike existing methods that show coverage deficits.
Problem
Flexible machine-learning methods for CATE generally perform poorly in uncertainty quantification, although reliable risk assessment is important for treatment decisions.
Method
The paper uses conformal inference to construct interval estimates for counterfactuals and individual treatment effects under the potential-outcome framework.
Results
Finite-sample coverage is guaranteed for completely randomized or stratified randomized experiments with perfect compliance, while broader settings have approximately controlled average coverage if either propensity scores or conditional potential-outcome quantiles are accurately estimated.
Takeaways & Limitations
The methods achieve desired coverage with reasonably short intervals, whereas existing methods can suffer significant coverage deficits.
Takeaways & Limitations
Inference on individual treatment effects is constrained because the two potential outcomes are never observed simultaneously, and their joint distribution is not identifiable without assumptions.
Abstract
from arXiv · showhide
Evaluating treatment effect heterogeneity widely informs treatment decision making. At the moment, much emphasis is placed on the estimation of the conditional average treatment effect via flexible machine learning algorithms. While these methods enjoy some theoretical appeal in terms of consistency and convergence rates, they generally perform poorly in terms of uncertainty quantification. This is troubling since assessing risk is crucial for reliable decision-making in sensitive and uncertain environments. In this work, we propose a conformal inference-based approach that can produce reliable interval estimates for counterfactuals and individual treatment effects under the potential outcome framework. For completely randomized or stratified randomized experiments with perfect compliance, the intervals have guaranteed average coverage in finite samples regardless of the unknown data generating mechanism. For randomized experiments with ignorable compliance and general observational studies obeying the strong ignorability assumption, the intervals satisfy a doubly robust property which states the following: the average coverage is approximately controlled if either the propensity score or the conditional quantiles of potential outcomes can be estimated accurately. Numerical studies on both synthetic and real datasets empirically demonstrate that existing methods suffer from a significant coverage deficit even in simple models. In contrast, our methods achieve the desired coverage with reasonably short intervals.
1. From Average Effects To Individual Effects
Average effects and CATE can miss individual-level variability that matters for treatment decisions, while reliable uncertainty quantification remains difficult. The paper addresses these gaps by constructing conformal intervals for individual treatment effects and counterfactuals.
- Motivation: ATE provides only a coarse summary of treatment-effect heterogeneity and may be insufficient or misleading for validating an intervention.The paper illustrates this with a treatment that helps 70% of patients but substantially worsens symptoms for the remaining 30%.
- Motivation: CATE is richer than ATE but can still neglect response variability and estimator variability unless covariates explain most heterogeneity and CATE is estimated nearly perfectly.The paper describes both conditional response variation and finite-sample variation in CATE estimators as immediate concerns.
- Motivation: Reliable uncertainty quantification is important for high-stakes decisions, yet modern machine-learning methods are typically under-studied and existing guarantees often rely on strong assumptions or asymptotics.The paper notes that confidence intervals or p-values are required for drug approval and that finite-sample reliability is difficult to establish.
- Approach: The paper uses conformal inference to construct prediction intervals for individual treatment effects under the potential-outcome framework.It targets subjects in the study with one missing potential outcome and subjects outside the study with both potential outcomes missing.
- Guarantees: For randomized trials with perfect compliance, the method provides finite-sample coverage without assumptions beyond i.i.d. sampling; broader settings receive coverage when either the outcome or treatment model is accurately estimated.The latter property is described as analogous to doubly robust guarantees for average-treatment-effect point estimates.
- Approach: Counterfactual intervals can be shifted into individual-treatment-effect intervals by contrasting the missing potential outcome with the observed outcome.A nested approach also generates intervals for in-sample subjects and trains a model to generalize them to subjects outside the study.
- Scope: The methods extend to generalizability or transportability settings with distributional shifts between study and target populations and to other causal-inference frameworks.These extensions are discussed for settings involving a different target-population covariate distribution and frameworks such as causal diagrams.
2. From Point Estimates To Interval Estimates
The paper frames individual treatment effects and counterfactuals as interval-inference targets because only one potential outcome is observed per unit and point-estimate uncertainty is difficult to quantify reliably. It distinguishes marginal from conditional coverage and accommodates target-population shifts through weighted criteria.
- Framework: Under the binary-treatment potential-outcome framework, each unit has two potential outcomes, but the observed data contain only the outcome corresponding to its treatment.The paper denotes treatment, potential outcomes, and covariates by T_i, Y_i(1), Y_i(0), and X_i.
- Framework: Because one potential outcome is missing for every unit, individual treatment effects are unobserved and must be inferred under strong ignorability.Strong ignorability rules out unmeasured confounders affecting both treatment assignment and potential outcomes.
- Inferential targets: Existing work mainly targets CATE, the conditional expectation of ITE, but the paper instead seeks prediction intervals for counterfactual outcomes and ITEs.The target intervals include C_1(x), C_0(x), and C_ITE(x), with marginal coverage at a prespecified level.
- Inferential targets: Estimated conditional quantiles may fail to deliver valid coverage because effective sample sizes are limited and the outcome model is imperfect.Oracle quantiles would provide conditional coverage, but substituting estimated quantiles can invalidate the interval.
- Coverage: Marginal coverage controls average validity, whereas guaranteed nontrivial conditional coverage is impossible without modeling assumptions.The paper compares this criterion with average-performance measures such as RMSE.
- Coverage: Coverage criteria can condition on treated or control units, or use a target-population covariate distribution for generalizability and transportability.The general criterion replaces the study covariate distribution with Q_X, which may represent the target population.
3. From Observables To Counterfactuals
The paper develops weighted conformal methods for counterfactual and individual treatment-effect intervals under strong ignorability, including covariate-shift adjustment. The methods provide finite-sample or approximate coverage guarantees and show strong empirical coverage across varied simulation settings, with reasonably short intervals.
- Counterfactual inference: Counterfactual intervals use treated or control observations under the strong ignorability assumption, while covariate shifts are handled through weighted conformal inference.The target and sampling distributions share the conditional outcome distribution but may differ in covariate distribution.
- Coverage guarantees: Conformalized regression provides finite-sample coverage guarantees without assumptions on the unknown joint distribution of covariates and outcomes.The guarantee applies regardless of the unknown data-generating mechanism.
- Coverage guarantees: For randomized trials with perfect compliance, weighted conformal inference uses known propensity scores and achieves coverage, including under stratified designs with covariate-varying assignment probabilities.Completely randomized experiments need no weighting when the propensity score is constant.
- Coverage guarantees: Weighted split-CQR retains approximate coverage when either the propensity score or conditional potential-outcome quantiles are estimated accurately.If the estimated weights are inaccurate, coverage is bounded using the total variation error Δw; accurate conditional quantiles can also suffice.
- Numerical experiments: In simulations, weighted split-CQR achieves almost exact ITE coverage across learning procedures, covariance structures, error variances, and dimensions.For CATE, it achieves coverage in all scenarios but may be conservative; its ITE intervals can be reasonably short and nearly oracle-length in homoscedastic settings with BART.
- Numerical experiments: Weighted split-CQR maintains conditional coverage as conditional variance increases, whereas Causal Forest, X-learner, and BART show decreasing conditional coverage in the heteroscedastic experiment.Quantile random forest and quantile gradient boosting perform best among the weighted split-CQR variants in this comparison.
- Scope: The reported comparisons concern interval estimates of ITE and CATE, not the accuracy of point estimates of CATE.Point-estimate accuracy does not determine interval coverage, and conformal intervals need not be derived from point estimates.
4. From Counterfactuals To Treatment Effects
The section develops two nested strategies for turning counterfactual intervals into ITE intervals for subjects with both potential outcomes missing. The exact strategy preserves a coverage guarantee, while the inexact strategy trades theoretical control for shorter intervals.
- From counterfactuals to ITE: Counterfactual intervals for both potential outcomes can be contrasted to form an ITE interval with the same 1−α coverage level.The interval uses the lower bound for Y(1) minus the upper bound for Y(0), and vice versa.
- Nested approach: The nested approach splits the data, trains counterfactual interval models on one fold, and constructs surrogate ITE intervals for units in the other fold.For treated units it predicts Y(0), while for controls it predicts Y(1), then combines each prediction with the observed outcome.
- Coverage guarantees: For randomized experiments, weighted counterfactual intervals satisfy the required conditional guarantees using propensity-based weights; observational settings are approximately valid when either propensity scores or conditional quantiles are well estimated.The procedure does not require splitting α between the two potential outcomes.
- Inexact and exact methods: The inexact method fits conditional quantiles of surrogate interval endpoints, whereas the exact method applies a second conformal calibration to expand them.The exact construction bounds the final ITE miscoverage by α + γ.
- Exact calibration: Jointly calibrating the left and right endpoints avoids the potentially conservative Bonferroni correction from calibrating them separately.Algorithm 2 uses a shared conformity score and calibration quantile to produce the expanded interval.
- Evaluation: The section evaluates the nested procedures in numerical experiments using synthetic observational data based on the NLSM study.The experiments compare interval coverage and length against competing approaches.
Step I: data splitting
The first step of the nested algorithm splits the data and estimates the propensity score on the first fold. These estimates determine the weights used for counterfactual inference on the second fold.
- Step I: data splitting: The algorithm estimates the propensity score ê(x) on the first data fold before constructing counterfactual intervals.This separates propensity estimation from the subsequent counterfactual inference step.
- Step I: data splitting: The second fold is reserved for counterfactual inference using the first fold as the training data.This fold structure supports the nested construction of surrogate ITE intervals.
- Step I: data splitting: For treated units, the method predicts the missing control outcome using weight w0(x)=ê(x)/(1−ê(x)) and combines it with the observed treated outcome.The resulting interval is a surrogate interval for the unit’s ITE.
- Step I: data splitting: For control units, the method predicts the missing treated outcome using weight w1(x)=(1−ê(x))/ê(x) and combines it with the observed control outcome.This is the treatment-side counterpart of the construction for treated units.
Step III: Interval of ITE on the testing point
The final step produces testing-point ITE intervals using either exact conformal calibration or inexact endpoint quantile modeling. Experiments show that inexact-CQR attains target coverage with comparatively short intervals, while coverage varies across methods and settings.
- Step III: Interval of ITE on the testing point: The exact version applies interval-outcome conformal calibration, while the inexact version fits conditional quantiles of surrogate interval endpoints.Both versions output an ITE interval for the testing point.
- Synthetic evaluation: All procedures use BART for conditional quantiles and gradient boosting for propensity-score estimation; the exact nested method sets α=γ=0.025.BART-based competitors are also evaluated without exact nested calibration.
- Synthetic results: Inexact-CQR reaches the desired coverage with intervals substantially shorter than those from naive or exact nested methods, while BART alone misses the target coverage.Causal Forest and X-learner also have poor coverage despite producing short intervals.
- Limitations: The oracle interval length cannot generally be attained because conditional ITE quantiles are unidentifiable without assumptions on the joint distribution of the two potential outcomes.Weighted split-CQR nevertheless produces reasonably short intervals.
- Synthetic results: Inexact-CQR has relatively even conditional coverage across conditional variance and CATE values, whereas inexact-BART performs worse.The conditional-coverage analysis is based on 100 replicates at a 95% target level.
- NLSM re-analysis: The NLSM re-analysis is exploratory because the ground truth is unavailable, and interval behavior is examined across 100 repeated data splits.Figure 7 reports average length and the fractions of intervals with exclusively positive or negative signs.
5. From Potential Outcomes to Other Causal Frameworks
The framework extends beyond potential outcomes by exploiting covariate shift and invariance of conditional outcome distributions. The paper applies this logic to do-interventions and outlines analogous extensions to invariant prediction, while leaving part of the latter development for future work.
- Extensions: Weighted conformal inference provides interval estimates for counterfactuals and ITEs, with finite-sample guarantees under randomized perfect-compliance designs and double robustness under broader assumptions.The broader guarantee holds approximately when either propensity scores or conditional potential-outcome quantiles are consistently estimated.
- Extensions: The key extension principle is covariate shift combined with invariance of the conditional distribution between observed and target populations.This structure lets the method reuse weighted prediction-interval machinery across causal settings.
- Causal diagram framework: In the graphical causal framework, variables satisfying the back-door criterion include confounders and exclude post-treatment variables.The paper restricts its discussion to treatment, outcome, and covariates meeting this criterion.
- Causal diagram framework: Under the graphical framework’s matching distributional structure, weighted split-CQR can produce doubly robust intervals for outcomes under a do intervention.The paper treats this as structurally analogous to the potential-outcome setting.
- Invariant prediction: Invariant prediction assumes Y is conditionally independent of environment E given X, while the covariate distribution may vary across environments.The target environment requires prediction under a shifted covariate distribution.
- Invariant prediction: With one source environment, weighted split-CQR has the same structure as the potential-outcome problem; multiple environments require more complex weights or a weighted pseudo-population.The paper describes these constructions but does not fully develop the multi-environment case.
A. Nonasymptotic Theory for Double Robustness of Weighted Split-CQR
This section establishes nonasymptotic double-robustness results for weighted split-CQR, covering two complementary conditions and a simpler limiting formulation.
- A. Nonasymptotic Theory for Double Robustness of Weighted Split-CQR: The analysis develops nonasymptotic results for two sides of weighted split-CQR's double robustness.A simpler asymptotic corollary is presented separately.
- A. Nonasymptotic Theory for Double Robustness of Weighted Split-CQR: Theorem 3 assumes an estimated covariate-shift ratio with finite conditional expectation and normalizes it to have expectation one.The resulting conformal interval is constructed from estimated conditional quantiles and the normalized weight.
- A. Nonasymptotic Theory for Double Robustness of Weighted Split-CQR: Theorem 4 adds regularity conditions involving conditional outcome probabilities, finite likelihood ratios, and moment or concentration parameters.Under these conditions, the theorem provides constants controlling the resulting bounds.
- A. Nonasymptotic Theory for Double Robustness of Weighted Split-CQR: When conditional quantile estimation errors satisfy a deterministic o(1) bound, the theorem bounds simplify as k and ℓ tend to infinity.This is stated in Remark 1 for δ ≥ 1.
A.1. Proof of Theorem 3
The proof of Theorem 3 constructs a weighted calibration argument under covariate shift, using normalized estimated weights and total-variation control to obtain coverage bounds.
- A.1. Proof of Theorem 3: The proof defines the target quantile through the generalized quantile operator Quantile(β; F).It uses the equivalent infimum and supremum characterizations of the quantile.
- A.1. Proof of Theorem 3: For absolutely continuous QX, the proof compares weighted calibration under QX with calibration under PX using the likelihood ratio w(x).The argument introduces a measure with density ˆw relative to PX and a corresponding sample under that measure.
- A.1. Proof of Theorem 3: The estimated weights are normalized, while the proof distinguishes the true probabilities pi from their estimated counterparts ˆpi.This enables the weighted exchangeability argument used for the calibration scores.
- A.1. Proof of Theorem 3: The resulting coverage control includes a penalty proportional to PX∼QX(E∞)EX∼PX|ˆw(X) −w(X)| when the likelihood ratio may be infinite.On the finite-weight region, the proof first establishes the desired bound and then extends it by conditioning and rescaling.
- A.1. Proof of Theorem 3: If w(Xn+1) is infinite, the correction becomes infinite and the conformal interval expands to (−∞, ∞).This handles the part of the target distribution outside the finite-weight region.
A.2. Proof of Theorem 4
The proof of Theorem 4 controls weighted split-CQR's finite-sample deviations through moment inequalities, concentration bounds, and quantile-estimation error terms.
- A.2. Proof of Theorem 4: The proof uses Rosenthal-type inequalities for independent sums with finite (1 + δ)-th moments.It invokes Rosenthal's inequality for δ ≥ 1 and von Bahr–Esseen's inequality for δ ∈ [0, 1).
- A.2. Proof of Theorem 4: The argument assumes finite likelihood ratios under the relevant distributions and analyzes a generic target draw from QX × PY |X.These assumptions ensure the weighted quantities and target construction are well-defined.
- A.2. Proof of Theorem 4: The proof bounds the probability that the conformal correction is too small by decomposing empirical fluctuation, weight, and conditional-quantile error contributions.The decomposition is controlled using sub-Gaussian concentration, moment bounds, and the function H measuring quantile-estimation error.
- A.2. Proof of Theorem 4: The quantile-estimation error enters through HN(x), while the proof controls its contribution using moment assumptions and Markov-type bounds.The resulting rate depends on r, b1, b2, δ, M, k, and ℓ.
- A.2. Proof of Theorem 4: P(˜Y ∈ ˆC(˜X)) ≥ 1 − α − (3B + 1)ϵn.This bound follows after combining the conditional coverage and target-distribution error terms.
A.3. An asymptotic result
Theorem 3 and Theorem 4 yield an asymptotic generalization of Theorem 1 when either of two stated conditions holds.
- A.3. An asymptotic result: The corollary combines Theorem 3 and Theorem 4 into an asymptotic result.
- A.3. An asymptotic result: The result applies when either B1 or B2, or both, is satisfied.
- A.3. An asymptotic result: Condition B2 consists of conditions (1)–(4) from Theorem 4.
B.1. Proof of Proposition 1
The proof extends special cases to a general weighted split-CQR result by normalizing the estimated weights and using an induced probability measure. It then establishes the bounds through exchangeability and auxiliary inequalities, with inverse-propensity conditions used in the corollary reduction.
- General proof strategy: The proof reduces the general result to normalized estimated weights and a probability measure induced by those weights.The normalization is justified by weighted split-CQR's invariance to rescaling; the induced sample is taken independently from the corresponding weighted distribution.
- Bounds: The lower bound follows from (A.1), while the upper bound follows from (A.9) and the subsequent argument.The supplied proof passage identifies these as the key steps for the two bounds.
- Assumption reduction: The corollary's assumptions are reduced using inverse propensity weights and integrability of the estimated and true propensity-score reciprocals.The reduction uses the definitions of ˆwN(x) and w(x), together with the conditions E[1/ˆeN(X) | Ztr] < ∞ and E[1/e(X)] < ∞.
- Exchangeability: Conditional on the training data, the calibration variables and the new variable are exchangeable.This exchangeability is stated for V1, . . . , Vn, Vn+1 after defining an independent copy of (X, C).
C. Additional Experimental Results
The additional experiments report estimated conditional ITE coverage across heteroscedastic settings and across CATE values. Results are shown for dimensions d = 10 and d = 100, with the higher-dimensional setting exhibiting behavior similar to the lower-dimensional one.
- Figure 8: Estimated conditional ITE coverage is evaluated as a function of conditional variance σ2(x) in heteroscedastic cases with d = 100 and α = 0.05.The figure summarizes medians and 95% and 5% quantiles across 100 replicates.
- Figure 9: Estimated conditional ITE coverage is evaluated as a function of CATE τ(x) for scenarios with d = 10.The remaining settings are stated to match Figure 8.
- Figure 10: The d = 100 scenarios show behavior very similar to the corresponding d = 10 scenarios when coverage is plotted against CATE τ(x).Figure 10 explicitly describes the comparison with Figure 9 as similar behavior.