Source-linked AI summary
Comment: Demystifying Double Robustness: A Comparison of Alternative Strategies for Estimating a Population Mean from Incomplete Data
Anastasios A. Tsiatis, Marie Davidian
TL;DR
The comment addresses continuing confusion about double robustness in estimating a population mean from incomplete data. It uses semiparametric theory and influence functions to compare estimators under alternative models, concluding that correctly specified doubly robust estimators share asymptotic variance while performance can differ under misspecification and small estimated observation probabilities. It also identifies limits of the parametric first-order framework.
Problem
The comment examines how double robustness should be understood when estimating a population mean with outcomes missing under a MAR mechanism.
Method
The comment analyzes statistical models for observed data and derives corresponding influence functions to compare estimator types and their large-sample properties.
Results
When both models are correctly specified, all doubly robust estimators have the same asymptotic variance, while performance differences can arise from model-specific estimation choices and small estimated observation probabilities.
Takeaways & Limitations
The comparison frames estimator performance through the assumptions and estimation choices embedded in statistical models, rather than through estimator forms alone.
Takeaways & Limitations
Without parametric assumptions beyond mild smoothness for E(y|x) and E(t|x), first-order asymptotic theory may no longer apply and higher-order theory may be needed.
Abstract
from arXiv · showhide
Comment on ``Demystifying Double Robustness: A Comparison of Alternative Strategies for Estimating a Population Mean from Incomplete Data'' [arXiv:0804.2958]
INTRODUCTION
The comment examines double robustness through the simple problem of estimating a population mean with MAR missing outcomes. It complements the authors’ presentation by emphasizing assumptions and estimator types from a semiparametric perspective.
- The discussion uses population-mean estimation with outcomes missing under MAR as a simple setting for examining double robustness.
- The missing-data problem is equivalent to estimating a potential-outcome mean under no unmeasured confounding.
- The literature distinguishes likelihood-oriented and weighting-based schools for addressing missing-data and causal-inference problems.
- The comment focuses on semiparametric assumptions embedded in statistical models leading to different estimator types rather than estimator forms.
SEMIPARAMETRIC THEORY PERSPECTIVE
The comment uses semiparametric models and influence functions to compare estimators for a population mean with incomplete outcomes. It argues that double-robust performance is determined by the models and influence-function forms, not simply by whether an estimator is AIPW.
- Influence-function framework: Influence functions determine the large-sample behavior of regular, asymptotically linear estimators through consistency, asymptotic normality, and asymptotic variance.The influence function has mean zero and finite second moment, and the asymptotic variance is E{ϕ^2(z)}.
- Statistical models: The analysis represents observed data as z = (t,x,ty) under missing at random and considers three semiparametric models for the outcome and observation mechanisms.Models I and II separately specify the conditional outcome mean or observation probability, while model III specifies both; p(x) remains unrestricted.
- Double robustness: Doubly robust estimators are consistent and asymptotically normal when either the outcome model or observation model is correct, and they share an influence function and asymptotic variance when both are correct.Under the intersection of models I and II, the result does not depend on how the outcome and observation-model parameters are estimated.
- Estimator comparisons: Under model II, performance depends on how closely γ* + m(x,β*) approximates m0(x) and on estimation of α, for which maximum likelihood is described as optimal.The comment therefore attributes differences among estimators to influence-function forms and nuisance-parameter estimation rather than AIPW membership.
- Estimator comparisons: Under model I, estimator performance depends on the form of ea(x) and on how the outcome-model parameter β is estimated through A(x,β).The comment derives regression and imputation estimators from particular choices of the influence-function component a(x).
- Estimator comparisons: When the observation model is incorrect but the outcome model is correct, bµSRR performs poorly because its design matrix becomes unstable for small bπi.Its parameter vector augments the outcome-model coefficients with the coefficient of bπ−1_i, producing this instability.
BOTH MODELS INCORRECT
When both working models are misspecified, estimator performance depends strongly on the missingness pattern and extrapolation quality. The discussion proposes several robustness strategies while cautioning that simulation findings are not broadly generalizable.
- Both-model misspecification is the setting for assessing doubly robust and other estimators outside their formal robustness class.
- Small estimated observation probabilities can make some doubly robust and inverse-probability-weighted estimators perform poorly because responses are missing from parts of the covariate space.
- In the KS scenario, estimators using estimated observation probabilities can fail badly, while outcome-model extrapolations remain minimally affected.
- A different scenario could instead make unobserved covariate regions highly influential for the outcome regression, harming outcome-model imputation; simulations therefore lack broadly applicable conclusions.
- Possible strategies include imposing extra assumptions through estimators like bµπ-cov or constructing hybrid influence functions that combine bµOLS and doubly robust estimators.
- The estimator in equation (11) spans doubly robust, imputation, and bµOLS forms as π(xi) varies, motivating shrinkage toward a common value and alternatives to logistic missingness models.
- Avoiding parametric assumptions for E(y|x) and E(t|x) may require higher-order asymptotic theory beyond first-order methods.
CONCLUDING REMARKS
The commentators regard the article as thoughtful and insightful and anticipate methodological developments addressing the challenges it highlights.
- The commentators appreciate the opportunity to offer perspectives on this important problem and look forward to methods that overcome challenges identified by KS.