Source-linked AI summary

Bootstrap Inference when Using Multiple Imputation

Michael Schomaker, Christian Heumann

arXiv:1602.07933v6stat.ME

TL;DR

Valid bootstrap inference after multiple imputation remains unclear, especially when analytic standard errors are unavailable or parameter distributions are non-normal. The paper introduces four combined methods and finds that three yield valid inference, with performance varying by imputed-dataset count and missingness.

  • Problem

    It is unclear how to obtain valid bootstrap confidence intervals after multiple imputation when analytic standard errors are unavailable or parameter distributions are non-normal.

  • Method

    The paper introduces four intuitively appealing methods that combine bootstrap estimation with multiple imputation for confidence-interval construction.

  • Results

    Both Boot MI and MI Boot are probably the best options for randomization-valid confidence intervals, with preference depending on imputation count and uncertainty.

  • Takeaways & Limitations

    Boot MI may be preferred for small M or high imputation uncertainty, whereas MI Boot may suit normal M and little or normal uncertainty.

  • Takeaways & Limitations

    Boot MI is computationally much more intensive than MI Boot, increasing computation time by a factor of 13 in the first simulation.

Abstract

from arXiv · show

Many modern estimators require bootstrapping to calculate confidence intervals because either no analytic standard error is available or the distribution of the parameter of interest is non-symmetric. It remains however unclear how to obtain valid bootstrap inference when dealing with multiple imputation to address missing data. We present four methods which are intuitively appealing, easy to implement, and combine bootstrap estimation with multiple imputation. We show that three of the four approaches yield valid inference, but that the performance of the methods varies with respect to the number of imputed data sets and the extent of missingness. Simulation studies reveal the behavior of our approaches in finite samples. A topical analysis from HIV treatment research, which determines the optimal timing of antiretroviral treatment initiation in young children, demonstrates the practical implications of the four methods in a sophisticated and realistic setting. This analysis suffers from missing data and uses the $g$-formula for inference, a method for which no standard errors are available.

1. Introduction.

The introduction identifies unresolved questions about valid bootstrap confidence intervals after multiple imputation, especially when analytic standard errors are unavailable or parameter distributions are non-normal. The article addresses this gap by introducing and evaluating four bootstrap confidence-interval methods for multiply imputed data.

  • Motivation: Valid bootstrap inference after multiple imputation remains unclear when analytic standard errors are unavailable or parameter distributions are non-normal.The concern is specifically how to obtain appropriate confidence intervals using nonparametric bootstrap estimation with missing data.
  • Methods: The proposed approaches either bootstrap each of M imputed data sets or create B bootstrap samples before imputing and combining results.Intervals can be based on MI combining rules or empirical percentiles from pooled estimates.
  • Research gap: Prior literature used bootstrap methods with missing or imputed data but did not study construction of bootstrap confidence intervals after multiple imputation.Existing analyses often provided little justification for their chosen method or incomplete implementation details.
  • Contribution: The article introduces four intuitively appealing bootstrap confidence intervals for data requiring multiple imputation and argues which methods should be preferred.It presents the methods as a novel contribution and illustrates their intrinsic features.
  • Evaluation: The methods are examined through theoretical considerations, numerical investigations, and a causal-inference analysis in HIV research.The motivating application concerns antiretroviral treatment research, while the paper’s sections evaluate and emphasize implications of the different approaches.

2. Motivation.

Earlier ART initiation guidelines make detailed randomized evaluation of treatment rules unethical, motivating observational estimation. For dynamic treatment rules, the g-computation formula compares treatment options but is computationally intensive and typically relies on non-parametric bootstrap confidence intervals, despite missing data in resource-limited settings.

  • Motivation: Earlier ART guidelines make trials evaluating treatment-initiation rules in detail ethically infeasible, so observational data are needed for relevant estimates.The treatment rules affect mortality, morbidity, and child development outcomes.
  • Motivation: For dynamic treatment rules based on time-varying variables such as CD4 lymphocyte count, the g-computation formula is an intuitive method for comparing treatment options.The method is computationally intensive.
  • Motivation: Confidence intervals for the g-computation formula are typically based on non-parametric bootstrap estimation, while resource-limited settings may have missing data for administrative, logistic, and clerical reasons.The passage identifies missingness as a practical complication of applying the method.
  • Motivation: The inferential target may include regression coefficients, odds ratios, factor loadings, or counterfactual outcomes estimated from data containing an outcome and covariates.The setup defines θ = (θ1, . . . , θk)′ with k ≥1.

3. Methodological Framework.

This section introduces four intuitive bootstrap–multiple-imputation approaches for constructing confidence intervals when analytic covariance estimates are unavailable, then motivates simulation-based comparison of their validity and intrinsic interval behavior.

  • Multiple imputation: Multiple imputation creates M augmented data sets by replacing missing values using posterior predictive draws, with estimates and variances combined across imputations.Reliable variance estimation requires M not to be too small.
  • Motivation: Missing data make bootstrap confidence intervals unclear because the target estimate comprises M estimates from separate imputed data sets.The problem arises when no analytic or ideal covariance estimate is available, including g-method treatment-effect estimation with time-varying confounding.
  • Four methods: The four proposed approaches differ in whether bootstrapping precedes or follows imputation and whether bootstrap estimates are pooled or combined through multiple-imputation variance rules.They are MI Boot (pooled sample [PS]), MI Boot, Boot MI (pooled sample [PS]), and Boot MI.
  • Four methods: MI Boot draws B bootstrap samples within each of M imputed data sets, whereas MI Boot estimates within-imputation standard errors or covariances for standard multiple-imputation inference.The latter yields M point estimates and M bootstrap-based standard errors.
  • Four methods: Boot MI draws B bootstrap samples with missing data and imputes each sample M times, either pooling all resulting estimates or combining estimates within bootstrap samples.Both variants apply the corresponding percentile or multiple-imputation construction to the resulting estimates.
  • Simulation plan: The methods are straightforward to implement, but their interval coverage validity is initially unclear, motivating Monte Carlo simulations across four increasingly complex settings.The settings include simple, dependent-variable, survival-analysis, and longitudinal time-dependent-confounding scenarios.

4. Simulation Studies.

Across four simulation settings, the three bootstrap–multiple-imputation approaches generally produced valid confidence-interval inference, with performance affected by missingness, pooling, and the number of imputations. MI Boot was substantially faster than Boot MI, while both methods showed exceptions under high missingness in Setting 2.

  • Computational cost: Boot MI computation time exceeded MI Boot time by a factor of 13 in Setting 1 and 1.3 in Setting 4.The relative computational disadvantage of Boot MI therefore varied substantially across simulation settings.
  • Simulation results: Point estimates for β were approximately unbiased in all settings, while no bootstrapping achieved about 95% coverage for all parameters and settings.The analytic standard errors served as the gold-standard reference for the first three settings.
  • Simulation results: MI Boot and MI Boot [PS] achieved about 95% coverage with similar interval widths, except under high missingness in Setting 2.MI Boot’s simulated and mean estimated standard errors were almost identical, indicating good standard-error estimation.
  • Simulation results: Boot MI and Boot MI [PS] generally produced coverage close to nominal levels, except under high missingness in Setting 2; pooled samples tended toward higher coverage and slightly different widths.Boot MI could perform well with M < 5, whereas the pooled approach tended toward coverage probabilities > 95%; with M = 1, Boot MI coverage was too large in the illustrated setting.
  • Number of imputations: Multiple imputation generally required a reasonable number of imputed data sets to perform well, regardless of whether bootstrapping was used.This pattern was consistent with multiple-imputation theory.
  • Bootstrap distributions: Boot MI [PS] produced a somewhat wider estimator distribution than MI Boot [PS], explaining the larger confidence interval in the first simulation setting.The comparison visualized bootstrap distributions within imputed data sets and estimator distributions across bootstrap samples.

5. Data Analysis.

The data analysis compares immediate ART initiation with a CD4-based ART strategy using the g-formula to estimate cumulative mortality differences. Three-year mortality under immediate ART initiation was estimated at 6.08%.

  • Data Analysis.: The analysis evaluates mortality over 3 years using treatment cohort collaborations from IeDEA-SA and IeDEA-WA.It builds on a recently published analysis by Schomaker et al.
  • Data Analysis.: The comparison examines immediate ART initiation versus assigning ART when CD4 count < 350 cells/mm3 or CD4% < 15%.The g-formula estimates the cumulative mortality difference between these strategies.
  • Data Analysis.: 6.08%: Three-year mortality was estimated at 6.08% for immediate ART initiation.The supplied passage reports this estimate but is truncated before reporting the comparator estimate.

MI Boot (PS) · Boot MI (PS)

In the HIV treatment analysis, Boot MI produced the shortest confidence interval and uniquely excluded zero, suggesting a beneficial effect of immediate ART initiation. Boot MI (PS) and MI Boot (PS) had nearly identical intervals because their overall estimate variation was similar, while MI Boot reflected substantial between-imputation uncertainty.

  • Boot MI (PS): 0.79% mortality difference was estimated between immediate treatment initiation and the delayed-treatment strategy.The delayed strategy’s estimated mortality was 6.87%.
  • MI Boot (PS): [−0.34%; 1.61%] was the confidence interval for Boot MI (PS), compared with [−0.31%; 1.63%] for MI Boot (PS).These intervals were nearly identical, consistent with similar overall estimate variation under the two approaches.
  • Boot MI (PS): The confidence intervals were [0.12%; 1.07%] for Boot MI and [−0.81%; 2.40%] for MI Boot.The other methods produced larger intervals and did not necessarily indicate a clear mortality difference.
  • Boot MI (PS): Boot MI produced the shortest confidence interval and was the only method whose 95% interval excluded 0%.It therefore suggested a beneficial effect of immediate treatment initiation for the estimated mortality difference.
  • Boot MI (PS): The distributions for Boot MI (PS), MI Boot (PS), and Boot MI were reasonably symmetric.The figure visualized the distributions of b,m for Boot MI (PS) and MI Boot (PS), and ˆ¯θ∗b for Boot MI.
  • MI Boot (PS): The overall variation of estimates was similar for MI Boot (PS) and Boot MI (PS).This similarity explains why their confidence intervals were almost identical.
  • MI Boot (PS): Large between-imputation uncertainty affected the MI Boot estimator.The point estimates varied substantially, possibly because of high missingness and a complex imputation procedure, producing MI Boot’s large confidence interval.
  • Boot MI (PS): Method 3, Boot MI, suggested a beneficial effect of immediate ART initiation versus delaying ART until CD4 count < 350 cells/mm3 or CD4% < 15%.The other methods produced larger confidence intervals and did not necessarily suggest a clear mortality difference.

6. Theoretical Considerations. · MI Boot (PS) · Boot MI (PS)

The theoretical discussion distinguishes how bootstrap and multiple-imputation uncertainty are combined, identifying three valid approaches and warning that Boot MI (PS) is inefficient and invalid. It also emphasizes that imputation bootstrapping does not replace the separate bootstrap required for analysis-model inference.

  • MI Boot (PS): MI Boot estimates bootstrap variance separately within each imputed dataset and then applies multiple-imputation combining rules.Its validity depends on the bootstrap variance being close to the corresponding analytical variance.
  • MI Boot and MI Boot (PS).: MI Boot and MI Boot pooled require reasonably large M, because variance estimates can be sensitive to the number of imputations.The discussion notes that M should often be much larger than 5 for good variance estimates.
  • MI Boot and MI Boot (PS).: Boot MI estimates the target in each bootstrap sample using multiple imputation, approximating its distribution validly under missing at random.This approach combines multiple imputation for each bootstrap sample with bootstrapping to estimate the target’s distribution.
  • Boot MI and Boot MI (PS).: Pooling all B × M Boot MI (PS) estimates effectively treats each as a separate estimator and is statistically inefficient.This corresponds to using multiple imputation with M = 1 for each bootstrap replicate.
  • Boot MI and Boot MI (PS).: 5% loss of efficiency occurs when the fraction of missingness is 0.25 and M = 5; lower M further increases variance and can produce incorrect coverage.Pooling estimates is inefficient, overestimates variance, and leads to confidence intervals with incorrect coverage.
  • Comparison.: Boot MI (PS) typically produces larger intervals than Boot MI, while Boot MI with M = 1 is not incorrect but often inefficient.General interval-width comparisons with MI Boot are difficult because multiple sources of imputation and bootstrap uncertainty contribute.
  • Bootstrapping as part of the imputation procedure.: Bootstrap procedures used inside imputation algorithms do not replace the additional bootstrap needed for analysis-model inference; when combined, the resampling steps are nested.Proper imputations are generated through random posterior-predictive draws, while the inferential bootstrap is a separate operation.
  • Bootstrapping as part of the imputation procedure.: Three approaches—Boot MI, MI Boot, and MI Boot (PS)—are correct, whereas Boot MI (PS) can yield overly large, invalid confidence intervals and is not recommended.The paper’s overall recommendation is based on the theoretical and simulation results.

7. Conclusion. · APPENDIX A: DETAILS OF THE G-FORMULA IMPLEMENTATION

The paper recommends Boot MI and MI Boot as the strongest options for randomization-valid confidence intervals, with their relative advantages depending on imputation conditions and computational or distributional considerations. The appendix defines the longitudinal g-formula framework and describes its application to 5,826 children followed for 36 months.

  • 7. Conclusion.: Boot MI and MI Boot are probably the best options for randomization-valid confidence intervals combining bootstrapping with multiple imputation.Boot MI may be preferred for small M or large imputation uncertainty, whereas MI Boot may suit normal M and little or normal imputation uncertainty.
  • 7. Conclusion.: 13-fold greater computation time made Boot MI substantially more intensive than MI Boot when the analysis model was simple relative to creating imputations.MI Boot naturally produces symmetrical confidence intervals, which may be undesirable when the estimator distribution is suspected to be non-normal.
  • APPENDIX A: DETAILS OF THE G-FORMULA IMPLEMENTATION: The g-formula framework represents longitudinal outcomes, interventions, time-dependent covariates, baseline variables, and censoring across baseline and discrete follow-up times.Treatment and covariate histories are represented explicitly, while remaining uncensored through time t is denoted by C̄_t = 0.
  • APPENDIX A: DETAILS OF THE G-FORMULA IMPLEMENTATION: The appendix distinguishes static treatment rules from dynamic rules that assign treatment as a function of each individual’s covariate history.Rules may also intervene on multiple variables, including censoring, without being written explicitly as separate interventions.
  • APPENDIX A: DETAILS OF THE G-FORMULA IMPLEMENTATION: 5,826 children were studied at baseline and follow-up times through 36 months, with death as the outcome and ART as the intervention.Time-dependent covariates included CD4 count, CD4%, WAZ, and indicators of whether these measurements were observed; baseline variables included CD4 measures, HAZ, sex, age, and region.
  • APPENDIX A: DETAILS OF THE G-FORMULA IMPLEMENTATION: The target quantity was cumulative mortality after 36 months under no censoring, regular 3-month follow-up, and treatment assignment according to the intervention rules.Identification relies on consistency, sequential conditional exchangeability, and positivity.
  • APPENDIX A: DETAILS OF THE G-FORMULA IMPLEMENTATION: Because no closed-form solution exists, the g-formula is approximated by fitting additive linear and logistic models and stochastically simulating covariates and death forward in time.For one rule, treatment at time 1 is assigned when generated CD4 count is below 350 cells/mm3 or CD4% is below 15%, followed by evaluation of 3-year cumulative mortality.
  • APPENDIX A: DETAILS OF THE G-FORMULA IMPLEMENTATION: The sequential g-formula used in simulations re-expresses the g-formula through sequential marginalization over covariate distributions under the intervention rule.This representation avoids integration with respect to L.

APPENDIX B: DATA GENERATING PROCESS IN THE SIMULATION STUDY

The simulation generated longitudinal baseline and follow-up data with structural equations, treatment and censoring indicators, and specified target quantities under treatment rules. Missingness was imposed through functions depending on baseline or follow-up outcomes.

  • Data-generating process: Structural equations generated baseline variables at t = 0 and follow-up CD4 count, CD4%, WAZ, HAZ, treatment, and censoring through t = 12.No deaths were assumed, and the process used simcausal along with Bernoulli, uniform, normal, and truncated normal distributions.
  • Generated values: The generated baseline data had region A = 75.5%, male sex = 51.2%, mean age = 3.0 years, mean CD4 count = 672.5, mean CD4% = 15.5%, mean WAZ = -1.5, and mean HAZ = -2.5.At t = 12, arithmetic means were CD4 count = 1092, CD4% = 27.2%, WAZ = -0.8, and HAZ = -1.5.
  • Target quantities: The target quantities ψ1 and ψ2 were expected Y values at time T under no censoring for specified treatment rules, with values −1.03 and −2.45 respectively.The passage defines the targets through treatment-rule interventions and reports these two resulting quantities.
  • Missingness mechanism: Missing baseline and follow-up data were generated using functions based on outcome magnitude, with follow-up missingness additionally depending on time.The supplied equations show baseline dependence on |Y0| and follow-up dependence on t and |Yt|.
Loading 1602.07933v6…