Source-linked AI summary
Multiple Imputation: A Review of Practical and Theoretical Findings
Jared S. Murray
TL;DR
Missing-data analyses need methods that support principled inference while accommodating uncertainty in both missing values and imputation models. This paper reviews MI theory, generation strategies, and empirical comparisons, concluding that flexible models are generally preferable to simple defaults, while further comparative evidence is needed.
Problem
Multiple imputation requires guidance for selecting and critiquing imputation procedures, while the comparative merits of newer imputation models remain relatively little known.
Method
The paper reviews MI mechanics, theoretical validity conditions, practical implications, imputation strategies, and empirical comparisons across model types.
Results
Theoretical and empirical evidence generally favors flexible, adaptive imputation models over simple default MVN, log-linear, or basic FCS specifications where possible.
Takeaways & Limitations
Practitioners should consider flexible imputation models rather than relying automatically on simple defaults, while comparing methods across realistic data settings.
Takeaways & Limitations
MI inference can be invalid under model misspecification, uncongeniality, or insufficiently proper imputations, and the review notes that further empirical comparisons are needed.
Abstract
from arXiv · showhide
Multiple imputation is a straightforward method for handling missing data in a principled fashion. This paper presents an overview of multiple imputation, including important theoretical results and their practical implications for generating and using multiple imputations. A review of strategies for generating imputations follows, including recent developments in flexible joint modeling and sequential regression/chained equations/fully conditional specification approaches. Finally, we compare and contrast different methods for generating imputations on a range of criteria before identifying promising avenues for future research.
1. INTRODUCTION
Multiple imputation addresses missing data by creating completed datasets, analyzing each, and pooling results while incorporating missing-data variability. This review focuses on imputation methods, their theoretical and empirical evaluation, and practical guidance for selecting procedures.
- Multiple imputation: Multiple imputation fills missing values by sampling from an imputation model to create a small number of completed datasets.Analysts estimate quantities in each completed dataset and combine them using rules that incorporate additional variability from missing data.
- Multiple imputation: In disseminated databases, MI supports approximately valid inference across many analyses while shifting missing-data handling from end-users to the imputer.Using one imputation set across analyses also ensures result differences are not attributable to different missing-data treatments.
- Developments and scope: User-created imputations have become more common as software has simplified imputation and pooling procedures.Available methods now range from parametric and resampling models to tree-based algorithms and flexible Bayesian nonparametric models.
- Review scope: The review emphasizes generating imputations and evaluating theoretical results and empirical evidence that guide the selection and critique of imputation procedures.It restricts attention to item missing data with independent observations, while noting broader applicability.
- Review scope: The paper reviews MI mechanics, inference conditions, practical model-selection implications, methods for one or multiple incomplete variables, model-choice considerations, and future discussion.These topics are organized across Sections 2 through 8.
2. MULTIPLE IMPUTATION: HOW DOES IT WORK?
MI generates multiple completed datasets under assumptions about the missing-data process, computes an estimate and variance in each, and combines them for inference on a scalar estimand. The framework separates within-imputation uncertainty from between-imputation uncertainty and uses repeated imputations to approximate the relevant quantities.
- Data notation: For each unit, Y records a p-dimensional vector of values and R records indicators showing which values are observed or missing.Rij equals 1 when Yij is observed and 0 otherwise; Yobs and Ymis collect the observed and missing values.
- Missingness assumptions: MI assumes missing at random, under which the response process need not be explicitly modeled for imputation.For non-MAR data, MI requires explicitly modeling the response mechanism or making other identifying assumptions.
- Combining imputations: For scalar Q, analysts compute an estimator and its variance in each of M completed datasets, then average the estimates to obtain the MI estimate.The completed datasets differ through imputations for Ymis, while each analysis produces Q-hat(m) and U-hat(m).
- Combining imputations: The MI variance estimator decomposes uncertainty into within-imputation variance and between-imputation variance, with a small-M bias adjustment.The within component estimates complete-data variance, while the between component estimates excess variance due to missing values.
- Inference: The Bayesian derivation of MI begins by decomposing uncertainty through the distribution of missing values conditional on observed values.The displayed identities motivate separating uncertainty attributable to missing data from uncertainty conditional on completed data.
- Inference: MI statistics are Monte Carlo estimates of the relevant quantities when imputations are generated from P(Ymis | Yobs).Rubin’s large-sample finite-M inference uses a reference t distribution, with degrees of freedom depending on M and the relative variance increase from nonresponse.
3. MULTIPLE IMPUTATION: WHEN DOES IT WORK?
MI’s validity depends on the relationship between the imputation and analysis models, proper-imputation conditions, and the validity of complete-data inference. These conditions may fail when model uncertainty is ignored or when variance estimates are inconsistent.
- Bayesian (in)validity under MI: When the imputer and analyst use the same missing-data model, MI draws can reproduce the analyst’s posterior, provided the posterior is approximately normal and M is not too small.Under these conditions, MI statistics reasonably approximate posterior inference.
- Bayesian (in)validity under MI: MI generally does not deliver valid Bayesian inference when the imputer’s and analyst’s models for missing data differ.A single set of imputations cannot generally support valid Bayesian inference for analysts with different beliefs about the missing-data distribution.
- Frequentist Validity: Conditions on complete data inference: Complete-data inference must itself be confidence valid, typically requiring a suitably normal sampling distribution and valid relationships between estimator bias and variance.MI cannot repair invalid complete-data inference; normality and validity conditions may hold only asymptotically or under particular modeling assumptions.
- Proper imputation for valid inference: Proper imputations combined with valid complete-data inference yield valid MI inference when M = ∞, but propriety is specific to an estimand and posited response mechanism.The paper emphasizes that imputations are not universally proper independently of the target estimand or response mechanism.
- Three essential conditions for proper imputation.: Correct specification of the missing-data model is sufficient for key propriety conditions, but misspecified models can still work when they capture features relevant to Q and U and missingness is not extreme.With modest missingness, observed data can dominate the effect of imperfect imputations unless the imputed values are sufficiently poor.
- Three essential conditions for proper imputation.: The between-imputation variance must reflect uncertainty in the imputation model; otherwise procedures such as MLE plug-in or empirical-distribution imputation may be improper.Model uncertainty can be incorporated by posterior parameter sampling or suitable bootstrap adjustments.
- Proper imputation for valid inference: MI’s variance estimate can be inconsistent for some estimands, although the resulting positive bias often has limited coverage impact when missingness is not extreme.The paper relates this issue to the broader concept of congeniality between analysis procedures and imputation models.
2. It matches the imputation model, i.e.,
MI inference depends on how the imputation and analysis models relate. Under congeniality, inference is confidence valid; under uncongeniality, validity depends on model saturation and may require adjusted variance estimates.
- Under congeniality, MI delivers samples from PA(Q | Yobs) that support confidence-valid inference.
- Even when the true model is nested within both models, standard MI inference may be invalid under uncongeniality.
- When the imputer’s model is more saturated than the analyst’s, usual MI inference is confidence valid and generally robust.
- When the imputer’s model is less saturated than the analyst’s, confidence validity is not guaranteed.
- Under the more-saturated imputer model, uncongenial analyses are generally safer because conservative inferences obtain.
- These theoretical results assume that the true model is nested within the imputation model class.
4. PRACTICAL IMPLICATIONS OF THEORETICAL RESULTS FOR IMPUTATION MODELING
Proper imputation must represent uncertainty about both missing values and the imputation model, while reflecting relevant features of the data. Practical model choices should be broad, predictive, and adaptable, but misspecification can still harm inference for sensitive estimands.
- MI targets valid inference about missing-data uncertainty rather than merely estimating or predicting the missing values.
- Proper imputations propagate intrinsic missing-value uncertainty and uncertainty about the imputation model.
- Misspecified imputation models can harm imputations and inference, especially for estimands sensitive to misspecification.
- In large samples, even small misspecification biases can become large relative to pooled standard errors.
- Including more completely observed variables can make the missing-at-random assumption more tenable, while omitted predictors can produce improper imputations and bias.
- Imputation models should track relevant joint-distribution features, including interactions, nonlinearities, and non-standard distributions.
5. GENERATING IMPUTATIONS FOR A SINGLE VARIABLE
Single-variable imputation methods range from familiar regression models and donor-based approaches to adaptive tree methods. Proper imputations must reflect parameter or model uncertainty, while method choice trades assumptions, flexibility, and tuning guidance.
- Regression imputation samples missing values from conditional models, including generalized linear models and extensions for zero inflation or truncation.
- Proper imputations require accounting for parameter uncertainty; fixing parameters at the observed-data MLE is generally improper.Posterior sampling under weakly informative priors is typically proper when the model fits well.
- Hot Deck/Nearest Neighbor Methods: Hot-deck methods borrow observed donor values within adjustment cells defined by cross-classifications, preserving plausible realized values but becoming difficult with many variables.
- Hot Deck/Nearest Neighbor Methods: Standard hot-deck MI is improper for a population mean because it ignores uncertainty in the implicit empirical imputation model.The approximate Bayesian bootstrap modifies donor sampling to produce proper imputations for the adjustment-cell mean.
- Hot Deck/Nearest Neighbor Methods: Predictive mean matching selects donors using distances between predicted means, generalizing the hot deck and accommodating more variables while remaining sensitive to model specification.
- Hot Deck/Nearest Neighbor Methods: CART assigns incomplete cases to leaves of a tree grown from complete cases and samples donors within leaves, with minimum leaf size controlling donor-pool size.Bootstrap-based trees, ABB sampling, and random forests are proposed to address donor-pool and tree uncertainty.
- Hot Deck/Nearest Neighbor Methods: Recursive partitioning methods can be fast and effective, especially for large sets of categorical variables, but comparative evidence and tuning guidance remain limited.
6. GENERATING IMPUTATIONS FOR MULTIPLE VARIABLES
Multivariate missing data can be handled with joint models or collections of univariate conditional models. The review covers parametric, mixture, Bayesian nonparametric, FCS, and sequential approaches, emphasizing coherence, flexibility, and computational trade-offs.
- Two basic strategies are joint modeling of missing variables and fully conditional specification using univariate models conditioned on the remaining variables.
- Joint specifications: Parametric joint models include multivariate normal, t, multinomial, log-linear, latent-class, and correspondence-analysis approaches, but often impose restrictive assumptions.
- Flexible joint models: Bayesian nonparametric models based mainly on infinite mixtures can increase complexity with sample size and have been extended to mixed data and constrained supports.
- Fully Conditional Specification: Under compatible conditional models, FCS and joint algorithms agree in finite samples under factorized priors, or asymptotically with a unique stationary distribution when priors do not factorize.
- Fully Conditional Specification: Finite-sample FCS inference may be inefficient when nonfactorized priors contain indirect parameter information that FCS ignores.
- Joint specifications: Sequential approach: Sequential specifications always define a coherent joint model when each univariate model is proper, but different variable orderings can produce different joint distributions and fits.
- Joint specifications: Sequential approach: Early variables in a sequential model may require flexible distributions because marginalizing over related covariates can produce complicated joint patterns.The review illustrates this issue with householder age and earnings distributions stratified by children in the household.
7. CHOOSING AND ASSESSING AN IMPUTATION STRATEGY
Choosing an imputation strategy requires balancing implementation, uncertainty representation, flexibility, computational cost, and empirical fit. The review favors flexible models while stressing convergence checks, realistic evaluations, and caution about limited comparative evidence.
- FCS convergence remains largely unresolved in general settings, particularly for non- or quasi-Bayesian procedures such as PMM.
- Current convergence guarantees for FCS rely on restrictive implicit joint models that may be too simple for realistic data.
- 7.2 Practical considerations derived from MI theory: Most reviewed methods include mechanisms for reflecting imputation-model uncertainty, but tree-based methods still need tuning and resampling strategies that reliably produce proper imputations.
- 7.2 Practical considerations derived from MI theory: Joint-sequential models may be easier to fit than FCS with many covariates because most univariate models contain fewer than p predictors.
- 7.2.2 Include as many variables as possible.: Non- and semiparametric methods can capture unanticipated data features and outperform default MI procedures in simulations, especially beyond simple parametric settings.The review calls for more realistic evaluations.
- Flexible imputation models make it harder to incorporate prior information manually or diagnose and correct model misfit.
- 7.3 Empirical comparisons between methods: Empirical comparisons are relatively rare because new models are commonly assessed on synthetic data generated from researcher-specified probability models.
- 7.3 Empirical comparisons between methods: Repeated-sampling studies from realistic populations evaluate bias, confidence-interval coverage, and efficiency under known missingness mechanisms.
8. CONCLUSION
The review concludes that multiple imputation is principled and practical, but the comparative merits of newer imputation models remain insufficiently understood. It supports flexible, carefully scrutinized models and identifies major needs in computation, theory, ordering, and realistic evaluation.
- Multiple imputation has become a principled and practical solution to missing-data problems, while comparative evidence for newer models remains limited.
- Theoretical considerations favor flexible imputation models, and the review advises avoiding or carefully scrutinizing simple default MVN, log-linear, and main-effects FCS procedures.
- Nonparametric Bayesian methods are promising but need scalable posterior computation and renewed theoretical study of why Bayesian MI tends to be proper.
- Joint-sequential approaches are understudied, including the consequences of variable ordering and algorithms for selecting useful orderings.
- FCS theory remains limited in scope and does not cover important effective variants such as PMM and CART.
- The field needs more empirical comparisons and a repository of realistic synthetic populations with shared incomplete samples for evaluating methods consistently.
- MI applications extend beyond item missingness to synthetic-data disclosure limitation, measurement-error adjustment, and statistical matching or data fusion.
APPENDIX A: SOFTWARE FOR MULTIPLE IMPUTATION
An online resource provides pointers to software implementations of multiple-imputation methods, while its December 2017 version lacked links for several nonparametric Bayesian joint-model packages.
- Software pointers for many multiple-imputation methods are available through an online resource.The resource is described as an updated version of Appendix A of Van Buuren (2012).
- As of December 2017, the resource was missing links to R packages for several nonparametric Bayesian joint models.
- The missing links included packages for imputing mixed continuous and categorical values and multivariate categorical data.