Source-linked AI summary
Multiple imputation of covariates by fully conditional specification: accommodating the substantive model
Jonathan W. Bartlett, Shaun R. Seaman, Ian R. White, James R. Carpenter
TL;DR
Missing covariate data are difficult to impute when the substantive model is nonlinear or contains interactions, because standard imputation models may be incompatible with it. The paper modifies fully conditional specification so each covariate model is compatible with the substantive model, and reports simulation evidence supporting consistent estimation and nominal confidence-interval coverage under the stated conditions.
Problem
Missing covariate data are common, but standard imputation models may be incompatible with nonlinear or interactive substantive models.
Method
The paper modifies fully conditional specification so that each covariate is imputed from a model compatible with the substantive model.
Results
The authors conjecture consistency, supported by simulations, with confidence-interval coverage attaining nominal levels in the studied cases.
Takeaways & Limitations
The proposed approach supports compatible covariate imputation across substantive models while avoiding the need to specify the full joint outcome-covariate distribution.
Takeaways & Limitations
The approach relies on missing-at-random data and correctly specified, mutually compatible imputation models; compatibility alone does not guarantee unbiased estimates.
Abstract
from arXiv · showhide
Missing covariate data commonly occur in epidemiological and clinical research, and are often dealt with using multiple imputation (MI). Imputation of partially observed covariates is complicated if the substantive model is non-linear (e.g. Cox proportional hazards model), or contains non-linear (e.g. squared) or interaction terms, and standard software implementations of MI may impute covariates from models that are incompatible with such substantive models. We show how imputation by fully conditional specification, a popular approach for performing MI, can be modified so that covariates are imputed from models which are compatible with the substantive model. We investigate through simulation the performance of this proposal, and compare it to existing approaches. Simulation results suggest our proposal gives consistent estimates for a range of common substantive models, including models which contain non-linear covariate effects or interactions, provided data are missing at random and the assumed imputation models are correctly specified and mutually compatible.
1 Introduction
Missing covariate data complicate analyses when imputation models conflict with nonlinear or interactive substantive models. The paper modifies fully conditional specification to make univariate imputation models compatible with the substantive model and evaluates the approach through simulation.
- Missing data are pervasive in medical research, and multiple imputation is commonly used to accommodate them.
- Fully conditional specification specifies separate regression models for each partially observed variable, allowing different models for continuous and binary variables.This flexibility makes FCS attractive when several variables with mixed measurement types are incomplete.
- Default covariate imputation models may be incompatible with substantive models containing nonlinear effects or interactions, causing biased parameter estimates.For a quadratic substantive model, a default linear imputation model for X|Y imposes linearity among imputed values, unless the quadratic coefficient is truly zero.
- Joint modeling can avoid incompatibility, but specifying a joint model is challenging with multiple partially observed continuous and discrete covariates.
- The proposed modification of FCS ensures that each univariate imputation model is compatible with the assumed substantive model.The paper develops the approach for normal linear, discrete-outcome, and proportional-hazards substantive models, then studies its performance by simulation.
- The approach is developed for settings analyzed under the missing-at-random assumption and is assessed using multiple completed datasets combined with Rubin’s rules.
2 Motivating example
The motivating example uses data from the Alzheimer’s Disease Neuroimaging Initiative, a public-private study of biomarkers and disease progression. Participants included cognitively normal controls, people with mild cognitive impairment, and people with early Alzheimer’s disease followed with clinical, cognitive, and MRI assessments.
- The Alzheimer’s Disease Neuroimaging Initiative was launched in 2003 as a 5-year public-private partnership.
- ADNI aimed to assess whether imaging and other biomarkers measure progression of mild cognitive impairment and early Alzheimer’s disease.
- The study aimed to recruit approximately 200 cognitively normal controls, 400 participants with mild cognitive impairment, and 200 with early Alzheimer’s disease.
- Participants received clinical and cognitive assessments plus MRI scans at baseline and specified intervals for up to 3 years.Assessments occurred every 6 or 12 months depending on the participant group.
3 Multiple imputation of partially observed covariates
The section introduces MI for partially observed covariates, explains how incompatibility with substantive models can cause mis-specification and inconsistency, and motivates compatible imputation models.
- 3.1 Setup: The setup includes a fully observed outcome Y, partially observed covariates X, fully observed covariates Z, and missingness indicators R under MAR.The substantive model f(Y|X,Z,ψ) is assumed correctly specified.
- 3.2 Multiple imputation of partially observed covariates: Bayesian MI draws imputation-model parameters from their posterior, imputes missing X values, fits the substantive model to each dataset, and combines estimates with Rubin’s Rules.When the imputation model is correctly specified, the MI estimator is consistent and confidence intervals can achieve nominal or higher coverage.
- 3.3 Compatibility and imputation model mis-specification: Default covariate imputation models can be incompatible with substantive models containing quadratic effects, interactions, or nonlinear survival effects.Such incompatibility can make the imputation model mis-specified and the MI estimator inconsistent, although bias may sometimes be small.
- 3.3 Compatibility and imputation model mis-specification: Compatible conditional models are those arising from a common joint model, while semi-compatible models can become compatible after restricting parameters.Incompatibility does not always imply mis-specification when a compatible restricted version is correctly specified.
- 3.4 Joint model imputation: The proposed strategy constructs a joint model retaining the substantive conditional and derives compatible imputation models, then factors X|Z into univariate conditional densities.This motivates modifying FCS so each univariate imputation model is compatible with the substantive model.
- 3.4 Joint model imputation: Specifying f(X|Z,δ) remains challenging for multivariate X, especially when continuous and discrete covariates are mixed.Compatibility alone does not guarantee correct specification; f(X|Z,δ) must also be correctly specified.
- 3.5 Fully conditional specification imputation: The paper proposes modifying FCS so each univariate imputation model is compatible with the substantive model.The modification is motivated by incompatibility between standard FCS imputation models and substantive models.
4 Review of fully conditional specification multiple imputation
Standard FCS MI repeatedly imputes each partially observed covariate conditional on the others, but its conditional models may be incompatible and need not yield Bayesian MI or valid variance inference.
- 4.1 The fully conditional specification algorithm: For each partially observed covariate Xj, FCS specifies an imputation model conditional on the other covariates, fully observed variables, and outcome.The covariate type determines the usual generalized linear model choice, with a non-informative prior for its parameters.
- 4.1 The fully conditional specification algorithm: FCS initializes missing values with observed values and repeatedly updates each variable using the latest imputations of the other variables.After convergence, final missing-data draws form one imputed dataset, and the process is repeated to create multiple datasets.
- 4.1 The fully conditional specification algorithm: Each update first draws imputation-model parameters from a posterior based on subjects with the covariate observed, then draws missing values using those parameters.The substantive model is fitted to every imputed dataset and results are combined using Rubin’s Rules.
- 4.2 Statistical properties: FCS is not equivalent to Bayesian MI even when conditional models are compatible because it may fail to use information about conditional-model parameters in marginal distributions.Compatibility alone is therefore insufficient for equivalence with a Bayesian joint model.
- 4.2 Statistical properties: When conditional models are correctly specified and semi-compatible, the MI estimator is consistent.If models are not semi-compatible, inconsistent estimates are generally expected, and the algorithm’s limiting distribution may be unclear.
- 4.2 Statistical properties: Because FCS need not correspond to Bayesian MI, Rubin’s variance rules are not guaranteed to provide valid inferences.This limitation applies even when the substantive parameter estimator is consistent.
5 Fully conditional specification imputation accommodating the sub-
SMC-FCS modifies FCS so each covariate imputation density is compatible with the substantive model, while its statistical validity depends on compatibility, correct specification, and convergence.
- 5.1 The SMC-FCS algorithm: The proposed substantive model compatible FCS (SMC-FCS) modifies standard FCS to make every univariate imputation model compatible with the substantive model.The method specifies a substantive-model prior and covariate models, then uses a compatible density for imputing each missing covariate.
- 5.1 The SMC-FCS algorithm: SMC-FCS uses all subjects, including those with imputed Xj values, when fitting the covariate model.This is necessary because fitting only observed Xj values could introduce bias when missingness depends on Y.
- 5.1 The SMC-FCS algorithm: The imputation density is constructed to be compatible with f(Y|X,Z,ψ), but it generally does not belong to a standard parametric family.This nonstandard form complicates direct simulation.
- 5.2 Statistical properties: When SMC-FCS corresponds to a well-defined Bayesian joint model, Rubin’s Rules can be applied for inference.Compatibility alone remains insufficient for Bayesian equivalence, including a logistic-normal conditional-model example.
- 5.2 Statistical properties: If the covariate models are semi-compatible and correctly specified, the authors conjecture that the SMC-FCS MI estimator is consistent.Simulation studies investigate this conjecture.
- 5.2 Statistical properties: Rubin’s variance estimator has no guaranteed nominal coverage when SMC-FCS is consistent but not equivalent to Bayesian joint-model imputation.When the covariate models are not semi-compatible, the estimator is generally not expected to be consistent.
- 5.3 Implementation: SMC-FCS may require more iterations than standard FCS because its model fitting uses the most recently imputed covariate values.The algorithm is assumed to converge to a stationary distribution when one exists.
6 Sampling from the imputation model
The paper uses rejection sampling to draw from SMC-FCS imputation densities, deriving bounds and acceptance rules for normal, discrete, and proportional-hazards substantive models.
- 6.1 Rejection sampling: Rejection sampling draws from an easy proposal density until a sampled covariate value satisfies an acceptance condition.Here the proposal is the covariate model f(Xj|X−j,Z,φj), and the target-to-proposal ratio must be bounded.
- 6.1 Rejection sampling: An upper bound c(Y,X−j,Z,ψ) for the substantive density in Xj is used to construct the rejection-sampling draw.Candidate Xj values and uniform random variables are repeatedly sampled until the acceptance inequality holds.
- 6.2 Normal regression models: For normal outcomes, the method samples a candidate missing covariate value and accepts it using the substantive model’s conditional mean structure.The normal-regression case is one of the substantive-model-specific bounds derived in this section.
- 6.3 Discrete outcome models: For discrete outcomes, the substantive conditional density is a probability bounded by one, simplifying the rejection-sampling bound.Binary outcomes modelled by logistic regression are included in this case.
- 6.4 Proportional hazards models: For censored event times, the method models W=min(T,C) and D=1(T<C), assuming noninformative censoring and censoring independent of X given Z.The proportional-hazards substantive model uses baseline hazard H0(t) and covariate effect function g(Xj,X−j,Z,β).
- 6.4 Proportional hazards models: The proportional-hazards derivation treats censored and observed-event subjects separately when constructing the substantive density used for imputation.The model may be parameterized with a finite-dimensional baseline hazard or Cox’s arbitrary baseline hazard.
7 Simulation study
The simulations evaluate imputation methods when standard covariate models are incompatible with quadratic substantive effects, under MCAR and MAR mechanisms. SMC-FCS generally performed well when its substantive-model-compatible imputation assumptions were appropriate, but misspecification could reduce coverage.
- Imputation methods: The compared approaches were linear imputation, JAV, polynomial combination, and SMC-FCS.Polynomial combination imputes the linear predictor combination involving X and X2, then solves a quadratic equation for X.
- Results: Under normal-mixture X with MCAR, JAV and polynomial combination were unbiased with coverage close to 95%, while linear imputation remained severely biased.SMC-FCS was somewhat biased toward the null and had 74% coverage despite its mis-specified marginal model for X.
- Results: Under normal X with MAR, SMC-FCS was unbiased and gave approximately 95% confidence-interval coverage for β2, whereas linear imputation had zero coverage.JAV had 18% coverage and polynomial combination had 75.9% coverage in this setting.
- Results: Under log-normal X with MAR, polynomial combination performed best, with unbiased β2 estimates and somewhat reduced confidence-interval coverage.SMC-FCS was biased because its assumed distribution for X was incorrect, and its coverage was extremely poor under the normal-mixture MAR scenario.
7.2 Linear regression with interaction
The interaction simulations compare standard FCS, JAV, and SMC-FCS across covariate distributions and missingness mechanisms. Standard FCS was biased, while compatible approaches were often unbiased but sensitive to distributional misspecification and nonlinear covariate relationships.
- Study design: The study simulated linear regression outcomes with two interacting covariates across bivariate-normal, log-normal, nonlinear-conditional, and binary–continuous settings.Covariates were made missing under both MCAR and MAR mechanisms.
- Imputation methods: Standard FCS included the outcome, the other covariate, and their interaction in separate linear or logistic imputation models.JAV additionally imputed the interaction variable X1X2 as a separate variable.
- Results: With bivariate-normal covariates under MCAR, JAV and SMC-FCS were unbiased with similar efficiency, while FCS was biased.JAV and SMC-FCS had confidence-interval coverage close to 95%; polynomial combination had slightly low coverage in the related quadratic setting.
- Results: With MAR covariates, SMC-FCS was unbiased for the bivariate-normal setting, whereas complete-case analysis and FCS remained biased.JAV showed a small bias toward zero for β3 and a larger bias for β1.
- Results: When covariate distributions or conditional relationships were misspecified, SMC-FCS generally had smaller biases than JAV but was not uniformly unbiased.SMC-FCS was biased under log-normal or quadratic conditional distributions, while FCS sometimes produced extreme imputations causing collinearity errors.
7.3 Cox proportional hazards models
The Cox simulations assess imputation for partially observed covariates in a proportional-hazards model at sample sizes n = 100 and n = 1,000. SMC-FCS achieved unbiased estimates and correct interval coverage in the reported n = 1,000 setting, while its variance treatment has limitations.
- Study design: Survival times followed a Cox model with two covariates, one binary and one conditionally normal, each observed with probability 0.7 under MCAR.Simulations used n = 100 and n = 1,000 subjects.
- Imputation methods: Standard FCS imputed covariates using logistic and linear regressions that included the event indicator and Nelson-Aalen marginal cumulative hazard.SMC-FCS instead used imputation models compatible with the Cox substantive model.
- Limitation: SMC-FCS requires draws involving the Cox baseline hazard H0(.), an infinite-dimensional parameter, and the implementation did not incorporate its uncertainty.The authors note uncertainty about asymptotically unbiased variance estimates in this semiparametric setting.
- Results: For n = 1,000, complete-case analysis was essentially unbiased, whereas SMC-FCS was unbiased and had correct confidence-interval coverage.The FCS biases were larger at n = 1,000 than at n = 100 in the reported simulations.
- Results: For n = 100, FCS estimates were somewhat biased but had approximately 95% coverage, while SMC-FCS showed slight upward bias and greater efficiency than complete-case analysis.SMC-FCS confidence intervals had correct coverage despite ignoring uncertainty in the baseline hazard function.
8 Analysis of data from ADNI
The ADNI analysis compares complete-case analysis, passive FCS, and SMC-FCS for time to Alzheimer’s conversion with partially observed covariates. SMC-FCS preserved the quadratic association seen in complete cases and improved precision relative to complete-case analysis.
- Analysis: The analysis modeled time to Alzheimer’s conversion using baseline variables including CSF Aβ1−42, P-tau, family history, and hippocampal volume.The complete-case analysis included 127 complete cases, of whom 61 converted to AD.
- Imputation methods: Passive FCS used linear regression for continuous variables and logistic regression for binary variables, incorporating the event indicator and Nelson-Aalen cumulative hazard.Fifty imputations were used.
- Results: Passive FCS attenuated the estimated quadratic and linear Aβ1−42 coefficients relative to complete-case analysis, with the linear coefficient no longer statistically significant.This matched the simulation finding that ignoring quadratic effects during imputation attenuates curvature estimates.
- Results: SMC-FCS produced linear and quadratic Aβ1−42 coefficients closer to complete-case estimates and preserved statistical significance of the quadratic coefficient.The other coefficients and confidence intervals were similar to those from FCS.
- Implications: Multiple imputation reduced standard errors and improved precision relative to complete-case analysis by including subjects with some missing values.SMC-FCS preserved the quadratic association between CSF Aβ1−42 and conversion hazard seen among complete cases.
9 Multiple imputation of covariates in practice
Multiple imputation requires care when the substantive model is unknown or when several analyses are planned. SMC-FCS can accommodate multiple putative outcome models and auxiliary variables, but one imputation set may not suit every later analysis.
- Application: Table 5 compares standard errors and log hazard-ratio estimates from complete-case, FCS, and SMC-FCS analyses of baseline predictors of Alzheimer’s disease conversion.The table concerns a Cox proportional hazards model relating conversion hazard to baseline risk factors.
- Multiple analyses: A single set of imputations may not be suitable for all possible subsequent analyses because validity depends on preserving the features later investigated.Imputation models are likely to be misspecified to some extent, although biases may be small when relevant data features are preserved.
- Multiple analyses: When several substantive outcome models are considered, SMC-FCS can use a larger model containing them as special cases.The resulting covariate models are compatible with the larger model and semi-compatible with substantive models nested within it.
- Auxiliary variables: Auxiliary variables can improve imputation efficiency or make the missing-at-random assumption more plausible, even when they are absent from the substantive model.Fully observed auxiliary variables can instead be incorporated as additional fully observed covariates.
- Auxiliary variables: Compatibility is not directly defined when imputation and substantive models involve different variable sets, as with auxiliary variables.Fully observed auxiliary variables can be included in the covariate set to address this distinction.
10 Discussion
The discussion argues that SMC-FCS addresses incompatibility between covariate imputation and substantive models, especially for nonlinear effects or interactions. Simulations support consistent estimation and nominal coverage under compatible, correctly specified models, while computational cost and model scope remain limitations.
- Motivation: Standard covariate imputation choices may be incompatible with substantive models, particularly when they contain nonlinear effects or interactions.Incompatibility can make the imputation model misspecified when the substantive model is correctly specified.
- Contribution: SMC-FCS modifies FCS so each covariate is imputed from a model compatible with the substantive model.Compatibility prevents conflicting assumptions between imputation and substantive models, although it does not guarantee correct specification.
- Results: When covariate models are mutually compatible and correctly specified, the authors conjecture that SMC-FCS is consistent, supported by simulation results.In these cases, confidence-interval coverage attained nominal levels even without equivalence to a Bayesian joint model.
- Results: Under misspecified covariate models, SMC-FCS estimates were still less biased than those from standard approaches.The comparison is reported as a simulation finding rather than a general guarantee.
- Comparisons: JAV is consistent under MCAR for linear substantive models with nonlinear effects or interactions, but gives biased estimates under MAR or for logistic regression.A polynomial-combination method performed better than JAV in the limited simulation study, with less bias and coverage closer to nominal.
- Limitations: SMC-FCS is more computationally intensive than standard FCS because it uses rejection sampling.It took six times longer than standard FCS to create 10 imputations in one simulated quadratic-effect scenario.
- Limitations: The approach requires further study of its statistical properties and defines imputation models for a single possibly multivariate outcome variable.The authors note this limits settings involving substantive models for multiple different outcomes.
- Extensions: SMC-FCS can be extended to impute missing outcomes by sampling from the assumed substantive model, although the paper assumes the outcome is fully observed.The paper notes that missing outcomes may provide little additional information without auxiliary variables.