Source-linked AI summary

What Is Meant by "Missing at Random"?

Shaun Seaman, John Galati, Dan Jackson, John Carlin

arXiv:1306.2812v1stat.ME

TL;DR

The paper addresses inconsistent definitions of MAR and unclear claims about when missingness mechanisms may be ignored. It standardizes MAR and MCAR definitions and relates them to ignorability across inferential frameworks, identifying sufficient conditions for valid likelihood-based inference while noting that these conditions need not be necessary.

  • Problem

    The literature is unclear about MAR definitions, ignorability, and what constitutes valid inference under different likelihood-based frameworks.

  • Method

    The paper formulates precise realised and everywhere MAR and MCAR definitions and analyzes ignorability under direct-likelihood, Bayesian, and frequentist frameworks.

  • Results

    Theorems show that under realised MAR and distinctness, likelihoods that account for or ignore the missingness mechanism can be proportional for direct-likelihood inference.

  • Takeaways & Limitations

    Clear terminology is needed when stating missing-data assumptions and distinguishing direct-likelihood from frequentist likelihood concepts.

  • Takeaways & Limitations

    The stated theorems provide sufficient rather than necessary conditions for ignoring the missingness mechanism.

Abstract

from arXiv · show

The concept of missing at random is central in the literature on statistical analysis with missing data. In general, inference using incomplete data should be based not only on observed data values but should also take account of the pattern of missing values. However, it is often said that if data are missing at random, valid inference using likelihood approaches (including Bayesian) can be obtained ignoring the missingness mechanism. Unfortunately, the term "missing at random" has been used inconsistently and not always clearly; there has also been a lack of clarity around the meaning of "valid inference using likelihood". These issues have created potential for confusion about the exact conditions under which the missingness mechanism can be ignored, and perhaps fed confusion around the meaning of "analysis ignoring the missingness mechanism". Here we provide standardised precise definitions of "missing at random" and "missing completely at random", in order to promote unification of the theory. Using these definitions we clarify the conditions that suffice for "valid inference" to be obtained under a variety of inferential paradigms.

1. INTRODUCTION

The literature used MAR inconsistently and left unclear how its definitions relate to valid inference and ignorability. The paper identifies these ambiguities and sets out precise definitions and framework-specific conditions.

  • MAR has been used inconsistently, including disagreement over whether it concerns only realised or all possible patterns and observed values.
  • Ambiguity also surrounds the distinction between direct-likelihood and frequentist likelihood inference.
  • Unclear terminology makes it difficult to interpret, compare, and apply missing-data theory in statistical practice.
  • The paper provides unambiguous MAR definitions and explains their relation to ignorability under different inferential frameworks.

2. TWO DEFINITIONS OF MAR AND MCAR

The paper formalizes missingness using data, missingness indicators, and observed-data functions, then distinguishes realised from everywhere MAR and MCAR. The stronger everywhere conditions cover all possible patterns and data values.

  • Y contains potentially observable data, M records whether each element is observed, and o(Y,M) extracts the observed elements.The length of o(Y,M) is the number of observed elements.
  • Realised MAR concerns the realised missingness pattern and realised observed data, requiring its probability not to depend on the realised missing values.
  • Everywhere MAR requires the analogous condition for all possible missingness patterns and observed-data values, and implies realised MAR.
  • In the i.i.d. setting, the single-unit formulation is equivalent to the general definitions, but it cannot apply when units’ missingness mechanisms are dependent.
  • Realised MCAR implies realised MAR, whereas everywhere MCAR means M is independent of Y and implies both MAR conditions.

3. MAR AND MCAR IN THE LITERATURE: A REVIEW

The review shows that MAR and related notation have supported conflicting interpretations in the literature. It also identifies ambiguity about parameter quantification and the meaning of observed and missing data.

  • Authors have used “MAR” for either realised MAR or everywhere MAR, sometimes citing Rubin’s definition while stating the other condition.
  • An exchange between Heitjan and Diggle demonstrates that the two definitions can produce opposite judgments for the same fully observed dataset.
  • The notation Yobs is ambiguous because it is itself a function of M, making expressions such as f(M | Yobs,φ) confusing.
  • Some definitions omit φ or fail to state whether the MAR equality must hold for every parameter value or only the true value.
  • Definitions also vary in whether equality is required for every possible Y or only values compatible with the realised observed data.
  • MCAR has likewise been used for conditions allowing dependence on covariates, sometimes called covariate-dependent MCAR.

4. DIRECT-LIKELIHOOD, BAYESIAN AND FREQUENTIST INFERENCE

The paper distinguishes Bayesian, direct-likelihood, general frequentist, and frequentist likelihood inference because they use different inferential targets and notions of validity. Likelihood methods eliminate nuisance parameters differently, while frequentist methods rely on repeated sampling.

  • Direct-likelihood inference uses likelihood maxima for point estimates and likelihood ratios to assess parameter plausibility.
  • Profile likelihood eliminates nuisance parameters by maximising over them at each fixed value of the parameter of interest.
  • Conditional likelihood instead specifies the distribution of Y given a function of Y, producing a likelihood with fewer parameters.
  • Bayesian inference specifies a prior and obtains parameter posteriors using Bayes’ theorem, with credible intervals derived from posterior quantiles.
  • Bayesian and direct-likelihood inference model the realised data Y, whereas frequentist inference evaluates hypothetical repeated-sampling properties.
  • The literature has not consistently distinguished direct-likelihood from frequentist likelihood inference, and Bayesian analyses may also be evaluated by frequentist properties.

5. IGNORABILITY OF THE MISSINGNESS MECHANISM

The paper specifies when the missingness mechanism can be ignored under different inferential frameworks, distinguishing what “the same” inference means. Under realised MAR, parameter distinctness, and related assumptions, likelihoods or posteriors based on observed data alone can coincide with joint-model results, though expected-information standard errors require everywhere MCAR.

  • Framework and definition: Ignorability means that inference from a model for the data alone agrees with inference from a joint model for data and missingness.The paper emphasizes that the meaning of “the same” depends on the inferential framework.
  • Direct-Likelihood Inference: Under realised MAR and distinct parameter spaces, the joint likelihood factorizes, and likelihoods conditional on the missingness parameter are proportional to the likelihood ignoring missingness.The proportionality requires a missingness model assigning positive probability to the realised pattern.
  • Direct-Likelihood Inference: Direct-likelihood inference about θ can therefore use the missingness-ignoring likelihood L2 when realised MAR holds and θ and φ are distinct.The result is based on proportionality between the joint and missingness-ignoring likelihoods.
  • Bayesian Inference: For Bayesian inference, realised MAR and a priori independent θ and φ make the posterior for θ from the joint model equal to the posterior from L2 and the marginal prior for θ.The factorized posterior separates the θ and φ terms.
  • Frequentist Likelihood Inference: Under everywhere MAR with distinct parameters, frequentist likelihood estimates, information-based variance estimates, intervals, and likelihood tests are the same using L1 or L2.This repeated-sampling result applies because the likelihoods remain proportional across repeated samples.
  • Bayesian Inference: Under everywhere MAR and independent priors, Bayesian point estimators and credible intervals have the same repeated-sampling properties whether missingness is modelled or ignored.The posterior distribution for θ is identical for every possible data vector and missingness pattern.
  • Caveat: Expected information from L2 should not be used naively under everywhere MAR; its use is appropriate only under everywhere MCAR, so observed information is recommended.This is the paper’s explicit caveat for standard-error calculation.

6. CONDITIONAL LIKELIHOOD AND REPEATED SAMPLING

The paper examines conditional likelihood and repeated-sampling settings, showing how MAR-related conditions determine when missingness can be ignored.

  • Conditional likelihood: Conditional likelihood can model outcomes given fully observed covariates without specifying a likelihood for all data.Here, the covariates form the conditioning variable X.
  • Conditional likelihood: Under everywhere MAR, conditional likelihoods that account for missingness are proportional to the likelihoods ignoring the missingness mechanism in realised samples and repeated samples conditional on X.The repeated sampling is conditional on X but not on the realised missingness pattern M.
  • Repeated sampling: For repeated sampling conditional on X and M, a weaker condition than everywhere MAR can replace realised MCAR in the relevant theorem.The condition requires the realised missingness probability to be invariant across complete-data values sharing the same observed components.
  • Repeated sampling: With fully observed covariates in repeated-measures data, the everywhere version of the weaker condition is called covariate-dependent MCAR.This terminology applies when X consists of the fully observed covariates.

7. DISCUSSION

The discussion distinguishes meanings of ignorability across inferential frameworks and emphasizes that precise MAR definitions are needed for valid practice. It also notes that repeated-sampling validity generally requires everywhere MAR when the missingness mechanism is not modeled.

  • DISCUSSION: Clearer definitions of missingness assumptions and sharper distinctions among inferential frameworks are presented as necessary for comparing methods and improving statistical practice.The discussion connects this need to confusion in the literature and to applications involving outcomes, covariates, and longitudinal data.
  • DISCUSSION: Ignorability can mean proportional likelihoods or equal sampling distributions, while validity of frequentist procedures is a separate interpretation.The paper treats ignorability as dependent on the assumed missingness model in one interpretation, but notes that usage is not universal.
  • DISCUSSION: Theorem 1 implies that using L2, the likelihood ignoring missingness, yields valid frequentist likelihood or frequentist Bayesian inference under the stated proportionality result.The result follows because L2 is proportional to the likelihood incorporating the true missingness mechanism.
  • DISCUSSION: MAR plus distinctness provides sufficient conditions for ignoring missingness in direct-likelihood and Bayesian inference, but not necessarily necessary conditions in every setting.The paper notes that restricted missingness-model parameter values may still make the likelihoods proportional without realised MAR.
  • DISCUSSION: Under distinct parameters and a complete class of data distributions, realised MAR is necessary and sufficient for ignorability in frequentist likelihood inference.This result identifies conditions under which the usual MAR requirement is exact rather than merely sufficient.
  • DISCUSSION: Incomplete-data methods that ignore the missingness mechanism cannot be guaranteed valid in repeated samples unless everywhere MAR holds.The discussion characterizes everywhere MAR as a restrictive assumption whose practical importance may be underappreciated.

APPENDIX

The appendix proves an equivalence between realised MAR and likelihood ignorability under parameter-space, completeness, and positivity assumptions.

  • APPENDIX: When the joint parameter space factors, the relevant conditional distribution is complete, and positivity holds, L1 is proportional to L2 for every φ if and only if realised MAR holds.This theorem establishes both directions under its stated assumptions.
  • APPENDIX: The appendix’s necessity argument relies on completeness of the conditional distribution of the missing components given the observed components.Without that assumption, the stated equivalence is not established by this proof.
  • APPENDIX: The proof’s only-if direction starts from proportionality of L1 and L2 and derives a quantity independent of θ.That quantity is then used to characterize the missingness probability.
  • APPENDIX: Completeness forces the missingness probability to equal the derived quantity for all admissible φ and compatible complete-data values.Therefore, the probability cannot depend on the missing components, which is precisely realised MAR.
Loading 1306.2812v1…