Source-linked AI summary

Assumptions of IV Methods for Observational Epidemiology

Vanessa Didelez, Sha Meng, Nuala A. Sheehan

arXiv:1011.0595v1stat.ME

TL;DR

Observational epidemiology needs methods that address unobserved confounding, but IV approaches differ in their causal targets and modeling assumptions. This paper compares those approaches and examines estimator bias under assumption violations. It finds that effect modification by unobserved confounders can seriously bias all IV methods and recommends routine sensitivity analysis.

  • Problem

    Observational IV methods differ in their targeted causal parameters and rely on assumptions whose appropriateness is important when unobserved confounding is present.

  • Method

    The paper theoretically compares common IV approaches in observational epidemiology and numerically evaluates their asymptotic bias when assumptions are violated.

  • Results

    Effect modification by unobserved confounders seriously increases the bias of all IV methods and can produce bias even under no causal effect.

  • Takeaways & Limitations

    Practical IV applications should be accompanied by sensitivity analyses, especially for interactions between exposure effects and unobserved confounders.

  • Takeaways & Limitations

    The conclusions about estimator bias are based on particular models and scenarios, and absence of effect modification by unobserved confounders cannot be empirically verified.

Abstract

from arXiv · show

Instrumental variable (IV) methods are becoming increasingly popular as they seem to offer the only viable way to overcome the problem of unobserved confounding in observational studies. However, some attention has to be paid to the details, as not all such methods target the same causal parameters and some rely on more restrictive parametric assumptions than others. We therefore discuss and contrast the most common IV approaches with relevance to typical applications in observational epidemiology. Further, we illustrate and compare the asymptotic bias of these IV estimators when underlying assumptions are violated in a numerical study. One of our conclusions is that all IV methods encounter problems in the presence of effect modification by unobserved confounders. Since this can never be ruled out for sure, we recommend that practical applications of IV estimators be accompanied routinely by a sensitivity analysis.

1. INTRODUCTION

Observational epidemiology often cannot use randomized trials and may suffer from implausible assumptions about unobserved confounding. The paper compares instrumental-variable approaches, their causal targets and assumptions, and evaluates bias when those assumptions fail.

  • Randomized trials may be unethical, impractical, or unrepresentative for exposures such as smoking, alcohol, nutrition, and exercise.
  • Standard observational adjustment assumes that sufficient confounders have been measured, an assumption that can produce misleading results when unobserved confounding remains.
  • An instrumental variable predicts exposure while having no direct effect on disease and remaining independent of unobserved confounders.
  • Suitable instruments are difficult to justify, although examples include prescription preference, cigarette price, and genetic variants used in Mendelian randomization.
  • IV methods can test or bound causal effects using defining properties, but point identification requires additional parametric and distributional assumptions.
  • The paper compares IV approaches in observational epidemiology and uses a numerical study to examine bias under assumption violations.

2. USING A GENETIC VARIANT AS AN IV

Mendelian randomization uses genetic variation as an instrumental-variable strategy to investigate causal effects when alcohol-related observational associations may be confounded. The ALDH2 example illustrates both the rationale for genetic instruments and the conditions required to interpret them.

  • Observed associations between alcohol consumption and several diseases are suspected to reflect confounding by diet, lifestyle, and socioeconomic factors.
  • The ALDH2*2 variant is associated with unpleasant acetaldehyde-related symptoms, and carriers tend to limit alcohol consumption.
  • Because genes are randomly assigned during meiosis, ALDH2 carriers should not systematically differ from other carriers in unobserved factors.
  • ALDH2 can provide evidence for a causal effect when genotype-associated differences in disease risk are interpreted under the assumption of no direct effect outside alcohol consumption.
  • A genetic variant associated with exposure does not automatically qualify as an instrument, because population structure can create differences in allele frequencies and disease prevalence.

3. CAUSAL INFERENCE

Causal inference concerns effects of intervening on exposure, but observational IV analyses must distinguish among population, subgroup, local, and individual causal parameters. The paper formalizes these distinctions and states the core independence conditions defining an instrument.

  • 3. CAUSAL INFERENCE: The do-operator distinguishes the outcome distribution under an imposed exposure value from the distribution observed among people with that exposure.
  • 3.1 Causal Parameters: Population causal parameters compare two exposure settings for the whole population, whereas conditional, local, and individual effects concern narrower targets.
  • 3.1 Causal Parameters: A local causal effect describes changing exposure for people who would normally be exposed, such as reducing alcohol consumption among naturally high consumers.
  • 3.1 Causal Parameters: Under effect modification, population, local, and individual effects can differ, and an estimator targeting one parameter can be biased for another.
  • 3.2 Instrumental Variables: IV methods offer an alternative when observational adjustment cannot address suspected unobserved confounding.
  • 3.2 Instrumental Variables: The core IV conditions require instrument independence from confounding, association with exposure, and conditional independence from outcome given exposure and confounder.
  • 3.2 Instrumental Variables: These conditions can be partially tested from observable data in categorical settings through inequality constraints on the joint distribution.
  • 3.2 Instrumental Variables: Under intervention, the exclusion restriction is expressed as independence between the instrument and outcome given the intervention.

4. SOME COMMON IV MODELS

Common IV models differ in the causal parameters they identify and in the additional assumptions required for point identification. The models also differ in data requirements and in how effect modification is handled.

  • Core identification: Core IV conditions can test for causal effects or bound them, but generally do not point-identify causal effects without additional assumptions.For general distributions, point identification from the core conditions alone occurs only in unusual situations.
  • Model restrictions: Additional parametric restrictions define estimands that equal the target causal parameter only when the statistical model is correctly specified.Under misspecification, the estimand may differ from the target because it relies on incorrect model assumptions.
  • Linear IV models: Linear models may be used as approximations for binary outcomes, because their conditional mean can otherwise take values outside the [0,1] range.Causal relative risks and odds ratios can nevertheless be identified under the linear model as described in the paper.
  • Linear IV models: In linear IV models, the LIVAE identifies individual, average, or local effects depending on whether homogeneity is assumed across individuals, unobserved confounders, or instrument levels.The linear model also identifies the causal relative risk through the corresponding LIVRR estimand.
  • Nonlinear Wald type methods: Wald-type methods use outcome and exposure associations with the instrument, with log-odds reasoning theoretically motivated by a rare-disease approximation to relative risk.The approach is described as heuristic before the loglinear structural model supplies a theoretical basis.
  • Multiplicative structural mean models: Multiplicative structural mean models identify a local causal relative risk when the exposure effect is constant across instrument levels, and identify the population effect under no exposure–confounder interaction on the multiplicative scale.The no-heterogeneity assumption across instrument levels is needed because only one parameter can be identified; analogous additive and multiplicative assumptions cannot both generally hold.

5. NUMERICAL ILLUSTRATION OF ASYMPTOTIC BIAS

The numerical study evaluates asymptotic relative bias for several IV approaches under realistic violations of their modeling assumptions, focusing on causal relative risk. It varies causal effects, confounding, and effect modification, and compares the resulting estimands with the true effect.

  • The study investigates whether IV methods remain consistent or avoid serious excess bias under no effect, no confounding, and related realistic violations.These desiderata motivate the comparison of asymptotic bias across concrete scenarios.
  • Asymptotic relative bias is defined from the difference between the targeted causal parameter and the model-specific estimand at the true distribution.It is zero when the model is correctly specified and identifies the causal parameter.
  • The comparison targets causal relative risk using the linear model, log-linear Wald approach, and multiplicative structural mean model.Each approach identifies CRR under its respective assumptions; WaldOR is omitted because it is slightly more biased for CRR than WaldRR.
  • The data-generating distributions are generally misspecified for the evaluated IV models because the outcome follows a logistic dependence on exposure and unobserved confounder.When α4 = 0, there is no effect modification by U on the logistic scale, although additive or multiplicative effect modification may remain.
  • Scenarios vary causal effects from CRR = 1.0 to 1.33 and 3.03, confounding through α3 ∈ {0, 0.1, 1, 2}, and interactions through β4 and α4 ∈ {−1, 0, 1}.Only combinations with |α4| ≤ |α3| are considered as realistic settings.
  • Nonparametric CRR bounds are extremely wide and include the null; for CRR = 3.03, they are approximately [0.2, 30].Stronger instruments can narrow the bounds, but the relative risk of 2.4 used here is considered about as strong as expected in Mendelian randomization.

5.2 Numerical Results

The numerical study compares IV estimators with the naïve relative risk across scenarios with no confounding, null effects, and causal effects with confounding. Bias varies substantially by method and interaction structure: MSMMRR is often least biased, while WaldRR can be severely biased.

  • No confounding: With no confounding, only WaldRR has nonzero bias because its binary-exposure assumption cannot be satisfied, whereas the other models’ assumptions hold.The naïve, linear, and multiplicative structural mean models are satisfied when X and Y are binary in these scenarios.
  • Causal effect and confounding: For CRR = 3.03, WaldRR has 40%–250% relative bias and seriously overestimates the true effect.Its bias is comparable to or larger than that of the other IV methods when CRR = 1.33.
  • Causal effect and confounding: The MSMMRR bias is similar for small and large CRR, with a maximum of 17%.This contrasts with the stronger deterioration observed for LIVRR and WaldRR as the causal effect becomes larger.
  • Interaction effects: When α4 = 0, LIVRR and MSMMRR are only slightly biased and much less biased than NRR; when β4 = 1 with α4 ≠ 0, MSMMRR bias reaches 17%.LIVRR bias ranges from 24% for small CRR to 45% for large CRR in the latter interaction settings.
  • Comparing methods: MSMMRR is much less biased in most settings, but no method is uniformly best; LIVRR outperforms it when α4 = −1, and NRR does so when α3 = 1 as well.Across the considered settings, the MSMMRR is least biased overall, while WaldRR is most biased in the null-effect scenarios.
  • Sign of bias: IV estimators can be negatively biased, and their bias does not always share the same sign, so they cannot generally be classified as over- or underestimating.By contrast, NRR is always positively biased under the study’s chosen coefficients for U.
  • Exposure frequency: With 50% exposure frequency, all IV methods show much less bias, WaldRR performs more sensibly, and MSMM remains clearly least biased.The MSMM is also not sensitive to interaction effects at this exposure frequency.

5.3 Practical Implications

The numerical study clarifies when common IV approaches are reliable and highlights the importance of untestable effect-modification assumptions. It supports method comparison and routine sensitivity analysis for specific applications.

  • Linear IV: Linear IV performed better than expected for binary outcomes, with relative asymptotic bias below 20% in all but six considered scenarios.Its approximate suitability is supported by a small true causal effect and restricted, balanced, or approximately normal exposure distributions.
  • Wald methods: Wald-type methods can be extremely biased when their strong assumptions are violated, including realistic settings, large causal effects, and even no confounding.Their bias can exceed that of the naïve approach and increases with the strength of the true causal effect.
  • Effect modification: All IV approaches except bounds assume no effect modification by the unobserved confounder, and violating this assumption can seriously increase bias, even under the null.Because the assumption concerns unobserved confounders, it is difficult to assess or justify in practice.
  • Practical recommendations: MSMM appears most recommendable for relative-bias performance in situations like those studied, especially with binary outcomes, but efficiency and other properties also matter.Because the numerical study covers specific scenarios, further comparisons and sensitivity analyses are recommended for each application.

6. CONCLUSION AND DISCUSSION

The comparison shows that IV approaches target different causal parameters and rely on different assumptions, while bias can become serious when effect modification by unobserved confounders is present. The authors therefore recommend explicit assumption assessment and routine sensitivity analyses.

  • Different IV approaches target individual, population, or local causal effects, rather than merely different effect scales.
  • The SMM approach makes the weakest assumptions because it does not require a model for exposure X given instrument G.Under stronger no-interaction assumptions, its local causal effect can equal the population causal effect.
  • All estimators encounter difficulties estimating population effects when exposure effects differ across levels of the unobserved confounder.The bias calculations apply only to the selected models and scenarios.
  • When interactions on the multiplicative scale are plausible, the MSMM estimator is closer to the local effect, whereas Wald relative risk is likely to be seriously biased.
  • Because absence of effect modification cannot be verified empirically, practical IV analyses should include sensitivity analyses, including for continuous outcomes analyzed with linear no-interaction models.
  • WaldRR and WaldOR require restrictive assumptions and may be unusable without joint information, while linear IV may be less biased but is poorly matched to rarely reported binary-outcome risk differences.
  • The study assesses asymptotic bias but not efficiency, whose variance depends strongly on instrument strength; measurement error can also bias point estimates and require additional modeling assumptions.

Justification of LIVAE

Under the additive structural mean model, the causal effect parameter is identified through a covariance ratio involving the outcome, exposure, and centered instrument.

  • The average causal effect equals model parameter β under the specified model.
  • The LIVAE estimand is β = Cov(Y,G)/Cov(X,G).
  • Risk ratios and odds ratios additionally require estimating the intercept of the relevant model.
  • Under the additive SMM, the exclusion restriction yields the moment condition E((Y − β_LX)G̃) = 0.
  • Solving that moment condition gives β_L = Cov(Y,G)/Cov(X,G).

Justification of WaldRR

The WaldRR identification argument uses a location-scale exposure model and conditional independence of the exposure residual from the instrument given the unobserved confounder.

  • The exposure residual ξ must satisfy ξ ⟂ G | U, an assumption automatically met by certain normal constant-variance models.
  • Location-scale exposure models satisfy the residual condition when only their location parameter depends on G and U.Restricted-support families such as Bernoulli distributions typically do not satisfy it.
  • Writing X = δG + k(U) + ξ separates the instrument-associated exposure component from confounding and residual components.
  • The loglinear regression coefficient of G on Y is γδ, while δ is recovered from a linear regression of X on G.
  • These regressions identify the causal relative risk through the WaldRR under the stated assumptions.

Justification of MSMMRR

The MSMMRR argument derives causal risk quantities from exclusion-restriction estimating equations and observed outcome-exposure distributions, with additional assumptions linking the local and population relative risks.

  • For the multiplicative model, the exclusion restriction induces E(Y exp(−γ_LX)G̃) = 0 to estimate γ_L.Unlike the linear case, this nonlinear equation generally lacks a simple closed-form solution.
  • With binary instrument G, the exclusion restriction equates E(Y exp(−γ_LX)|G = 1) and E(Y exp(−γ_LX)|G = 0).
  • When X and Y are also binary, the conditional expectation can be expanded into observable moments and rearranged to obtain the model expression.
  • Under additional assumptions, integrating over G and X yields E(Y | do(X̃ = 0)) from observed outcome means and exposure probabilities.
  • If the Y-X relative risk is constant within levels of U, exp(γ_L) is the population causal relative risk.
  • Substitution then gives an expression for E(Y | do(X̃ = 1)) using observed outcome means, exposure probabilities, and exp(γ_L).

Relations Between Assumptions

The paper establishes a hierarchy among IV-related assumptions: stronger structural or linear models imply additive or multiplicative structural mean models, but the implications generally do not reverse.

  • Relations Between Assumptions: Under the IV conditions, the linear model implies the additive structural mean model (SMM).
  • Relations Between Assumptions: The log-linear model implies the multiplicative structural mean model (MSMM), but the reverse implication does not hold.
  • Relations Between Assumptions: The structural equation model specifies individual potential responses as Y_i(x) = β_Ix + ξ_i, with ξ_i fixed within individuals but varying across individuals.
  • Relations Between Assumptions: That structural equation model implies the linear model and therefore the additive SMM, whereas the reverse is not true.
Loading 1011.0595v1…