Source-linked AI summary

Structural Nested Models and G-estimation: The Partially Realized Promise

Stijn Vansteelandt, Marshall Joffe

arXiv:1503.01589v1stat.ME

TL;DR

Applied use of structural nested models and G-estimation remains relatively infrequent despite their advantages for estimating joint treatment effects. This paper reviews their development, explains their properties and extensions, and concludes that they offer greater flexibility and typically better-performing estimators than alternatives.

  • Problem

    Estimating joint effects of sequential treatments is challenging when confounders are affected by prior treatment, while SNMs and G-estimation remain infrequently applied.

  • Method

    The paper reviews structural nested models and G-estimation, including identification under ignorability, doubly robust estimation, censoring, sequential treatments, and direct and indirect effects.

  • Results

    SNMs and G-estimation allow greater flexibility than marginal structural models and typically yield better-performing estimators, especially for continuous or strongly confounded exposures.

  • Takeaways & Limitations

    These methods are particularly useful when treatment is continuous or strongly correlated with subject characteristics, where inverse-probability-weighted estimators may perform poorly.

  • Takeaways & Limitations

    Censoring-related estimating equations may be discontinuous in the parameters, causing optimization difficulties or no finite-sample solution.

Abstract

from arXiv · show

Structural nested models (SNMs) and the associated method of G-estimation were first proposed by James Robins over two decades ago as approaches to modeling and estimating the joint effects of a sequence of treatments or exposures. The models and estimation methods have since been extended to dealing with a broader series of problems, and have considerable advantages over the other methods developed for estimating such joint effects. Despite these advantages, the application of these methods in applied research has been relatively infrequent; we view this as unfortunate. To remedy this, we provide an overview of the models and estimation methods as developed, primarily by Robins, over the years. We provide insight into their advantages over other methods, and consider some possible reasons for failure of the methods to be more broadly adopted, as well as possible remedies. Finally, we consider several extensions of the standard models and estimation methods.

1. INTRODUCTION

Structural nested models were developed to estimate joint effects of treatment sequences when time-varying confounding makes standard simultaneous regression inappropriate. The introduction frames their advantages and broad extensions against their relatively infrequent use in applied research.

  • Motivating problem: SNMs address confounding by variables affected by earlier treatment when estimating the joint effect of a treatment sequence.Such a variable is associated with the outcome, predicts subsequent treatment, and is affected by earlier treatment.
  • Motivating example: In the EPO example, hematocrit is affected by earlier EPO, predicts later EPO, and confounds the later treatment effect on mortality.This creates confounding by a treatment-affected variable in observational studies of extended EPO dosing.
  • Motivating problem: Standard methods estimating treatment effects simultaneously are inappropriate whether or not they adjust for or condition on the time-varying confounder.Adjusting for the intermediate variable can block the pathway from early treatment through the confounder to the outcome.
  • Methods: Robins introduced the parametric G-formula, structural nested models with G-estimation, and related approaches to handle such confounding.These methods commonly rely on assumptions including no unmeasured confounders or sequential ignorability and can face positivity violations.
  • Scope and adoption: SNMs include models for treatment effects on outcome means and models for effects on entire outcome distributions, yet applied use has remained relatively infrequent.The first category includes SNMMs and links to SNCFTMs; the second includes SNDMs.

2. STRUCTURAL MODELS FOR POINT TREATMENTS

This section introduces structural models for point treatments, using potential outcomes to define causal effects and parameterizing treatment removal through mean- or distribution-based models. It also describes multivariate-treatment extensions and cautions that distributional and rank-preserving assumptions can be substantially restrictive.

  • Potential outcomes: Potential outcomes Y^a define causal effects by comparing outcomes under different treatments for the same subject or population, linked to observed outcomes by consistency.Y^a is observed as Y when A=a and otherwise is counterfactual.
  • Structural Mean Models: Structural Mean Models parameterize average causal effects through a known link and treatment-effect function, with ψ*=0 typically representing no treatment effect.With treatment absence coded as a=0, SMMs express the effect of removing treatment on the outcome mean.
  • Structural Distribution Models: Structural Distribution Models map treated conditional-outcome percentiles to untreated percentiles, thereby modeling treatment effects on the outcome distribution rather than only its mean.Their parameterization typically sets ψ*=0 as the no-effect model and can include location-shift specifications.
  • Model limitations: Rank-preserving SDMs are easier to communicate but seldom plausible because they require subjects’ outcome rankings to remain unchanged after treatment removal.Location-shift SDMs also impose stronger assumptions than correspondingly parameterized SMMs by requiring the same shift at every percentile.

3. IDENTIFICATION AND ESTIMATION IN STRUCTURAL MODELS FOR POINT TREATMENTS

Under ignorability, SMM and SDM parameters are identified, with the weaker assumption sufficient for treatment-effect-on-the-treated parameters and blip functions nonparametrically identified. G-estimation solves estimating equations using working models, is doubly robust in key models, but failure-time analyses face censoring and nonsmooth-equation challenges.

  • Identification: Ignorability is empirically unverifiable and sufficient for identifying treatment effects among subjects receiving the corresponding treatment level.For binary treatments, the weaker assumption identifies the effect of treatment on the treated.
  • Identification: Under ignorability, SMM and SDM blip functions are nonparametrically just identified from the law of the observables without restrictions or parameterization.The contrast compares outcomes under observed treatment with outcomes that would have been seen without treatment for each treatment and covariate level.
  • Estimation: Estimating equations are based on setting conditional covariances or independence relations involving blipped-down outcomes to zero, with locally efficient variants requiring additional variance and model specifications.The equations can also accommodate repeated-measures outcomes.
  • Estimation: G-estimators solve SMM or SDM estimating equations after substituting consistent estimates from working models for exposure and residual-outcome distributions.In SDMs and linear or loglinear SMMs, the estimator is doubly robust when either working model is correctly specified, alongside the structural model.
  • Failure-time outcomes: Censoring complicates failure-time analyses: inverse probability of censoring weighting handles random censoring, while type I censoring requires different treatment for SAFTMs and can induce nonsmooth estimating equations.Nonsmoothness can cause optimization problems or leave estimating equations without finite-sample solutions.

4. PROPERTIES OF G-ESTIMATION IN STRUCTURAL MODELS FOR POINT TREATMENTS UNDER IGNORABILITY

For point treatments under ignorability, G-estimation can be less efficient than correctly specified ordinary regression but is more extensible and can outperform IPW, especially with limited overlap or outcome-model misspecification. Its interpretation as a weighted average of treatment effects remains useful even when homogeneity fails.

  • Efficiency relative to ordinary regression: Correctly specified ordinary regression estimators are at least as efficient as the previously considered G-estimators, but their advantage is usually not much larger under homoscedasticity.The OLS estimator has smaller asymptotic variance, though the difference is described as usually not much smaller.
  • Efficiency relative to ordinary regression: Ordinary regression estimators lack G-estimators’ extensibility to sequential treatments and rely explicitly on modeling the outcome–covariate association.This reliance can be disadvantageous when treated and untreated subjects are very different.
  • Comparison with IPW estimation: The locally efficient IPW estimator has strictly larger variance than the locally efficient G-estimator unless treatment and covariates are independent, when the estimators are equally efficient.The models define the same parameter, but the marginal structural model is less restrictive because it does not require treatment-effect homogeneity.
  • Comparison with IPW estimation: Limited propensity-score overlap can make IPW variance large, whereas G-estimation retains a useful weighted-average interpretation that emphasizes strata with more treatment-effect information.When Var(A | L) is close to zero, 1/Var(A | L) can become large; the G-estimator’s weighted average gives most weight to informative strata.
  • Robustness to misspecification: When the outcome model is misspecified in regions of little overlap, the misspecification has minor impact on G-estimator variance but particularly strong impact on locally efficient IPW variance.The contrast may be even greater for sequential treatments and simple, inefficient inverse-probability weighting methods.

5. STRUCTURAL NESTED MODELS FOR TIME-VARYING TREATMENTS

This section defines structural nested mean models for sequential treatments as conditional blip-effect models that recursively remove treatment effects over time. It also introduces structural nested distribution models, which generalize rank-preserving percentile mappings to restrictions holding in distribution.

  • Structural nested mean models: SNMMs model the effect of a treatment blip at time tm on subsequent outcome means after removing later treatments and holding future treatments at reference level 0.They parameterize conditional contrasts of potential outcomes given treatment and covariate histories through tm.
  • Structural nested mean models: The parameterization can encode the null hypothesis ψ = 0, while example models distinguish short-term effects from long-term effects across treatment times.In the two-time-point example, short-term effects are assumed constant across time points, whereas other parameters encode long-term effects.
  • Structural nested mean models: SNMMs construct transformed outcomes whose means equal those expected if treatment were suspended from time tm onward, thereby formalizing sequentially blipping down treatment effects.For identity and log links, the transformation applies across subsequent outcomes or only the end-of-study outcome when that is the target.
  • Link-function limitations: For links other than identity and log, the transformed outcome can depend complicatedly on the observed data distribution, including the conditional distribution of later covariates and treatments.With a logit link and two time points, calculating U0(ψ) requires both a conditional outcome mean and the distribution of (L1,A1) given (L0,A0).
  • Structural nested distribution models: SNDMs map percentiles of counterfactual outcome distributions and relax rank-preserving restrictions by requiring the modeled equality to hold only in distribution conditional on observed history.They yield recursively constructed variables that predict outcomes after treatment is suspended from time tm onward.

6. IDENTIFICATION AND ESTIMATION IN STRUCTURAL NESTED MODELS FOR SEQUENTIAL TREATMENTS

SNMs support identification and G-estimation under a broader range of assumptions than marginal structural models, including sequential ignorability, specified departures from it, and instrumental variables. Their G-estimators can be doubly robust and accommodate non-linear outcomes, but model specification is difficult and IV analyses sacrifice identification, precision, and finite-sample performance.

  • Identification: SNMs offer a broader array of useful identifying assumptions because instrumental-variable and future-ignorability assumptions can support inference when they do not identify marginal treatment effects indexing MSMs.The section highlights this broader identification framework as an important advantage of SNMs.
  • Sequential ignorability: Under sequential ignorability, G-estimation identifies SNMM parameters by solving estimating equations based on conditional covariances between blipped-down outcomes and specified functions of treatment histories.The estimating equations sum these conditional covariances across treatment times.
  • Estimation: G-estimators are doubly robust when treatment and outcome models are variation-independent: they remain consistent if the SNM and either model A or model B are correctly specified.This property holds regardless of which of the two auxiliary models is correctly specified.
  • Estimation: Specifying model B can be thorny and nontrivial when covariates are high-dimensional or strongly associated with treatment, despite alternative estimators being more efficient under correct specification.The concern is consequential because model B may be difficult to postulate and misspecification threatens the preferred robustness strategy.
  • Departures from sequential ignorability: When sequential ignorability fails, sensitivity analyses can vary a departure function over a plausible range, while future ignorability may sometimes eliminate residual confounding bias.One example varies γ between −1 and 1; future ignorability can address intermittent measurement of confounding covariates in continuous-time treatment processes.
  • Instrumental variables: IV-based G-estimation includes two-stage least squares and extends to censored failure-time and dichotomous outcomes, but IV analyses lose nonparametric identification and typically have lower precision and greater finite-sample bias.The IV framework also extends to sequential treatments, while residual dependencies or nonlinear treatment effects may remain unidentified without additional untestable assumptions.

7. PREDICTING THE EFFECTS OF INTERVENTIONS

The section describes how structural nested models can estimate expected counterfactual outcomes under alternative treatment regimes. It highlights coherence and computational complications when comparing regimes, and explains how current treatment interaction functions can address some of these complications.

  • Estimating counterfactual outcomes: Expected outcome without treatment can be estimated by averaging U ∗ 0 ( ˆψ) in SNMMs or U0( ˆψ) in SNDMs.These estimands follow from conditional mean identities involving U ∗ 0 (ψ∗) and U0(ψ∗).
  • Comparing treatment regimes: Using separate structural nested models for different treatment regimes may produce models that fail to imply coherent comparisons between expected counterfactual outcomes.A separate model can use the alternative regime as its reference, but coherence across models remains a concern.
  • Dynamic treatment regimes: Current treatment interaction functions can overcome complications in dynamic-regime prediction and transport effects in the treated to population-averaged treatment effects.These functions supplement the SNM, but the data carry no information about them.
  • Computational complications: Evaluating the transformed untreated outcome may require modeling the distribution of L1 conditional on A0,L0, which is cumbersome when L1 is high-dimensional.This complication is avoided in simple structural models without effect modification by post-treatment variables when nondynamic regimes are considered.
  • Identifying assumptions: The no-current-treatment-interaction assumption can follow from mild strengthenings of sequential ignorability or instrumental-variables assumptions together with corresponding structural-model conditions.The supplied passages describe these strengthenings for treatment histories and, in one case, binary A1 (0/1).

8. DIRECT AND INDIRECT EFFECTS

This section explains how SNMs represent controlled direct effects and why direct parameterization is needed when mediators are fixed at nonzero values. It also describes inverse-weighted estimation, instrumental-variable alternatives, and the identification challenges of natural direct effects.

  • Controlled direct effects: SNMs represent treatment effects with subsequent treatments fixed at reference levels, corresponding to controlled direct effects.These effects concern an exposure’s effect on an outcome while controlling subsequent treatments at specified reference levels.
  • Controlled direct effects: Knowing γ∗0(a0,l0;ψ∗) for all a0,l0 does not establish that A0 has no direct effect when A1 is controlled at a nonzero value.The parameterization may encode the controlled direct effect only when A1 is uniformly controlled at zero.
  • Controlled direct effects: Robins proposed directly parameterizing the controlled direct effect, allowing A1 to modify A0’s effect while modeling only the effect of A0.The model uses a known smooth function m(a0,a1,L0;ψ) satisfying m(0,a1,l0;ψ) = 0.
  • Estimation: Inverse weighting by f(A1 | A0, L1) addresses selection bias after treating subjects with other A1 levels as censored, assuming sequential ignorability.This approach applies SMM techniques to the counterfactual outcome Y a1, which is observed only among subjects receiving A1 = a1.
  • Identification: Initial randomization or instrumental-variable assumptions can identify controlled direct effects with SNFTMs when mediator ignorability is violated.Such violations can occur even in randomized trials because the evolution of subsequent mediators may be poorly understood.
  • Natural direct effects: Natural direct effects have appealing subject-specific interpretations, but extensions to sequential treatments or mediators remain undeveloped because of identification difficulties.Natural direct effects allow A1 to take the counterfactual level it would have under a specified A0, unlike controlled direct effects with a common fixed mediator level.

9. CONCLUDING REMARKS

SNMs and G-estimation address treatment-confounder feedback while retaining regression-like parameterization and conditioning on measured histories. Despite limited adoption and computational barriers, they offer flexible, often better-performing estimation and motivate continued software development and extensions.

  • Core approach: SNMs parameterize conditional treatment effects while avoiding conditioning on post-treatment variables by removing later-treatment effects and conditioning on prior treatment and covariate history.G-estimation controls measured confounders through conditioning, retaining close connections with ordinary regression methods.
  • Barriers and remedies: Limited adoption largely reflects the absence of off-the-shelf software and difficulties solving estimating equations for censored survival times with SNFTMs.SAS and Stata macros are available, and newer SNCFTMs can overcome the censored-survival difficulties.
  • Advantages: SNMs and G-estimation allow greater flexibility than MSMs and typically yield better-performing estimators, especially for continuous exposures or strongly confounded binary exposures.IPW estimators typically perform poorly when treated and untreated subjects differ substantially in characteristics.
  • Extensions: The literature also connects SNMs with effect modification, mediation, and models for identifying optimal treatment sequences through regrets or treatment effects relative to no treatment.These variants model treatment blips on utility functions when subsequent treatments are optimal.
  • Future directions: The authors hope continued development of computational algorithms and software programs will make SNMs accessible to a wider audience.They aim to make the literature and methods more accessible while pointing readers to related work on extensions.
Loading 1503.01589v1…