Source-linked AI summary
A causal framework for classical statistical estimands in failure time settings with competing events
Jessica G. Young, Mats J. Stensrud, Eric J. Tchetgen Tchetgen, Miguel A. Hernán
TL;DR
Classical competing-risks estimands lacked a formal causal framework for interpreting effects and identifying assumptions. This paper uses counterfactual definitions and causal diagrams to show that risk contrasts can represent total or direct effects, whereas hazard contrasts generally lack causal interpretations.
Problem
Classical competing-risks estimands and their estimating procedures lacked a formal framework for characterizing causal effects and identifying conditions.
Method
The paper uses a counterfactual framework and causal diagrams representing competing events as time-varying covariates to define estimands and formalize identification conditions.
Results
Risk contrasts can quantify total or direct treatment effects depending on whether competing events are treated as censoring, but hazard contrasts generally cannot be interpreted as causal effects.
Takeaways & Limitations
Competing events may need to be treated as time-varying covariates for exchangeability in total-effect analyses, including randomized trials.
Takeaways & Limitations
Identifying direct effects may require exchangeability assumptions about censoring and changing prognostic factors that are often implausible, especially when competing events are censoring events.
Abstract
from arXiv · showhide
In failure-time settings, a competing risk event is any event that makes it impossible for the event of interest to occur. For example, cardiovascular disease death is a competing event for prostate cancer death because an individual cannot die of prostate cancer once he has died of cardiovascular disease. Various statistical estimands have been defined as possible targets of inference in the classical competing risks literature. Many reviews have described these statistical estimands and their estimating procedures with recommendations about their use. However, this previous work has not used a formal framework for characterizing causal effects and their identifying conditions, which makes it difficult to interpret effect estimates and assess recommendations regarding analytic choices. Here we use a counterfactual framework to explicitly define each of these classical estimands. We clarify that, depending on whether competing events are defined as censoring events, contrasts of risks can define a total effect of the treatment on the event of interest, or a direct effect of the treatment on the event of interest not mediated through the competing event. In contrast, regardless of whether competing events are defined as censoring events, counterfactual hazard contrasts cannot generally be interpreted as causal effects. We illustrate how identifying assumptions for all of these counterfactual estimands can be represented in causal diagrams in which competing events are depicted as time-varying covariates. We present an application of these ideas to data from a randomized trial designed to estimate the effect of estrogen therapy on prostate cancer mortality.
1 Introduction
Competing risks research has multiple classical estimands, but prior reviews lacked a formal causal framework for interpreting their effects and identifying assumptions. This paper uses counterfactual definitions and causal diagrams to clarify these estimands and analytic choices.
- A competing event makes the event of interest impossible; cardiovascular disease death therefore competes with prostate cancer death.
- Classical competing-risks estimands include marginal and cause-specific cumulative incidence, marginal and subdistribution hazards, and cause-specific hazard.
- Prior reviews described these estimands and their estimation procedures without formally characterizing causal effects or identifying conditions.
- The paper uses a counterfactual framework to define classical estimands, distinguish total and direct effects, and represent identifying assumptions in causal diagrams.
2 Observed data structure
The observed-data structure follows randomized treatment assignment and longitudinal measurement of event histories, competing events, covariates, and censoring. Competing events deterministically preclude later observation of the event of interest.
- Individuals are randomly assigned to estrogen therapy or placebo at baseline, with equally spaced follow-up intervals through a prespecified maximum.
- Y_k and D_k record histories of the event of interest and competing event, while L_0 and L_k represent baseline and time-varying characteristics.
- Within each follow-up interval, competing-event, event-of-interest, and covariate measurements are temporally ordered as (D_k, Y_k, L_k).
- The analysis assumes variables are measured without error and initially assumes no loss to follow-up.
- After a competing event occurs without prior event-of-interest occurrence, all future event-of-interest indicators are known and deterministically zero.
3 Counterfactual estimands when competing events do not exist
Without competing events, the paper defines counterfactual risks and hazards under treatment interventions. Risks support causal contrasts, whereas hazard contrasts generally do not because treatment can alter who remains at risk.
- The no-competing-event case treats death from any cause as the event of interest before extending the framework to competing events.
- For each treatment level, the counterfactual outcome Y^a_{k+1} indicates whether the event of interest occurs by interval k+1 under assignment A=a.
- Counterfactual risk is the population probability of the event of interest by interval k+1 under an intervention assigning all individuals to treatment level a.
- The discrete-time hazard measures event occurrence in interval k+1 conditional on surviving to k, and approaches the continuous-time hazard as interval length shrinks.
- Hazard contrasts generally lack causal interpretation because treatment may change the composition of individuals surviving to k, even when the contrasts are identifiable.
4 Counterfactual estimands when competing events exist
The paper defines competing-risk estimands through counterfactual risks and hazards, distinguishing effects that do or do not operate through competing events. Risk contrasts can represent total or direct effects, whereas hazard contrasts generally lack causal interpretations.
- 4.1 Direct effects: Counterfactual risks distinguish hypothetical elimination of competing events from risks observed without eliminating them.The former is the marginal cumulative incidence or net risk; the latter is the subdistribution, cause-specific cumulative incidence, or crude risk.
- 4.1 Direct effects: A risk contrast under elimination of competing events quantifies the treatment effect on the event of interest not mediated through competing events.The paper identifies this contrast as a controlled direct effect.
- 4.2 Total effects: A risk contrast without eliminating competing events quantifies the total treatment effect on the event of interest, including pathways mediated by competing events.The competing event may mediate treatment effects through paths such as A → D_k → Y_k.
- 4.2 Total effects: For mutually competing outcomes, the treatment effect on the composite outcome equals the sum of the effects on the event of interest and the competing event.The composite example is death from any cause rather than prostate cancer death alone.
- 4.3 Counterfactual hazards: Three counterfactual hazard contrasts are defined, including marginal, subdistribution, and cause-specific hazards.The subdistribution hazard retains individuals who experienced a competing event in its risk set, whereas the cause-specific hazard conditions on being event-free.
- 4.3 Counterfactual hazards: None of the counterfactual hazard contrasts can generally be interpreted as causal effects because treatment can alter who survives into later risk sets.Thus hazards may differ because the treatment groups contain different survivors at time k.
5 Definition of a censoring event
The paper defines censoring relative to the counterfactual outcomes required by the chosen estimand. Consequently, competing events may be censoring events for direct-effect estimands but not for total-effect estimands.
- 5 Definition of a censoring event: A censoring event is an event that makes future counterfactual outcomes of interest unknown, even for an individual receiving intervention a.The definition is tied to the investigator-chosen estimand and its required counterfactual outcomes.
- 5 Definition of a censoring event: Competing events are censoring events when the estimand depends on outcomes under treatment and an intervention eliminating competing events.This applies to direct effects and corresponding hazard contrasts involving the eliminated-competing-event outcome.
- 5 Definition of a censoring event: Competing events are not censoring events when the estimand depends on outcomes under treatment without eliminating competing events.Individuals who experience a competing event cannot subsequently experience the event of interest by definition.
- 5 Definition of a censoring event: Loss to follow-up is always a censoring event, and identification generally requires indexing counterfactual outcomes by an intervention that eliminates it.If loss to follow-up does not affect future events, results may also apply without that intervention index.
- 5 Definition of a censoring event: Administrative study-end survival is not a censoring event under the paper’s definition when follow-up ends at or before the administrative study end.Left-censored failure times require particularly strong assumptions for causal inference.
- 5 Definition of a censoring event: The framework does not require a potential censoring time for individuals observed to fail during the study period.After a competing event, the absence of a later event preventing knowledge of the event of interest does not create subsequent censoring.
6 Identification of estimands when competing events exist
The paper identifies competing-risk estimands under assumptions represented with causal diagrams. Risk contrasts can identify direct or total effects, whereas hazard contrasts may be identified without generally having causal interpretations.
- Identification conditions depend on whether competing events are treated as censoring events.The paper considers risks when loss to follow-up and competing events may occur, with separate conditions for different causal estimands.
- 6.1 Direct effects: Identifying the risk under elimination of competing events requires consistency, exchangeability, and positivity assumptions.Consistency may be problematic when elimination requires an unspecified intervention on death from other causes.
- 6.1 Direct effects: The g-formula identifies risk under elimination of competing events and has algebraically equivalent inverse probability weighted representations.The weights account for remaining free of loss to follow-up and competing events conditional on measured history.
- 6.1 Direct effects: Baseline-only exchangeability requires censoring to be independent of changing prognostic factors during follow-up, an assumption often considered implausible when competing events are censored.This stronger assumption is especially problematic because censoring and competing events are not randomly assigned by design.
- 6.2 Total effects: The same assumptions identify direct effects and hazards under elimination of competing events, but these assumptions can fail when the competing event shares an unmeasured cause with the event of interest.The causal DAGs distinguish assumptions supporting direct-effect identification from those supporting total-effect identification.
- 6.3 Counterfactual hazards and the “hazards of hazard ratios” revisited: All three counterfactual hazard contrasts may be identified under the observed data structure, yet hazard contrasts do not generally have causal interpretations.Conditioning on prior survival can induce a non-causal path through a common cause of treatment and the event of interest.
7 Choosing a definition of causal effect when competing events exist
The paper distinguishes total and direct effects when competing events exist. Total effects include pathways through competing events, while direct effects exclude those pathways but may require difficult or unrealistic interventions.
- Total and direct effects are both contrasts in counterfactual risks, differing in whether competing events mediate treatment effects.The direct effect excludes mediation through the competing event; the total effect includes all causal pathways.
- In the extreme example, the total effect appears protective against prostate cancer death because estrogen therapy instantly causes cardiovascular death.The example produces a negative risk difference for prostate cancer death while the treatment increases risk of the competing event.
- Reporting effects on both the event of interest and the competing event can reveal when a total-effect contrast has an uninformative interpretation.In the example, the effect on prostate cancer death is negative while the effect on the competing event is positive.
- A negative direct effect cannot be explained by a harmful treatment effect on the competing event, but defining it requires well-defined interventions eliminating competing events.Such interventions are often difficult to imagine for deaths from other causes.
- When competing-event interventions are impossible or unrealistic, the untestable assumptions required for direct-effect identification may be especially questionable.Positivity can also fail when recovered patients have essentially no probability of avoiding discharge.
- Composite outcomes avoid interventions on competing events and can be identified by randomization without loss to follow-up, but they may answer a different question.For prostate cancer death, the composite outcome instead quantifies all-cause mortality.
8 Estimation of direct and total effects
The paper describes parametric g-formula and inverse probability weighting approaches for estimating direct and total effects. These approaches differ in their modeling requirements and representations of censoring and competing events.
- The parametric g-formula estimates conditional components and uses Monte Carlo simulation to approximate sums or integrals over risk-factor histories.It can use discrete-time or continuous-time hazard models, including pooled logistic, proportional hazards, or additive hazards models.
- Inverse probability weighting estimates g-formula contrasts through weighted product-limit or weighted hazard procedures.The paper presents representations based on loss-to-follow-up weights and, for some estimands, competing-event hazards.
- IPW consistency requires correct specification of weight-denominator models, whereas the parametric g-formula generally requires correct outcome-hazard and covariate-distribution models.This distinction motivates alternative estimators with fewer model assumptions.
- Estimating risk without eliminating competing events requires models for the observed cause-specific hazards of both the event of interest and the competing event.This distinguishes the corresponding g-formula from the version that treats competing events as censoring.
- When exchangeability is assumed given only baseline covariates, the estimation algorithms simplify by replacing functions of time-varying covariates with functions of baseline covariates.The paper notes this simplification for both parametric g-formula and IPW approaches.
9 Example: A randomized trial on prostate cancer therapy
A randomized DES-versus-placebo trial illustrates estimation of total and direct effects on prostate cancer death, competing-event mortality, and composite mortality. The results suggest protective total effects on prostate cancer death, potentially offset by harmful effects on other-cause mortality, while direct effects are smaller and require strong assumptions.
- Analysis and estimands: The analysis used parametric g-formula and IPW estimators based on pooled logistic models for prostate cancer death, other death, and loss to follow-up.The models estimated observed cause-specific hazards for the event of interest, competing event, and censoring.
- Analysis and estimands: Before month 50, no loss to follow-up made the competing-event-free risk estimators identical to the nonparametric cumulative proportion estimator.Both IPW estimators incorporated the zero censoring hazard before month 50.
- Total effects: By 60 months, all point estimates suggested a total protective effect of estrogen therapy versus placebo on prostate cancer death, although confidence intervals were wide.Treatment and placebo risk curves began diverging shortly after follow-up began.
- Competing-event effects: Estimates of other-cause mortality were examined because deaths from competing causes could partly explain the apparent protective total effect on prostate cancer death.The competing-event risk was estimated under treatment and placebo using the parametric g-formula and two IPW versions.
- Composite outcome: The composite all-cause mortality risk stayed close to the null through most follow-up, consistent with opposing total-effect directions for prostate cancer and other-cause death.The composite estimate was formed by summing corresponding risks for prostate cancer death and other causes of death.
- Direct effects: The direct effect on prostate cancer death, treating other-cause death as censoring, was smaller by 60 months than the total effect and varied over time.These estimates relied on the same pooled logistic model assumptions as the preceding analysis.
- Direct effects: Direct-effect estimates require caution because confidence intervals were wide, model misspecification was possible, and the identifying assumptions were especially strong and unverifiable in real-world studies.The exchangeability assumption based only on baseline covariates was particularly strong because post-randomization prognostic factors were not measured.
10 Discussion
The paper uses counterfactual definitions to distinguish total and direct effects in competing-risk analyses and argues that hazard contrasts generally lack causal interpretation. It formalizes identification conditions, illustrates them with causal diagrams and estimators, and applies them to an estrogen trial.
- Causal estimands: When competing events are censoring events, risk contrasts quantify direct effects; otherwise, they quantify total effects through all treatment-to-outcome pathways.The distinction depends on whether the competing event is included among censoring events.
- Identification: Identification of total effects with loss to follow-up requires competing events as time-varying covariates, alongside baseline and other time-varying covariates, even in randomized trials.The paper links identifying functions for total and direct effects to different versions of Robins’s g-formula.
- Identification: Causal diagrams with competing events as time-varying covariates can assess exchangeability assumptions using established graphical identification rules.The framework makes the estimand explicit before evaluating the assumptions needed for identification.
- Application: The estrogen-trial application had only baseline covariates, making its exchangeability assumptions less plausible, although the framework extends to observational studies.The trial application was used to motivate the ideas rather than restrict their scope to randomized studies.
- Hazards: Contrasts of counterfactual hazards generally lack causal interpretation, even when assumptions identifying corresponding contrasts or direct effects hold.The paper does not recommend routine use of any hazard as a causal-effect measure when competing events exist.
- Causal estimands: Competing events may mediate treatment effects, so total effects can have uncertain interpretation when treatment also affects the competing event.The paper therefore considers effects that do not capture treatment effects on competing events, while noting their drawbacks.
- Implications: Existing effects that exclude treatment’s effect on competing events require either ill-defined interventions or well-defined interventions in an unknowable subpopulation.The paper identifies new effect definitions avoiding these problems as an important area for future work.
A Single-world intervention graphs
The section constructs SWIG templates by splitting intervention nodes into natural and fixed values, then uses these graphs to represent counterfactual interventions and assess exchangeability. Different causal DAGs imply different exchangeability conclusions.
- SWIG construction: SWIGs transform a causal DAG by splitting each intervention node into its natural value and its fixed intervention value.The intervention variables here include treatment and censoring at each time, with censoring potentially including loss to follow-up and competing events.
- SWIG construction: The transformed graph indexes post-treatment random variables as counterfactuals under the intervention while preserving the corresponding causal arrows.This yields a graphical template for evaluating identification assumptions under specified interventions.
- Figure 8: Figure 8 represents an intervention setting treatment to a and eliminating loss to follow-up and competing events, with exchangeability implied under the Figure 1 DAG.The graph is a template for interventions in which the censoring indicators are set to fixed values.
- Figure 8: Under the Figure 1 data-generating assumptions, exchangeability follows from the absence of specified unblocked backdoor paths on the SWIG template.The relevant paths concern treatment, competing events, loss to follow-up, outcomes, and measured covariates.
A Identification Proofs
The proofs establish that counterfactual risks and hazards under specified interventions equal corresponding g-formula expressions when the stated consistency, positivity, and exchangeability assumptions hold.
- Main identification result: Theorem 1 provides the central identification result from which the paper’s risk and hazard corollaries follow.The proofs use the stated identifying assumptions and repeated probability and expectation arguments.
- Risk identification: The g-formula expressions can be rewritten using vectors of censoring indicators and measured time-varying covariates.The proof derives equivalent representations through lemmas and iterative conditioning.
- Risk identification: Under assumptions (20)–(22), risk under elimination of competing events equals g-formula (23), while risk without elimination equals g-formula (29) under assumptions (26)–(28).These corollaries connect the two counterfactual risk definitions to distinct g-formula expressions.
- Hazard identification: Under the corresponding assumptions, counterfactual hazards with and without elimination of competing events equal expressions (25) and (32), respectively.The hazard results follow from the identified risk representations by selecting the relevant intervention definitions.
- Extended identification: Additional corollaries identify hazards conditioned on competing events and the competing-event hazard itself under joint assumptions.The results also accommodate mutually competing and semi-competing-risk settings.
A.1 Parametric g-formula and IPW estimators of the risk of the competing event
The appendix describes estimators and data structures for competing-event risks, including parametric g-formula and IPW approaches, and discusses alternative direct-effect estimands and their population limitations.
- A.1 Parametric g-formula and IPW estimators of the risk of the competing event: Parametric g-formula and IPW estimators of competing-event risk use estimated observed cause-specific hazards, with analogous algorithms for mutually competing events.An IPW estimator can alternatively use weighted estimates of these cause-specific hazards.
- A.1 Parametric g-formula and IPW estimators of the risk of the competing event: In semi-competing-risk settings, the original competing event is treated differently, and methods for settings without competing events can be used when its indicator is set to zero.The appendix distinguishes risk estimation for mutually competing and semi-competing events.
- A.2.1 Structure of input data sets: The data application used person-time data, with one dataset for g-formula and cause-specific-hazard IPW estimators and another retaining individuals after competing events for subdistribution-hazard IPW.A third dataset reversed the event and competing-event roles to estimate competing-event effects.
- A.2.2 Model assumptions: The estimators were implemented using different person-time datasets depending on whether competing events were removed or retained in the risk set.Records retaining failed individuals require model fitting restricted to those who have not yet experienced the modeled competing event.
- A.2.2 Model assumptions: The observed prostate-cancer and other-death hazards were modeled with pooled over-time logistic regressions using treatment, time functions, and baseline covariates.The prostate-cancer hazard used a third-degree month polynomial; the other-death hazard used a second-degree polynomial.
- A.2.2 Model assumptions: Loss-to-follow-up hazards were set to 1 before month 50 because no patient was lost before that time, then modeled for later records.The model included treatment and selected baseline covariates; hemoglobin was excluded because of convergence failures in bootstrap samples.
- A Cross-world alternatives to controlled direct effects: A controlled direct effect can target effects not mediated by competing events, but it concerns an unknown subset because each person’s competing-event status is observed under only one treatment level.The appendix contrasts this limitation with the survivor average causal effect’s restriction to a different population.
- A Cross-world alternatives to controlled direct effects: The survivor average causal effect avoids interventions on competing events by restricting to individuals who would avoid them under either treatment, but that subset is unknown and its size is not known.This is a major drawback even when identifying assumptions for the estimand hold.