Source-linked AI summary

SurvHTE-Bench: A Benchmark for Heterogeneous Treatment Effect Estimation in Survival Analysis

Shahriar Noroozizadeh, Xiaobin Shen, Jeremy C. Weiss, George H. Chen

arXiv:2603.05483v1cs.LGcs.AIstat.ML

TL;DR

Right-censored survival data make heterogeneous treatment effects difficult to estimate and compare because censoring, counterfactuals, and identification assumptions interact. SurvHTE-Bench provides a unified benchmark across controlled and realistic datasets, finding that estimator performance depends strongly on context, with survival-oriented methods gaining advantages as censoring or assumption violations increase. The benchmark supports fairer, reproducible evaluation while requiring domain-specific validation and attention to its scope boundaries.

  • Problem

    No standardized benchmark evaluates survival HTE methods under right-censoring, leaving comparisons and robustness assessment inconsistent.

  • Method

    SurvHTE-Bench unifies 53 survival HTE methods and evaluates them across synthetic, semi-synthetic, and real-world datasets with varied causal and survival conditions.

  • Results

    Performance is strongly context-dependent: outcome imputation methods excel in low-censoring randomized settings, whereas survival meta-learners and Causal Survival Forests gain advantages as censoring or assumptions violations increase.

  • Takeaways & Limitations

    The benchmark provides a platform for controlled stress-testing and realistic validation of estimator strengths and weaknesses under systematic assumption violations.

  • Takeaways & Limitations

    The synthetic suite omits some complexities and represents assumption violations as binary rather than across a continuum of severity.

Abstract

from arXiv · show

Estimating heterogeneous treatment effects (HTEs) from right-censored survival data is critical in high-stakes applications such as precision medicine and individualized policy-making. Yet, the survival analysis setting poses unique challenges for HTE estimation due to censoring, unobserved counterfactuals, and complex identification assumptions. Despite recent advances, from Causal Survival Forests to survival meta-learners and outcome imputation approaches, evaluation practices remain fragmented and inconsistent. We introduce SurvHTE-Bench, the first comprehensive benchmark for HTE estimation with censored outcomes. The benchmark spans (i) a modular suite of synthetic datasets with known ground truth, systematically varying causal assumptions and survival dynamics, (ii) semi-synthetic datasets that pair real-world covariates with simulated treatments and outcomes, and (iii) real-world datasets from a twin study (with known ground truth) and from an HIV clinical trial. Across synthetic, semi-synthetic, and real-world settings, we provide the first rigorous comparison of survival HTE methods under diverse conditions and realistic assumption violations. SurvHTE-Bench establishes a foundation for fair, reproducible, and extensible evaluation of causal survival methods. The data and code of our benchmark are available at: https://github.com/Shahriarnz14/SurvHTE-Bench .

1 INTRODUCTION

Heterogeneous treatment effects can reveal patient-level variation that population averages miss, but right-censored survival data add censoring and identification challenges. SurvHTE-Bench addresses the lack of standardized evaluation with a comprehensive benchmark spanning controlled and realistic settings.

  • HTE estimation can be more useful than ATE estimation when treatment effectiveness varies across individuals.
  • Right-censored survival data combine unobserved counterfactuals and confounding with incomplete observation of event times.
  • Existing causal survival studies use bespoke simulations or limited real datasets with differing censoring, survival, and causal assumptions.
  • These inconsistent evaluation practices hinder standardized comparisons, robustness assessment, and measurement of progress.
  • SurvHTE-Bench introduces a comprehensive benchmark with unified methods, 40 synthetic datasets, and semi-synthetic and real-data evaluations.

2 BACKGROUND AND RELATED WORK

The paper frames survival HTE estimation around censored potential outcomes, CATEs defined through survival-time transformations, and assumptions required for identification. It reviews imputation, direct-survival, and meta-learning approaches while emphasizing that existing evaluations remain difficult to compare.

  • For each unit, the observed data include covariates, binary treatment, a possibly censored event time, and an event indicator.
  • The target is a CATE defined through a transformation of potential event times, with RMST as the paper’s primary estimand.Other possible estimands include median survival time and survival probability at a fixed time.
  • Identification relies on consistency, ignorability, treatment positivity, ignorable censoring, and censoring positivity, whose violations are examined in the benchmark.
  • Prior evaluations vary in censoring levels and causal assumptions, leaving cross-study comparability and robustness under simultaneous violations unclear.
  • Survival HTE estimators are grouped into outcome imputation methods, direct-survival causal methods, and survival meta-learners.

3 SURVHTE-BENCH

SurvHTE-Bench evaluates survival CATE methods across synthetic, semi-synthetic, and real-world data, with synthetic settings systematically varying causal assumptions and survival dynamics. Its evaluation reports CATE accuracy, ATE bias, and auxiliary model-quality metrics across repeated splits.

  • The benchmark tests estimators under both satisfied and violated identification assumptions using synthetic, semi-synthetic, and two real-world datasets.
  • Synthetic data: The synthetic suite contains 40 datasets formed by crossing eight causal configurations with five survival scenarios.Configurations vary treatment, positivity, confounding, and censoring; scenarios vary event-time distributions and censoring rates.
  • Evaluation: Evaluation reports CATE RMSE, ATE bias, imputation MAE, regression MAE, propensity AUC, and time-dependent C-index over 10 random splits.
  • Synthetic data: The AFT noise distribution is Gaussian, producing a model that does not satisfy proportional hazards.
  • Methods: The benchmark implements 53 survival CATE variants spanning outcome imputation, direct-survival, and survival meta-learner families.

4 BENCHMARKING RESULTS

Performance is strongly context-dependent: survival meta-learners and Causal Survival Forests gain advantages as censoring intensifies or causal assumptions are violated, while outcome imputation methods excel in favorable settings. Semi-synthetic results broadly corroborate these patterns, with stability becoming especially important when RMSE differences are compressed.

  • Context dependence: In low-censoring randomized settings, outcome imputation methods excelled, whereas survival meta-learners and Causal Survival Forests gained advantages under intensified censoring or assumption violations.This overall pattern varied by causal configuration and survival scenario.
  • Overall performance: S-Learner-Survival and Matching-Survival led overall rankings, with average ranks of 5.17 and 5.42 across 53 estimator variants.At the family level, their average ranks were 3.30 and 3.48 across 11 families.
  • Causal-assumption violations: Under imbalanced treatment, Double-ML remained strong at 1.80, while T-Learner-Survival fell to last place at 9.00 because treated samples were sparse.In balanced randomized trials, Double-ML scored 3.60 and Causal Forest 5.60.
  • Causal-assumption violations: Under positivity violations alone, Double-ML and X-Learner outperformed survival meta-learners, while combining violations restored the survival meta-learners' advantage.Causal Survival Forests dropped substantially under isolated positivity violation.
  • Censoring effects: Informative censoring favored survival meta-learners and Causal Survival Forests over outcome imputation, but degraded every method and increased CATE RMSE variability.The results indicate greater estimation difficulty under dependent censoring.
  • Censoring effects: By Scenario D, S-Learner-Survival and Matching-Survival achieved average ranks of 1.6 and 2.4 under high censoring, outperforming other approaches.Survival-aware methods also gained Top-k coverage as censoring increased.
  • Semi-synthetic results: Double-ML achieved the lowest ACTG RMSE at 10.65, while survival-oriented methods were frequently among the top performers across MIMIC-i–v.The semi-synthetic results broadly corroborated the synthetic benchmark's performance patterns.
  • Semi-synthetic results: Across 53%–88% censoring in MIMIC-i–v, S-Learner-Survival remained stable with RMSE 7.897–7.921, while T-Learner-Survival became more variable under extreme censoring.In realistic covariate spaces, compressed RMSE differences make stability, interpretability, and computational cost relevant selection considerations.

5 DISCUSSION

SURVHTE-BENCH provides an extensible platform for benchmarking survival HTE estimators across synthetic, semi-synthetic, and real datasets. Its evaluations expose estimator-family strengths and weaknesses, while the benchmark remains limited in scenario coverage and treatment structure.

  • 5 DISCUSSION: SURVHTE-BENCH benchmarks survival HTE estimators across synthetic, semi-synthetic, and real datasets.The platform supports controlled stress-testing under assumption violations and validation in realistic clinical-like settings.
  • 5 DISCUSSION: The benchmark’s empirical evaluations reveal strengths and weaknesses across estimator families.
  • 5 DISCUSSION: Synthetic scenarios do not cover all complexities, and binary violations may miss the continuum of real-world assumption-violation severity.The authors suggest graded sensitivity analyses for confounding and overlap violations.
  • 5 DISCUSSION: The benchmark currently excludes time-varying treatments, instrumental variables, dynamic covariates, and multi-valued or continuous treatments.Future extensions are proposed to address more complex clinical and policy settings.

ETHICS STATEMENT

The benchmark is intended to support systematic evaluation in high-stakes healthcare, but responsible use requires safeguards beyond benchmark performance. The authors emphasize human oversight, fairness assessment, domain-specific validation, and clinical-trial evidence.

  • ETHICS STATEMENT: SURVHTE-BENCH could support systematic evaluation of survival methods for personalized medicine and clinical decision-making.The authors connect standardized evaluation and practical guidance with more informed treatment selection.
  • ETHICS STATEMENT: Misinterpreting benchmark results or overtrusting algorithms could reduce necessary human oversight in clinical contexts.
  • ETHICS STATEMENT: Performance differences across demographic groups could exacerbate healthcare disparities if ignored.
  • ETHICS STATEMENT: The benchmark should not replace domain-specific validation, fairness assessment, or clinical-trial evidence.The authors recommend fairness audits, appropriate safeguards, and validation in future applications.

REPRODUCIBILITY STATEMENT

The project provides reproducibility resources spanning benchmark design, data generation, estimator implementations, training details, and experimental results. Its modular repository and datasets are intended to support replication and extensible community evaluation.

  • REPRODUCIBILITY STATEMENT: The benchmark includes 40 synthetic datasets from 8 causal configurations and 5 survival scenarios.The design and evaluation protocol are described in the main text, with generation details in Appendix A.
  • REPRODUCIBILITY STATEMENT: The repository provides code, scripts, and READMEs for reproducing experiments across synthetic, semi-synthetic, and real-data settings.
  • REPRODUCIBILITY STATEMENT: Released materials include the complete synthetic suite, ACTG semi-synthetic data, and real-data materials for Twins and ACTG 175.For MIMIC-IV, credentialed access is required, so the project provides dataset-generation code rather than raw data.
  • REPRODUCIBILITY STATEMENT: The benchmark is modular and extensible, allowing new estimators or datasets to be added while preserving comparability.
  • REPRODUCIBILITY STATEMENT: SURVHTE-BENCH implements 53 estimator variants across outcome imputation, direct-survival, and survival meta-learner families.Appendices document estimator construction, causal methods, imputation strategies, training, and hyperparameters.
  • REPRODUCIBILITY STATEMENT: Additional materials cover synthetic, semi-synthetic, and real-world analyses, including survival-probability and RMST sensitivity results.
  • REPRODUCIBILITY STATEMENT: An appendix adds a dataset with informative censoring caused by unobserved confounding, extending the benchmark beyond its main design.

A.5 OBSERVED DATA CONSTRUCTION

Observed synthetic survival data are constructed by combining factual event times with censoring times, then recording the minimum time and an event indicator. The benchmark varies causal configurations and survival scenarios to produce 40 datasets with distinct treatment, censoring, and effect regimes.

  • A.5 OBSERVED DATA CONSTRUCTION: Observed survival data combine event and censoring times, recording the earlier time and whether the event occurred.The factual event time is T = T(W), determined by the observed treatment assignment.
  • A.5 OBSERVED DATA CONSTRUCTION: Five survival scenarios crossed with eight causal configurations yield 40 synthetic datasets testing specific combinations of survival dynamics and causal violations.
  • A.5 OBSERVED DATA CONSTRUCTION: Table 4 summarizes how event-time and censoring-time processes are generated across survival scenarios.
  • A.5 OBSERVED DATA CONSTRUCTION: Synthetic censoring rates are reported for 50,000-sample datasets and differ under informative censoring because the censoring distribution changes.
  • A.5 OBSERVED DATA CONSTRUCTION: Generator parameters calibrate censoring severity, treatment balance, and effect heterogeneity across interpretable regimes.Treatment is generally near-balanced except where imbalance is intentional, while coefficients control treatment effects and covariate interactions.
  • A.5 OBSERVED DATA CONSTRUCTION: Figure 5 displays event-time and censoring-time survival curves by causal configuration and survival scenario, with empirical censoring rates and treatment probabilities.

B IMPUTATION METHODS DETAILS

The benchmark uses three strategies to impute event times for censored subjects, converting survival outcomes into targets for evaluation or modeling. Their reliability depends on censoring-time information and assumptions about censoring and sample size.

  • Margin imputation: Margin imputation uses the Kaplan-Meier estimator to assign a conditional-mean surrogate event time after censoring.The estimate is based on the survival curve from the training dataset.
  • Limitations: Margin imputation is highly uncertain after very early censoring but is more likely to approach the true event time near maximum follow-up.Its reliability therefore varies with the censoring time.
  • IPCW-T imputation: IPCW-T imputation averages observed event times from uncensored subjects whose observed times exceed the censoring time.It uses subsequent uncensored individuals as empirical evidence about the unobserved event timing.
  • Pseudo-observation imputation: Pseudo-observation imputation estimates each subject’s contribution using a full-sample estimator and the corresponding leave-one-out estimator.The pseudo-observation is defined as N · ˆθ − (N − 1) · ˆθ−i.
  • Pseudo-observation imputation: Pseudo-observations are substituted for true event times after being computed for censored subjects.They can approximate E[Ti | Xi] under certain assumptions and large samples, supporting nonparametric imputation of censored times.
  • Implementation: Imputed event times are constrained to be at least as large as the censoring time, reflecting that the event must occur afterward.This constraint is manually enforced in the implementation.

C LIST OF CATE ESTIMATORS IN SURVHTE BENCHMARK

The benchmark evaluates survival-CATE estimators across outcome-imputation, direct-survival, and survival-meta-learning families. These configurations combine imputation strategies, causal learners, and regression or survival base models, totaling 53 variants.

  • Estimator families: The benchmark evaluates 53 survival-CATE method configurations across three estimator families.The families are outcome imputation, direct-survival CATE models, and survival meta-learners.
  • Outcome imputation: Outcome-imputation methods combine three imputation strategies with four meta-learners and three regression models, producing 36 variants.The meta-learners are S-, T-, X-, and DR-Learners; the base models are Lasso, Random Forest, and XGBoost.
  • Direct-survival CATE: Direct-survival models include Causal Survival Forests and SurvITE, which handle right-censored data without separate imputation.SurvITE learns balanced representations and optimizes a survival-specific loss.
  • Survival meta-learners: Survival meta-learners comprise S-, T-, and matching-learners paired with Random Survival Forests, DeepSurv, or DeepHit, yielding 9 variants.These frameworks are extended to handle censored data directly.
  • Meta-learners: The T-Learner fits separate outcome models for treated and control groups before estimating treatment effects.Its conceptual simplicity can be accompanied by high variance with unequal treatment-group sizes or poor extrapolation under limited overlap.
  • Meta-learners: The S-Learner fits one model using all data and includes treatment assignment as an additional feature.Its performance depends on capturing treatment-feature interactions with the base learner.
  • Meta-learners: The X-Learner combines imputed treatment effects with propensity-score weighting, which is useful for unequal treatment-group sizes or heterogeneous effects.The propensity score is typically estimated with logistic regression.
  • Meta-learners: The DR-Learner combines outcome and propensity modeling through doubly robust scores that remain consistent if either model is correctly specified.Its procedure includes explicit propensity-score modeling and doubly robust outcome construction.

E.4 COMPUTATION TIME OF SURVIVAL CATE METHODS

The benchmark measures computational cost as average runtime per dataset and experimental repeat, alongside broad performance summaries across causal and survival settings. Neural-network survival models are reported to be substantially more computationally expensive than classical or tree-based methods.

  • Computation time: Runtime is averaged across 40 synthetic datasets and 10 random seeds, excluding imputation time.Measurements use Python’s time.time() and report mean seconds with standard deviation.
  • Computation time: Neural-network survival models incur substantially higher computational costs than classical or tree-based methods.The comparison is reported in the runtime analysis accompanying Table 13.
  • Performance comparison: Win-rate analyses report how often method families reach Top-1, Top-3, and Top-5 positions for CATE RMSE and ATE Bias.These rates complement average rankings by emphasizing frequency of strong performance.
  • Performance comparison: Borda rankings aggregate test-set CATE RMSE across every causal-configuration and survival-scenario combination.All 53 methods are ranked for each configuration-scenario pair before aggregation.
  • Results: S-Learner-Survival achieves CATE RMSE Top-3 and Top-5 rates of 67.5% and 85.0%, respectively.It also has ATE Bias Top-3 and Top-5 rates of 57.5% and 77.5%.
  • Results: Matching-Survival has a 92.5% Top-5 rate on CATE RMSE, while Causal Survival Forests has a 25.0% Top-1 rate among direct-survival methods.Classical outcome-imputation meta-learners attain Top-1 positions only rarely in the reported settings.

F.2 RANKING OF CAUSAL METHODS FOR DIFFERENT SURVIVAL SCENARIOS

Method rankings vary with censoring, survival dynamics, causal assumptions, and data structure. Survival-aware methods generally gain an advantage as censoring or assumption violations intensify, while no family dominates every setting.

  • Higher censoring shifts the highest rankings toward direct-survival methods and survival meta-learners, while Double-ML declines from rank 1.5 in Scenario A to rank 6.9 in Scenario E.S-Learner-Survival, Matching-Survival, and Causal Survival Forests dominate later scenarios.
  • Under low censoring, winners depend on the survival scenario: Double-ML leads CATE RMSE in Scenario A, whereas Causal Survival Forests leads both metrics in Scenario B.Scenario A gives Double-ML 62.5% Top-1 CATE RMSE; Scenario B gives Causal Survival Forests 37.5% Top-1 CATE RMSE and 25.0% Top-1 ATE Bias.
  • Under medium and high censoring, survival-aware methods occupy more top-ranked positions, with Causal Survival Forests leading Scenario C and S-Learner-Survival leading Scenario D on CATE RMSE.Matching-Survival and T-Learner-Survival also achieve strong coverage across the high-censoring scenarios.
  • Outcome imputation performs best in randomized settings, but survival meta-learners and Causal Survival Forests rise under unmeasured confounding, informative censoring, and multiple simultaneous violations.In RCT-5%, Double-ML achieves top rank 1.80; under multiple violations, S-Learner-Survival and Matching-Survival take the top two spots.
  • In randomized configurations, CATE RMSE is split between imputation baselines and survival meta-learners, whereas ATE Bias tends to favor survival-aware approaches.Under RCT-50, several methods tie at 20.0% Top-1 CATE RMSE, while survival-aware methods provide stronger Top-k coverage.
  • Across semi-synthetic datasets, method rankings depend on covariate structure, censoring, mechanism complexity, and horizon, with variability becoming important under extreme censoring.ACTG and MIMIC show different ranking patterns, while MIMIC-i–v and MIMIC-vi–ix indicate that stability and mechanism complexity affect evaluation.

G.4.4 DETAILED ANALYSIS AND PRACTICAL IMPLICATIONS

The semi-synthetic analysis finds no universally best method across ACTG and MIMIC. Performance differences depend on data structure, censoring, stability, estimand, and time horizon.

  • No single method dominates across ACTG and MIMIC, because rankings shift with covariate support, censoring regime, and treatment-assignment mechanisms.Flexible doubly robust estimators perform strongly in ACTG, while survival-oriented approaches are frequently competitive in MIMIC.
  • Under extreme MIMIC censoring, variability across repetitions can distinguish methods even when mean RMSE values are close.Some approaches show noticeably higher variability, while others remain stable across censoring levels.
  • Earlier horizons separate methods more clearly than later horizons across MIMIC variants, while RMST targets can compress performance differences by integrating over time.Shortening the RMST horizon changes error scale but rarely overturns broad rankings.
  • In trial-like ACTG settings, flexible causal estimators can be accurate; in highly censored, high-dimensional MIMIC settings, survival-oriented and stable meta-learner methods are often competitive.The practical guidance emphasizes stability under censoring when selecting methods for EHR-like data.
  • CATE RMSE comparisons in Table 35 report mean ± standard deviation over 10 repeats for RMST horizons h = Tmax and h = Tmed.

H.1 TWINS DATASET

The Twins benchmark uses twin birth and first-year mortality data to construct an observational survival dataset with approximate individual-level ground truth. The ACTG HIV trial provides randomized treatment comparisons under baseline and heavily injected censoring, where Causal Survival Forests remains comparatively stable.

  • The Twins dataset treats being born the heavier twin as treatment and first-year mortality time as the outcome, administratively censored at 365 days.The dataset includes more than 11,000 twin pairs after restricting weight and missingness.
  • Twin siblings provide paired potential outcomes for the treatment effect, but the resulting ground truth assumes the unobserved outcome equals the sibling’s observed outcome.This approximation may not fully capture genetic or environmental heterogeneity.
  • The observational Twins construction uses logistic treatment assignment and exponential censoring, producing a 68.1% treatment rate and 84.8% censoring rate.
  • Twins experiments use 50/25/25 training, validation, and testing splits, repeat experiments across 10 random splits, and report test-set CATE RMSE.Results are shown for horizons h = 30 and h = 180 days, with similar findings reported for the two horizons.
  • ACTG 175 compares ZDV against ZDV+ddI, ZDV+Zal, and ddI, with additional censoring increasing rates from below 15% to above 90%.Baseline CATE estimates are formed by averaging 10 Causal Survival Forests runs.
  • Across HIV comparisons, Causal Survival Forests estimates cluster near baseline after censoring injection, whereas outcome-imputation methods show greater variation.Figures 22 and 23 present the ZDV versus ZDV+Zal and ZDV versus ddI comparisons.

I ADDITIONAL INFORMATIVE CENSORING VIA UNOBSERVED CONFOUNDING

This extension evaluates survival HTE estimators when censoring depends on an unobserved confounder, using a modular OBS-UConf design in Scenario C. Causal Survival Forests and matching-based survival meta-learners tend to perform best, while broader extensions remain future work.

  • Data generation: The alternative informative-censoring mechanism makes censoring depend on an unobserved variable that also affects treatment and outcomes.This violates ignorable censoring because observed covariates alone cannot explain the censoring–survival dependence.
  • Data generation: The experiment combines OBS-UConf with survival Scenario C, which uses Poisson hazards and medium censoring.The observed covariates and latent confounder are uniformly distributed, while treatment is assigned observationally.
  • Summary statistics: The setup includes up to 50,000 samples, a 53.9% treatment rate, 39.7% censoring rate, and population-level ATE of 0.7737.The censoring rate is driven by the unobserved confounder U.
  • Experimental results: CATE RMSE and ATE bias are evaluated across 10 random splits using representative estimators from all three method families.Figure 24 reports mean ± standard error for both metrics under informative censoring induced by unobserved confounding.
  • Experimental results: Causal Survival Forests and survival meta-learners with matching tend to perform best under this setting.The pattern is consistent with findings from the main synthetic datasets.
  • Extensibility to other settings: The same censoring mechanism can be extended to other causal configurations and survival scenarios, but systematic exploration is left for future work.Examples include randomized trials with imbalance and AFT or Cox models.
Loading 2603.05483v1…