Source-linked AI summary
High-dimensional variable selection for Cox's proportional hazards model
Jianqing Fan, Yang Feng, Yichao Wu
TL;DR
High-dimensional survival analysis requires selecting important covariates when their number can exceed the sample size, especially when marginal screening misses jointly informative variables. The paper extends sure independence screening to Cox’s model through iterative conditional screening and penalized selection, with simulations showing strong performance, particularly for correlated covariates and challenging joint effects.
Problem
Variable selection in Cox’s proportional hazards models is difficult when dimensionality exceeds sample size and important covariates may be marginally unrelated or obscured by correlated variables.
Method
The method iteratively combines conditional marginal screening with penalized partial-likelihood selection, using variants that control how aggressively variables are screened before penalization.
Results
Simulations found that ISIS performed well for dependent covariates and challenging jointly informative predictors, while LASSO often retained substantially larger models and had larger estimation errors.
Takeaways & Limitations
The iterative sure independence screening scheme provides a computationally efficient and specific variable-selection approach for high-dimensional Cox regression.
Takeaways & Limitations
The setup assumes conditionally independent, noninformative censoring, and the simulations use exponentially distributed censoring with mean 10.
Abstract
from arXiv · showhide
Variable selection in high dimensional space has challenged many contemporary statistical problems from many frontiers of scientific disciplines. Recent technology advance has made it possible to collect a huge amount of covariate information such as microarray, proteomic and SNP data via bioimaging technology while observing survival information on patients in clinical studies. Thus, the same challenge applies to the survival analysis in order to understand the association between genomics information and clinical information about the survival time. In this work, we extend the sure screening procedure Fan and Lv (2008) to Cox's proportional hazards model with an iterative version available. Numerical simulation studies have shown encouraging performance of the proposed method in comparison with other techniques such as LASSO. This demonstrates the utility and versatility of the iterative sure independent screening scheme.
2. Cox’s proportional hazards models
The Cox proportional hazards model represents covariate-dependent hazard through a baseline hazard and regression coefficients, while estimating the baseline hazard nonparametrically from observed failure times. Cox’s partial likelihood enables estimation of the regression parameters before constructing the baseline cumulative hazard estimate.
- Model and data: Observed survival data include event times, censoring indicators, and covariates under conditionally independent, noninformative censoring.The observed time is Y=min{T,C}, with δ indicating whether failure occurs before censoring.
- Model and data: The proportional hazards model specifies h(t|x)=h0(t) exp(xT β), with unknown baseline hazard h0(t) and regression vector β.Both components must be estimated from the censored survival data.
- Baseline hazard estimation: Breslow’s nonparametric construction represents cumulative baseline hazard as jumps at observed failure times.The jump sizes are obtained from the likelihood-based estimation scheme.
- Estimation: Cox’s partial likelihood results after substituting the estimated baseline-hazard increments into the likelihood.This removes the unspecified baseline hazard from the regression-parameter optimization.
- Estimation: Maximizing the log-likelihood estimates β, after which fitted hazard increments produce a nonparametric estimate of the baseline cumulative hazard.The regression estimate is plugged into the estimated increments before cumulative-hazard construction.
3. Variable selection for Cox’s proportional hazards model via penalization
Penalized partial likelihood provides variable-selection capability for Cox regression, addressing the failure of ordinary estimation to set coefficients exactly to zero and to handle p>n. The paper focuses on SCAD-based extensions of SIS and ISIS, solved with local linear approximation when needed.
- Motivation: Ordinary coefficient estimation leaves all covariates in the final model, so it cannot select important variables or handle p>n.This motivates variable-selection methods for high-dimensional Cox regression.
- Penalized selection: Penalization regularizes the objective with variable-selection-capable penalties, including LASSO, SCAD, elastic net, adaptive L1, and minimax concave penalties.The paper treats penalized likelihood as equivalent to penalized partial likelihood.
- SCAD penalty: The paper uses the SCAD penalty for its SIS and ISIS extensions when necessary.SCAD is a symmetric quadratic spline with a recommended parameter value a=3.7.
- SCAD penalty: SCAD is non-convex, so the paper uses a local linear approximation algorithm for SCAD-penalized optimization when needed.The non-convexity distinguishes SCAD optimization from convex penalization procedures.
4. SIS and ISIS for Cox’s proportional hazard model
The paper extends SIS and ISIS to Cox regression by ranking covariates through marginal or conditional partial-likelihood utility, then applying penalized selection iteratively. Variants based on sample splitting and intersection aim to reduce false selection, while ISIS addresses predictors whose importance emerges only jointly.
- Marginal screening: Sure screening means that a selected model of size o_p(n) contains the true model with probability tending to one.This property formalizes retention of all important variables during screening.
- Marginal screening: SIS ranks covariates by marginal utility and retains the top d variables, aiming to include the sparse true model with high probability.The marginal utility is the maximized partial likelihood for a single covariate.
- Iterative selection: ISIS addresses jointly related predictors that are marginally unrelated or obscured by other predictors with larger marginal correlations.It uses conditional information beyond the marginal ranking used by SIS.
- Iterative selection: Each ISIS iteration screens variables conditionally on the current selected set, then applies penalized partial likelihood to update the sparse model.The conditional utility measures a candidate’s additional contribution given previously selected variables.
- Iterative selection: The iterative update can remove variables selected earlier, and iterations stop after reaching d covariates or unchanged successive models.This allows later conditional information to revise earlier screening decisions.
- False-selection control: Random-splitting variants apply SIS or ISIS to two sample halves and intersect the resulting sets to reduce false selections.Under exchangeability conditions, the associated probability bound for including r unimportant covariates decreases with dimensionality.
- False-selection control: The second sample-splitting variant chooses larger half-sample screening sets so their intersection contains d covariates before penalization.It is therefore less aggressive than the first variant.
5. Simulation
The simulations compare SIS variants, ISIS, SCAD-refined models, and LASSO across increasingly difficult Cox-model settings. ISIS performs especially well when covariates are correlated or important variables are marginally uninformative, while LASSO retains larger models and incurs larger estimation errors.
- Simulation design: The study compares SIS and ISIS variants with LASSO across six Cox-model simulation cases, using 100 Monte Carlo repetitions and several estimation and screening measures.The reported measures include median L1 and squared L2 estimation errors, sure-screening proportions before and after SCAD, and median final-model size.
- Simulation results: Vanilla-SIS performs reasonably with independent predictors but poorly with dependent predictors, whereas vanilla-ISIS and its second variant perform very well in both cases.ISIS improves substantially over SIS under covariate dependence in both inclusion of all true variables and estimation error.
- Simulation results: LASSO has the sure screening property but a median model size ten times larger than ISIS, accompanied by larger L1 and squared L2 estimation errors.The simulations attribute the larger model to many small nonzero coefficients selected for unimportant variables.
- Simulation results: ISIS selects X4 in every repetition for the specially designed Cases 3 and 4, while LASSO rarely includes it; ISIS also performs well when X5 has a small coefficient.These results indicate that ISIS uses joint covariate information in settings that defeat marginal screening.
- Simulation results: Adding 20 noisy variables increases difficulty, with the largest impact in Cases 3, 4, and 6.Cases 1 and 2 are relatively easy in the oracle model, whereas Cases 3 and 4 are harder.
6. Real data
The Neuroblastoma analysis applies the proposed selection method to gene-expression and survival data from 246 patients, identifying eight genes whose model contribution exceeds that of randomly added genes.
- Data: The Neuroblastoma dataset contains gene-expression measurements at 10,707 probe sites and survival information for 246 patients after removing five outlier arrays.The 251 patients were diagnosed between 1989 and 2004; the analysis focuses on overall survival.
- Selected-gene model: The eight-gene Cox model has log-(partial)likelihood -129.3517, compared with -215.4561 for the null model, with an obvious χ2-test significance.Estimated coefficients are reported in Table 5, and the estimated baseline survival function is plotted in Figure 5.
- Selected-gene model: Removing one selected gene reduces the average log-likelihood to -141.2948, a decrease of 11.9431, and all selected genes except AP31816 and Hs150167.1 show significant χ2 tests.The eight leave-one-gene-out log-likelihoods average -141.2948; the significance pattern matches the p-values in Table 5.
- Random-gene comparison: Adding two randomly selected unchosen genes raises the average log-likelihood only to -128.3933, an increase of 0.9584, with no significant χ2 test in 20 repetitions.The comparison is against the model containing the eight selected genes.
- Interpretation: These experiments indicate that deleting a selected gene substantially worsens the model, whereas adding two random genes provides little improvement.The analysis therefore supports the importance of the eight selected genes relative to unselected genes in this dataset.
7. Conclusion
The paper develops iterative sure independence screening for survival analysis when dimensionality can greatly exceed sample size. Simulations report sure screening with very small false selection, greater computational efficiency and specificity than the compared LASSO version, and smaller absolute and mean square errors.
- Contribution: The proposed method targets survival analysis with dimensionality that can be much larger than the sample size.Its focus is iterative sure independence screening for high-dimensional Cox-model variable selection.
- Method: Vanilla-ISIS iteratively combines large-scale screening by conditional marginal utility with moderate-scale penalized partial-likelihood selection.The screening filters unimportant variables before penalized selection further narrows the model.
- Results: Simulations demonstrate sure independence screening with very small false selection for the proposed vanilla-ISIS procedure.The conclusion characterizes this as methodological power demonstrated through carefully designed simulation studies.
- Results: Compared with the evaluated LASSO version, the method is more computationally efficient and more specific in selecting important variables.The reported comparison also includes smaller absolute deviation error and mean square error.