Source-linked AI summary
Sure independence screening in generalized linear models with NP-dimensionality
Jianqing Fan, Rui Song
TL;DR
Ultrahigh-dimensional variable selection requires screening methods that remain useful when the number of covariates greatly exceeds the sample size and existing SIS theory is restrictive. The paper ranks marginal likelihood-based estimates or likelihoods in generalized linear models, establishes sure screening with vanishing false selection, and derives an exponential QMLE inequality; tuning-parameter choice remains beyond its scope.
Problem
Existing SIS theory motivates extending independent screening to generalized linear models and less restrictive assumptions, while marginal fitting can miss jointly important variables.
Method
The paper fits p componentwise marginal regressions and ranks maximum marginal likelihood estimates or maximum marginal likelihood values for screening.
Results
The proposed rankings have sure screening properties with vanishing false selection rates, quantify dimensionality reduction, and include an exponential tail bound for the QMLE.
Takeaways & Limitations
Independent screening provides a computationally feasible way to reduce NP-dimensional generalized-linear-model problems to a smaller variable set.
Takeaways & Limitations
Choosing the screening threshold γn is an important practical problem left beyond the scope of the paper.
Abstract
from arXiv · showhide
Ultrahigh-dimensional variable selection plays an increasingly important role in contemporary scientific discoveries and statistical research. Among others, Fan and Lv [J. R. Stat. Soc. Ser. B Stat. Methodol. 70 (2008) 849-911] propose an independent screening framework by ranking the marginal correlations. They showed that the correlation ranking procedure possesses a sure independence screening property within the context of the linear model with Gaussian covariates and responses. In this paper, we propose a more general version of the independent learning with ranking the maximum marginal likelihood estimates or the maximum marginal likelihood itself in generalized linear models. We show that the proposed methods, with Fan and Lv [J. R. Stat. Soc. Ser. B Stat. Methodol. 70 (2008) 849-911] as a very special case, also possess the sure screening property with vanishing false selection rate. The conditions under which the independence learning possesses a sure screening is surprisingly simple. This justifies the applicability of such a simple method in a wide spectrum. We quantify explicitly the extent to which the dimensionality can be reduced by independence screening, which depends on the interactions of the covariance matrix of covariates and true parameters. Simulation studies are used to illustrate the utility of the proposed approaches. In addition, we establish an exponential inequality for the quasi-maximum likelihood estimator which is useful for high-dimensional statistical learning.
1. Introduction.
The paper extends independent screening from Gaussian linear models to generalized linear models by ranking marginal likelihood-based quantities. It establishes sure screening, vanishing false selection, dimensionality reduction, and an exponential QMLE bound for high-dimensional learning.
- Ultrahigh-dimensional regression arises in applications where p can grow much faster than n, making many models unidentifiable.NP-dimensionality is defined by log p = O(n^a) for some a > 0.
- Fan and Lv’s correlation-ranking SIS method motivated research on screening procedures for more general models and less restrictive assumptions.
- The paper proposes ranking maximum marginal likelihood estimates or maximum marginal likelihoods for generalized linear models.The framework includes Fan and Lv’s procedure as a special case and addresses a technical challenge in likelihood ranking by exploiting ranking invariance under monotone transforms.
- Independent screening is fast but crude: marginal rankings can miss jointly important variables or overrank jointly unimportant ones.Iterative and multistage procedures are discussed as extensions, while the simulations focus on vanilla SIS.
- The proposed methods possess sure screening properties with vanishing false selection rates under surprisingly simple conditions.
2. Generalized linear models.
The paper formulates generalized linear models using an exponential-family response and a covariate vector with an intercept. It allows the number of covariates to increase with sample size and treats ordinary Gaussian linear regression as a special case.
- The response follows an exponential-family model with canonical parameter, mean b′(θ), and optional dispersion excluded from the presentation.The paper focuses on canonical links for simplicity.
- The generalized linear model estimates a (p + 1)-vector of coefficients from covariates including the intercept x0 = 1.
- The data are assumed to be i.i.d., with p allowed to grow with n.
- Ordinary linear regression with standard normal errors is a special case, and standardized designs make MMLE and marginal-correlation rankings equivalent.
3. Independence screening with MMLE.
MMLE-based independence screening fits separate componentwise marginal regressions and ranks variables by marginal coefficient magnitude. Thresholding these estimates reduces the parameter dimension while retaining joint nonsparsity under a mild condition.
- For each covariate, the MMLE is obtained by fitting a componentwise marginal regression with the response, that covariate, and an intercept.The empirical objective is based on the marginal likelihood and can be rapidly computed.
- The screening rule selects variables whose estimated marginal coefficients exceed a predefined threshold γn.
- Ranking by marginal coefficient magnitude can reduce pn variables, potentially hundreds of thousands, to a much smaller computationally feasible set.
- Although marginal-model interpretations are biased relative to the joint model, nonsparse information can pass to marginal models under a mild condition.
- Under suitable thresholding and conditions, the selected set asymptotically contains the true model with probability one.
4. An exponential bound for QMLE.
The paper derives an exponential tail bound for the quasi-maximum likelihood estimator under general conditions, providing a technical tool for its sure-screening analysis. The framework uses i.i.d. data, compact parameter spaces, positive-definite information, and Lipschitz likelihoods.
- An exponential bound for the QMLE is established for use in subsequent high-dimensional screening theory.The result is also presented as independently useful because it holds under general conditions.
- The setup allows covariates with discrete or continuous components and dimensionality that depends on n.
- The population parameter lies in the interior of a sufficiently large compact convex parameter set.
- The assumptions require finite positive-definite Fisher information and a Lipschitz property for the likelihood.These conditions support identifiability, QMLE existence, and the concentration argument.
- The proof controls the tail of n∥β̂ − β0∥ through empirical-process arguments together with convexity and Lipschitz continuity.
5. Sure screening properties with MMLE.
This section characterizes when marginal coefficients preserve the nonsparsity of jointly important variables, establishing sure screening for MMLE-based independence learning in generalized linear models. It also gives uniform-convergence guarantees and bounds the selected-model size in terms of covariate dependence.
- Population aspect: A jointly important variable remains marginally important when its covariance with the mean response function is nonzero under the stated screening framework.The marginal regression parameter measures the covariance between the marginal covariate and the mean response function.
- Population aspect: Partial orthogonality makes jointly unimportant variables have zero marginal coefficients, enabling model selection consistency at a suitable threshold.The condition requires variables outside the active set to be independent of variables inside it.
- Uniform convergence and sure screening: The MMLE screening property does not directly depend on the covariance matrix operator norm, unlike screening based on the full likelihood.The result is therefore compatible with NP-dimensional settings under the theorem’s assumptions.
- Uniform convergence and sure screening: For normal covariates, the theorem is weaker than Fan and Lv’s dimensionality range but allows nonnormal covariates and other error distributions.The stated comparison is log p_n = o(n^{1−2κ}) for Fan and Lv versus log p_n = o(n^{(1−2κ)/4}) here.
- Uniform convergence and sure screening: The selected-set size is O(n^{2κ}λmax(Σ)), so covariate correlation controls dimensionality reduction and can make false discoveries negligible.More generally, the size is of order ||Σβ⋆||²/γ_n²; when n^{2κ}λmax(Σ)/p → 0, the selected fraction is negligible.
6. A likelihood ratio screening.
This section introduces screening by ranking marginal likelihood contributions and compares it with MMLE screening. Both methods have sure screening properties under related conditions, while their selected-model sizes are of the same order.
- Likelihood ratio screening: Likelihood-ratio screening and MMLE screening are equivalent for sure screening and produce selected sets of the same order of magnitude.The likelihood-ratio method incorporates both estimator magnitude and associated variation, whereas MMLE screening uses estimator magnitudes.
- Likelihood ratio screening: Marginal likelihood screening ranks features by their marginal contributions to likelihood increments after solving p_n two-parameter optimization problems.The procedure sorts the marginal likelihood values in descending order and retains variables exceeding a predefined threshold ν_n.
- Likelihood ratio screening: The marginal likelihood increment also measures the correlation between a marginal covariate and the mean response function.This provides the population-level foundation for establishing its sure screening property.
- Likelihood ratio screening: Under the stated regularity conditions, likelihood-ratio screening admits a sure screening theorem and a corresponding bound on the number of selected variables.The paper separately develops results for minimum and total signal sizes, with the stochastic-noise challenge handled through ranking invariance under strict monotone transforms.
7. Numerical results.
The simulations evaluate SIS based on marginal MLE and marginal likelihood rankings across generalized linear-model settings with varying correlation, sparsity, sample size, and dimensionality. SIS generally achieves sure screening, while stronger correlation and larger nonsparse sets increase screening difficulty and can expose weaknesses in competing methods.
- Simulation design: Three simulation settings vary correlation structures, sample sizes from 80 to 600, and dimensionalities p = 2000, 5000, and 40,000.Settings S1, S2, and S3 are used to study SIS behavior under different covariance structures.
- Covariance behavior: Empirical maximum covariance eigenvalues increase with correlation, q, and p, and with decreasing sample size.The empirical minimum eigenvalue is zero and sample covariance condition numbers are infinite because p > n.
- Evaluation: The marginal MLE and marginal likelihood-ratio procedures screen variables, with minimum model size measuring the number needed to contain the true model.LASSO and SCAD provide reference comparisons for smaller p because of computational burden.
- Signal difficulty: Minimum |t|-statistics in oracle models become smaller as within-covariate correlation increases, sometimes reaching three decimals.This indicates that recovering all significant variables can remain difficult even when the true model size is known.
- Screening performance: As correlation or nonsparse-set size increases, MMMS and RSD usually increase for SIS, LASSO, and SCAD.The reported results are based on 200 simulations per scenario.
- Screening performance: SIS performs well across the designed scenarios, while LASSO and SCAD occasionally fail under high correlations or larger nonsparse sets.In the third setting, LASSO and SCAD usually fail to select important variables, whereas SIS remains reasonably effective.
8. Concluding remarks.
The paper establishes sure screening for two marginal-likelihood-based methods in generalized linear models, while identifying scope boundaries and extensions for future work.
- The proposed methods rank maximum marginal likelihood estimators or maximum marginal likelihood in generalized linear models.
- Sure independence screening and vanishing false selection rate are established for the proposed methods, with Fan and Lv (2008) as a special case.
- The current theoretical results require concavity of the log-likelihood in regression parameters, excluding some noncanonical-link generalized linear models.
- The authors identify broader future extensions to regression models beyond generalized linear models and to grouped-variable screening.
- Choosing the tuning parameter γ_n remains an important practical problem; suggested approaches include retaining n or n/log(n) features initially and using later-stage criteria.
9. Proofs.
The proofs combine empirical-process concentration tools with convexity, Taylor expansion, and covariance arguments to establish estimator bounds and screening results.
- The proof of the quasi-maximum likelihood estimator bound uses symmetrization, contraction, and concentration theorems.
- The estimator proof first bounds ∥ˆβ − β0∥ on a suitable neighborhood, then applies a lemma to conclude the result.
- Covariance arguments show that zero marginal utility implies a zero marginal regression coefficient under the stated monotonicity conditions.
- The proofs bound the number of variables with marginal coefficients exceeding εn^-κ by O{n^2κλmax(Σ)}.
- The screening theorem follows by transferring lower bounds from population marginal signals to estimated signals and applying union bounds.