Source-linked AI summary
Variable selection in semiparametric regression modeling
Runze Li, Hua Liang
TL;DR
The paper addresses the computationally difficult problem of selecting variables across both parametric and nonparametric components of semiparametric regression models. It proposes nonconcave penalized likelihood and a semiparametric GLRT, establishing oracle-type estimator behavior and a nuisance-free chi-square null distribution for the test.
Problem
Semiparametric variable selection must select nonparametric models and significant parametric variables, while traditional submodel procedures require separate smoothing selection for each submodel.
Method
The paper uses nonconcave penalized likelihood for parametric selection and extends generalized likelihood ratio tests to semiparametric models for nonparametric selection.
Results
The estimator has oracle properties under suitable penalty and regularization choices, while the GLRT has a nuisance-parameter-independent chi-square limiting null distribution.
Takeaways & Limitations
The procedures provide computationally less costly variable selection and allow GLRT critical values from either asymptotic chi-square theory or bootstrap methods.
Takeaways & Limitations
The oracle properties require bandwidth conditions such as nh4 → 0 and nh2/(log n)2 → ∞, together with the stated regularity conditions.
Abstract
from arXiv · showhide
In this paper, we are concerned with how to select significant variables in semiparametric modeling. Variable selection for semiparametric regression models consists of two components: model selection for nonparametric components and selection of significant variables for the parametric portion. Thus, semiparametric variable selection is much more challenging than parametric variable selection (e.g., linear and generalized linear models) because traditional variable selection procedures including stepwise regression and the best subset selection now require separate model selection for the nonparametric components for each submodel. This leads to a very heavy computational burden. In this paper, we propose a class of variable selection procedures for semiparametric regression models using nonconcave penalized likelihood. We establish the rate of convergence of the resulting estimate. With proper choices of penalty functions and regularization parameters, we show the asymptotic normality of the resulting estimate and further demonstrate that the proposed procedures perform as well as an oracle procedure. A semiparametric generalized likelihood ratio test is proposed to select significant variables in the nonparametric component. We investigate the asymptotic behavior of the proposed test and demonstrate that its limiting null distribution follows a chi-square distribution which is independent of the nuisance parameters. Extensive Monte Carlo simulation studies are conducted to examine the finite sample performance of the proposed variable selection procedures.
1. Introduction.
The paper develops variable-selection procedures for generalized varying-coefficient partially linear models, addressing selection in both nonparametric and parametric components. It combines nonconcave penalized likelihood with generalized likelihood ratio testing and studies the resulting estimators and test.
- Motivation: Variable selection in the GVCPLM must identify significant variables in both nonparametric and parametric components.This makes the problem more challenging than selection in standard linear models.
- Motivation: Traditional stepwise and best-subset procedures require smoothing-parameter selection for every submodel, creating substantial computational difficulty.
- Contributions: The proposed procedures use nonconcave penalized likelihood for parametric-component selection and generalized likelihood ratio tests for nonparametric-component selection.
- Contributions: The paper studies convergence rates, asymptotic properties, and oracle performance of the resulting estimator.
- Contributions: The proposed GLRT has a nuisance-parameter-independent chi-square limiting null distribution with diverging degrees of freedom.Critical values can be obtained from the asymptotic chi-square distribution or bootstrap methods.
- Organization: The paper presents Monte Carlo studies and a real-data application after developing the procedures and their asymptotic theory.
2. Selection of significant variables in the parametric component.
The paper estimates parametric coefficients by combining local likelihood estimation of the nonparametric component with penalized likelihood, producing exact zeros for variable selection. Under suitable penalties, tuning parameters, and bandwidth conditions, the estimator has convergence, sparsity, and asymptotic-normality properties while remaining computationally less costly than fully iterated alternatives.
- Penalty and tuning selection: The penalty may be prespecified, with regularization parameters selected separately for coefficients using criteria such as cross-validation or generalized cross-validation.Different penalties and tuning parameters can preserve coefficients regarded as important by leaving them unpenalized.
- Penalty choices: The framework considers L0, L1, Lq, and SCAD penalties, with L0 corresponding to best subset selection but requiring exhaustive subset search and high computational cost in large dimensions.LASSO and bridge penalties avoid the discontinuity of L0, while SCAD is among the nonconcave penalties considered.
- Penalized likelihood: Local likelihood first estimates the nonparametric coefficient functions, after which penalized likelihood estimates the parametric coefficients and can set some coefficients exactly to zero.Exact zero coefficients correspond to excluding the associated variables from the final model.
- Computational properties: The proposed estimator has the asymptotic performance of an oracle estimator and is less computationally costly than fully iterated backfitting or profile likelihood estimation.Ridge regularization may stabilize local likelihood estimation when the Hessian is nearly singular for high-dimensional covariates.
- Sampling properties: The penalized likelihood estimator has rate OP(n^-1/2 + an) under an → 0, bn → 0, nh^4 → 0, and nh^2/log(1/h) → ∞.These conditions are stated for a local maximizer of the penalized likelihood.
- Sampling properties: Under additional regularization conditions, the estimator is root-n consistent, sets the zero-coefficient block exactly to zero, and is asymptotically normal.Theorem 2 also states that undersmoothing is necessary for root-n consistency and asymptotic normality.
3. Statistical inferences for nonparametric components.
The paper estimates nonparametric components by local likelihood and uses a generalized likelihood ratio test to select significant variables. Bandwidth conditions support oracle properties and a chi-square null distribution for the test.
- Estimation: Local likelihood estimation replaces β with its estimate and maximizes over local parameters to obtain α̂(u).
- Estimation: The estimated nonparametric component has the same asymptotic bias but smaller asymptotic covariance than the corresponding estimator without estimated β.
- Bandwidth selection: An optimal local bandwidth can be selected by minimizing conditional mean squared error, while global selectors can use standard univariate nonparametric bandwidth techniques.
- Bandwidth selection: O(n^-1/3) bandwidth is recommended because the usual O(n^-1/5) order does not satisfy the theoretical conditions for the earlier results.The proposed order satisfies nh^4 → 0 and nh^2/(log n)^2 → ∞, supporting the oracle property.
- Variable selection: The GLRT selects nonparametric variables by testing whether selected coefficient functions are jointly zero against the alternative that at least one is nonzero.The procedure can apply a sequence of such tests for subsets of variables.
- Variable selection: Under the null hypothesis, TGLR has an asymptotic χ^2 distribution with degrees of freedom determined by the kernel, bandwidth, number of tested functions, and support length.The result is presented as a new Wilks phenomenon extending generalized likelihood ratio theory to semiparametric modeling.
4. Simulation study and application.
The simulations evaluate variable-selection, estimation, testing, and computational performance in semiparametric models, while the application examines burn-survival data. The proposed procedures show favorable finite-sample behavior and identify significant covariates.
- Extensive Monte Carlo simulations assess the finite-sample performance of the proposed variable-selection procedures.
- The backfitting estimate of α(u) performs as well as estimation using the true β in the Poisson and logistic examples.This comparison uses RASE.
- The proposed GLRT has a finite-sample null distribution close to χ2, with power increasing rapidly as δ increases.For Example 4.1, bootstrap-based Type I error rates at significance levels 0.25, 0.10, 0.05, and 0.01 are 0.2250, 0.0875, 0.05, and 0.0125.
- SCAD performs best in the logistic example and is very close to the oracle procedure.The study compares penalized likelihood procedures with best subset variable selection using GMSE and model complexity.
- The burn-center application uses cross-validation to select bandwidth and tests whether the age-dependent coefficient for prior respiratory disease is zero.The selected bandwidth is 48.4437, and Figure 2 displays the resulting coefficient estimates and confidence intervals.
- The GLRT statistic for prior respiratory disease is 15.7019 with P value 0.015, supporting significance at level 0.05.The test is based on 1,000 bootstrap samples.
- The SCAD-selected application model retains three z-variables: Z3, Z5, and Z7.Their estimates and standard errors are −1.9388(0.4603), −0.0035(0.0054), and −0.0007(0.0006), respectively.
5. Proofs.
The proofs establish uniform asymptotic behavior for local likelihood estimates, then derive asymptotic normality, variable-selection properties, and a chi-square limit for the generalized likelihood ratio test under stated regularity conditions.
- Regularity conditions: The proof relies on concavity, smoothness, moment, kernel, support, and density conditions that ensure uniqueness, uniform convergence, and control of approximation errors.Concavity ensures a unique local-likelihood solution, while some bounded-support assumptions are imposed mainly to simplify the proofs.
- Asymptotic expansions: Under bandwidth conditions h → 0 and nh → ∞, the local likelihood estimator has a uniform expansion with error O_P{h^2 + c_n log^1/2(1/h)}.The expansion holds uniformly over u in the support of U, with c_n = (nh)^−1/2.
- Asymptotic normality: The asymptotic normality of the local estimates follows from the central limit theorem and Slutsky theorem applied to sums of independent and identically distributed random vectors.The penalized likelihood is then used to improve the estimate of β̃.
- Variable-selection properties: The penalized estimator sets the inactive parametric coefficients to zero with probability tending to one, while its active-coefficient analysis proceeds at the n^−1/2 rate.For inactive coefficients, the derivative and coefficient have opposite signs near zero, so the penalized likelihood is maximized at zero.
- Variable-selection properties: The proof establishes asymptotic normality for the active coefficients after showing that the inactive coefficients vanish and that the relevant Hessian converges to its limiting matrix.The Hessian expansion is given by −B_1 + o_P(1).
- Generalized likelihood ratio test: Under the null hypothesis, the generalized likelihood ratio statistic converges to a chi-square distribution with degrees of freedom d_f,n tending to infinity.The proof shows that the auxiliary terms are negligible relative to the leading term before applying the likelihood-ratio result.