Source-linked AI summary

Scale-invariant Optimal Sampling for Rare-events Data with Sparse Models

Jing Wang, HaiYing Wang, Qiang Zhang, Hao Helen Zhang

arXiv:2608.22597v1stat.MLcs.LG

TL;DR

Rare-events subsampling must reduce computation without losing estimation efficiency, but existing optimal probabilities depend on covariate scales and can be distorted by inactive features. This paper develops scale-invariant sampling for sparse models, establishes adaptive-lasso theory, and reports desirable theoretical and practical properties for the resulting procedures.

  • Problem

    Existing optimal subsampling probabilities for rare-events data depend on covariate scales, and inactive features can therefore distort sampling and variable-selection results.

  • Method

    The paper combines adaptive lasso for rare-events variable selection with prediction-error-based scale-invariant optimal sampling and a maximum sampled conditional likelihood estimator.

  • Results

    The practical subsampling estimator has asymptotic variance reaching the lower bound of a large class of subsample estimators.

  • Takeaways & Limitations

    The proposed sampling procedure produces small, balanced subsamples with little information loss and reduced computational burden.

  • Takeaways & Limitations

    The asymptotic prediction-error results are limited to fixed numbers of active variables and may not generalize to dense or over-parameterized models.

Abstract

from arXiv · show

Subsampling is effective in tackling computational challenges for massive data with rare events. Overly aggressive subsampling may adversely affect estimation efficiency, and optimal subsampling is essential to mitigate the information loss. However, existing optimal subsampling probabilities depend on data scales, and some scaling transformations may result in inefficient subsamples. This problem is more significant when there are inactive features, because their influence on the subsampling probabilities can be arbitrarily magnified by inappropriate scaling transformations. We tackle this challenge and introduce a scale-invariant optimal subsampling function in the context of sparse models, where inactive features are commonly assumed. Instead of focusing on estimating model parameters, we define an optimal subsampling function to minimize the prediction error, using adaptive lasso to outline the estimation procedure and study its theoretical guarantee. We first introduce the adaptive lasso estimator for rare-events data and establish its oracle properties, thereby validating the use of subsampling. Then we derive a scale-invariant optimal subsampling function that minimizes the prediction error of the inverse probability weighted (IPW) adaptive lasso. Finally, we present an estimator based on the maximum sampled conditional likelihood (MSCL) to further improve the estimation efficiency. We conduct numerical experiments using both simulated and real-world data sets to demonstrate the performance of the proposed methods.

1 Department of Statistics, University of Connecticut, Storrs, CT 06269, USA

The paper lists an affiliation with Wills Eye Hospital and Thomas Jefferson University in Philadelphia, USA.

  • Wills Eye Hospital is affiliated with Thomas Jefferson University in Philadelphia, Pennsylvania, USA.
  • The listed address is 840 Walnut St, Philadelphia, PA 19107, USA.
  • The affiliation is associated with the Vickie and Jack Farber Vision Research Center.

3 Department of Mathematics, University of Arizona, Tucson, AZ 85721, USA

The paper’s keywords emphasize massive data, oracle properties, prediction error, scaling, and variable selection.

  • The paper concerns massive data and rare-events subsampling challenges.
  • Prediction error and scaling are central analytical themes.
  • Variable selection and oracle property are identified as key topics.

1 Introduction

The paper addresses scale-dependent optimal subsampling for rare-events data, especially when sparse models contain inactive features. It develops scale-invariant sampling with adaptive-lasso-based estimation and theoretical guarantees.

  • Motivation: Rare-events data are highly imbalanced, with zeros potentially hundreds or thousands of times more numerous than ones.
  • Scale dependence: Existing A-OS and L-OS prediction errors can change substantially when covariate scales are transformed.In the example, A-OS can resemble uniform sampling at s = 0.01, while L-OS can do so at s = 100.
  • Scale dependence: Scale dependence can distort variable selection because inactive variables may be arbitrarily rescaled without changing the underlying model.
  • Contributions: It establishes oracle properties for adaptive lasso in rare-events data and proposes a unified estimator combining variable selection with optimal sampling.
  • Contributions: The paper proposes scale-invariant optimal subsampling to support parameter estimation and variable selection.
  • Contributions: The practical subsampling algorithm reduces computational burden, while the estimator’s asymptotic variance reaches the lower bound for a broad class of subsample estimators.

2 Background and model setup

The model setup treats rare-events binary responses in which the expected number of ones is much smaller than the expected number of zeros. The background motivates retaining ones and sampling zeros to control computational cost and variance inflation.

  • Model setup: The data consist of N observations (x_i, y_i) with binary responses and covariates from their joint distribution.
  • Rare-events assumption: Rare-events asymptotics assume the intercept tends to −∞, implying that the expected number of ones is negligible relative to the expected number of zeros.
  • Efficiency: The full-data maximum likelihood estimator has asymptotic variance of order 1/N1 rather than 1/N, so estimation efficiency is determined by the number of rare ones.
  • Subsampling: Keeping all ones and sampling zeros can reduce computational cost, although aggressive subsampling may inflate variance.
  • Scale-dependent methods: Earlier A- and L-optimal sampling functions were designed to reduce variance inflation but depend on covariate scale.
  • Research gap: Because scale changes can affect optimal probabilities and variable-selection results, the paper aims to develop scale-invariant alternatives and address variable selection for rare-events data.

3 Nonuniform sampling with variable selection for rare-events data

The paper develops subsampling for rare-events variable selection with adaptive lasso, addressing scale dependence that can make inactive variables distort sampling probabilities. It establishes theoretical guarantees and proposes prediction-oriented, scale-invariant probabilities based on active features.

  • Adaptive lasso for rare-events data: Adaptive lasso is introduced for rare-events data, with oracle properties established for variable selection and estimation.The estimator supports selecting active variables despite scarce positive observations.
  • Subsampling algorithm: Poisson subsampling retains all ones and samples zeros with inclusion probability π(x_i,y_i)=y_i+(1−y_i)ρφ(x_i).The algorithm records inverse-probability weights for sampled zeros before fitting adaptive lasso.
  • Theoretical guarantees: Theorem 1 shows asymptotic estimation efficiency is predominantly determined by the number of ones rather than the full data size.The variance also includes a subsampling inflation term, and sufficient zero sampling can avoid asymptotic information loss.
  • Limitations of scale-dependent functions: Existing A- and L-optimal sampling functions are scale-dependent and can be influenced by inactive variables when the active set is estimated from pilots.Rescaling inactive covariates can inflate or suppress their contributions, potentially misleading subsampling and variable selection.
  • Scale-invariant optimal function: The P-optimality criterion minimizes prediction error and yields sampling probabilities invariant to covariate rescaling for a class of functions g.Its construction uses the probability term, which remains unchanged under rescaling in logistic regression, and improves active-variable contributions in the numerical example.

4 Penalized MSCL estimator

The penalized MSCL estimator combines sampled conditional likelihood with adaptive lasso to support sparse variable selection while improving estimation efficiency over the IPW estimator. The paper establishes oracle properties, variance dominance, and a practical two-step implementation with reduced computational complexity.

  • 4 Penalized MSCL estimator: The penalized MSCL estimator extends sampled conditional likelihood with an adaptive lasso penalty to ensure model sparsity.The MSCL objective is based on the conditional likelihood of the selected subsample, with penalization added for variable selection.
  • 4 Penalized MSCL estimator: Under stated regularity and tuning conditions, the penalized MSCL estimator has consistent variable selection and asymptotic normality for active parameters.The fixed-dimensional case is covered by a corresponding corollary under a consistent pilot estimator and additional tuning-rate conditions.
  • 4 Penalized MSCL estimator: The MSCL estimator has asymptotic variance no larger than the IPW estimator, with equality when c = 0.When c = 0, the subsample contains sufficiently many zeros and the two estimators can be equally efficient; otherwise MSCL is recommended for improved efficiency.
  • 4 Penalized MSCL estimator: The MSCL estimator achieves the lower bound of asymptotic variances among a class of asymptotically unbiased subsample estimators that includes the IPW estimator.The variance Vmscl(A) is interpreted as the Cramér–Rao bound for this estimator class.
  • 4.1 Practical algorithm and computational complexity: The practical method uses a pilot stage to estimate active variables and sampling probabilities, followed by a second-stage subsample and penalized MSCL optimization.The algorithm screens with a pilot sample, computes approximate optimal probabilities, then applies the estimated probabilities to obtain the final subsample.
  • 4.1 Practical algorithm and computational complexity: With a subsample size of the same order as N1, the two-step algorithm has complexity O(ζout×inNeαtq), reported as significantly faster than the full-data estimator.The reduction comes from using smaller sample size and dimension in the optimization, despite the two-step procedure.

5 Numerical experiments

Numerical experiments on simulated and real data show that optimal subsampling generally improves estimation, prediction, variable selection, and computational efficiency over uniform or full-data baselines. The scale-invariant P-OS method is generally the strongest or most robust optimal-subsampling approach.

  • Estimation and prediction efficiency: All optimal sampling estimators outperform uniform sampling for parameter estimation, and eventually outperform full-data lasso as sampling rates increase.These comparisons use empirical median squared error across three simulated parameter settings.
  • Estimation and prediction efficiency: Optimal sampling estimators achieve lower empirical median squared prediction error than uniform sampling, with P-OS performing best among the three optimal methods.The experiments also show that sampling estimators can eventually outperform full-data lasso despite using fewer observations.
  • Variable selection and computational complexity: Second-stage screening selects numbers close to the true active-variable counts for all subsampling methods, whereas uniform sampling selects fewer active variables at low sampling rates.Uniform sampling has higher rates of excluding active variables than optimal subsampling procedures.
  • Computational time: Subsampling algorithms substantially reduce computational time relative to full-data estimators; optimal sampling uses about 0.77% of full-data adaptive-lasso time.The reduction includes the cost of calculating sampling probabilities.
  • Comparison of different pilot estimators: The lasso is a robust general-purpose pilot estimator, while SIS can outperform it only when its selected-variable count is appropriately specified.An inappropriate SIS selection size may lead to significantly worse performance.
  • Real-data experiments: On real data, nonuniform sampling generally outperforms uniform sampling, and P-OS is significantly better than the other optimal-sampling estimators on the covtype dataset.A font-data exception occurs where A-OS is worse than uniform sampling at a high sampling rate.
  • Real-data experiments: Scale-dependent L-OS performs worse than uniform subsampling in the IRIS analysis, whereas scale-invariant P-OS is more robust and is never the worst method.The authors note that scale-dependent probabilities can either increase or decrease efficiency unpredictably before subsampling.
  • Real-data experiments: For the IRIS Registry analysis, variable-selection patterns are similar across sampling probabilities, with gender consistently selected and region and smoking-status indicators selected as possible risk-associated factors.The study analyzes thyroid eye disease using registry data from patients aged 18–90 years.

6 Conclusion and limitations

The paper concludes that scale-invariant optimal subsampling provides desirable theoretical and numerical properties for rare-events variable selection while reducing information loss and computation. It recommends lasso pilots and P-OS importance functions, but identifies limitations concerning the optimized criterion, asymptotic scope, and model specification.

  • Conclusion: The proposed scale-invariant probabilities produce small, balanced subsamples with little information loss and facilitate computation in practice.The practical procedure requires selecting a pilot estimator, sampling rate ρ, and importance function φ(x).
  • Practical recommendations: The authors recommend lasso pilot estimators and P-OS importance functions, with sampling rate ρ larger than but of the same order as N1/N0.P-OS is recommended because it usually performs well and is robust without prior knowledge of the true model.
  • Limitations: The optimized criterion targets asymptotic mean squared error for rare-event probability estimation rather than variable-selection quality.The authors call for future probabilities optimized directly for variable-selection metrics.
  • Limitations: The theoretical results may not generalize to dense or over-parameterized models, and prediction-error results are limited to a fixed number of active variables.These boundaries arise because the analysis relies on asymptotic normality and fixed active-variable counts.
  • Limitations: The analysis assumes a correctly specified sparse full model and does not account for model misspecification or feature counts vastly exceeding observations.The authors identify these settings as directions for further research.

A Details of mathematical proofs

The appendix supplies proofs for the adaptive-lasso estimator's consistency, variable-selection consistency, asymptotic normality, and optimal sampling-function results under stated assumptions. It uses Taylor expansions, concentration inequalities, KKT conditions, and optimization arguments.

  • General assumptions and setup: The proofs begin by stating general notation and assumptions for fixed and diverging dimensions and active-variable counts.The appendix also defines asymptotic comparison notation and assumes bounded cN without loss of generality.
  • Asymptotic normality: The asymptotic-normality proof derives the limiting IPW target function and analyzes adaptive-lasso penalties separately for active and inactive variables.Taylor expansion and the stated convergence results yield the limiting quadratic representation.
  • Fixed-dimensional asymptotics: For fixed dimensions, the proof decomposes the weighted score, controls remainder terms, and applies KKT conditions to establish adaptive-lasso variable-selection consistency.The inactive-coordinate penalty dominates the score, while active coordinates converge to their true values.
  • Diverging-dimensional asymptotics: For diverging dimensions, the proof controls score terms with Bernstein-type bounds and verifies the weighted Hessian and KKT conditions under sparsity assumptions.The argument establishes the required bounds uniformly over inactive coordinates and supports uniqueness of the estimator.
  • Optimal-function derivations: A general optimization lemma derives minimizers of expectation criteria under normalized importance functions using the Cauchy–Schwarz inequality.The resulting optimizer is proportional to h(x) and normalized by E{h(x)}.
  • Optimal-function derivations: The appendix applies the general lemma to derive optimal functions for trace-based variance and information criteria.Separate calculations target tr(Vw(A)) and tr(Mw(A)).

A.4 Proof of Theorem 2

The proof establishes scale invariance of the leveraging term and develops the two-step adaptive-lasso subsampling procedure, including its theoretical properties and implementation conditions.

  • Scale invariance: Theorem 2 proves that the leveraging term remains invariant under covariate rescaling when parameters are reparameterized accordingly.The transformed covariance-related matrix and gradient preserve the relevant quadratic form.
  • Scale invariance: The probability term is unchanged by scaling inactive variables because the model value g(x; θt) remains unchanged.This completes the scale-invariance argument for the proposed sampling function.
  • Two-step algorithm: The algorithm uses a pilot sample for screening, estimates optimal sampling probabilities, draws a second subsample, and fits an adaptive lasso.The first-stage screening can use lasso or other estimation-and-selection methods, while the second stage operates on the selected variables.
  • Implementation conditions: The procedure can handle p > N when first-stage screening reduces the selected-variable dimension below N.Sure independence screening is given as another possible first-stage method.

B.2 Computational complexity

The computational analysis compares full-data adaptive lasso with the proposed two-step procedure and shows how sparse screening reduces the cost of optimal-probability approximation and estimation.

  • Baseline complexity: Full-data lasso and adaptive lasso have complexity O(ζout×inNp) and require computation over all N observations.The adaptive-lasso complexity uses the corresponding outer and inner iteration counts.
  • Subsample estimation: The average subsample size under sampling rate ρ is O{N(e^αt + ρ)}, reducing the data processed during subsequent estimation.The stated order applies to the expected subsample size.
  • Overall procedure: Using optimal probabilities, the two-step algorithm combines pilot estimation, probability approximation, and subsample fitting into the complexity expressions stated for (5), (8), and (6).The resulting cost depends on pilot iterations, q, Npl, and the expected subsample size.

C.1 Simulation details for Section 1

The simulations use logistic-regression data to test scale-dependent optimal subsampling under sparse and non-sparse parameter settings, with multiple covariate rescalings and baseline estimators.

  • Simulation design: The simulations generate logistic-regression covariates from a six-dimensional lognormal distribution with covariance entries Σij = 0.5|i−j|.The full simulated data contain N = 500000 observations.
  • Parameter settings: The non-sparse setting uses βt = (−1, −1, −0.01, −0.01, −0.01, −0.01), while the sparse setting uses βt = (−1, 0, 0, 0, 0, 0) and αt = −5.These settings contrast weakly nonzero effects with one active variable.
  • Scale transformation: To test scale dependence, the sixth covariate is multiplied by s ∈ {0.01, 0.1, 1, 10, 100} while its coefficient is divided by s.This preserves xTβt and therefore leaves the logistic-regression model unchanged.
  • Baselines and estimation: The evaluation compares uniform sampling, full-data lasso, and full-data adaptive lasso, using lasso for first-stage screening and γ = 1 for adaptive-lasso weights.With γ = 1, the weights are ˆwj = 1/|ˆβpl(j)|.

D.1 Results on variable selection omitted in Section 5

The variable-selection results compare optimal subsampling with uniform sampling and standardization, while related figures assess estimation and prediction errors across sampling settings.

  • Reported metrics: The tables report mean selected-variable counts, active-variable exclusion rates, and true-model selection rates for the compared methods.Table 5 covers selected-variable counts, Tables 6 and 9 cover false negative rates, and Tables 7 and 8 cover true-model selection.
  • Variable selection: Uniform sampling has higher rates of excluding active variables than optimal subsampling procedures, although it may select the true model more often in some cases.Because uniform sampling more often excludes important variables, the authors state that optimal sampling may be preferable in practice.
  • Estimation and prediction error: The study compares eMSE and eMSPE for P-OS without standardization and sP-OS with data standardization.Figure 9 reports eMSE and Figure 10 reports eMSPE, using the same pilot sample size Npl = 500.
  • Comparison with standardization: P-OS and sP-OS have similar eMSE and eMSPE performance, but standardization may reduce the rate of selecting true models.The comparison uses the same pilot estimation methods for fairness.
Loading 2608.22597v1…