Source-linked AI summary
ARISE: An adaptive residual-informed stability ensemble for feature selection in small-sample biomedical omics
Zardad Khan, Amjad Ali, Naz Gul, Sheema Gul, Saeed Aldahmani
TL;DR
Small-sample molecular classification needs feature selectors that balance predictive relevance, stability, redundancy control, and multiclass discrimination. ARISE integrates these signals through adaptive nested selection and ranked first in all 15 dataset–metric cells across five datasets, classifiers, and metrics.
Problem
Small-sample molecular classification needs feature selectors that are predictive, stable, nonredundant, and sensitive to minority classes.
Method
ARISE combines complementary relevance evidence, class-balanced robustness, residual-based redundancy control, multiclass pair coverage, and nested profile selection.
Results
ARISE ranked first in all 15 dataset–metric cells across five datasets, three fixed classifiers, and three reported metrics.
Takeaways & Limitations
ARISE provides a transparent framework for small-sample molecular classification, while dataset-specific feature-budget optima support evaluating multiple feature-set sizes.
Takeaways & Limitations
The benchmark used fixed classifiers and dataset-level resampling, motivating future evaluation of cohort transfer and feature-selection stability for biomarker interpretation.
Abstract
from arXiv · showhide
Objective: Small-sample molecular classification requires feature selectors that identify predictive, stable, and nonredundant subsets for binary and multiclass outcomes. We propose ARISE (Adaptive Residual-Informed Stability Ensemble), which integrates complementary relevance signals, class-balanced stability assessment, residual-informed redundancy control, and multiclass pairwise coverage. Methods: ARISE combines seven percentile-normalized relevance components through 15 predefined profiles, adaptively weighted by nested inner cross-validation. It was evaluated on five molecular datasets, eight feature-set sizes, three fixed classifiers (k-nearest neighbours, support vector machine, and random forest), and six filter comparators. Generalization was estimated by five-fold outer cross-validation repeated 50 times using balanced accuracy, macro-F1, and Cohen's kappa. Results: Across 210,000 held-out assessments, ARISE ranked first in all 15 dataset-metric combinations. Equal-dataset means were 0.793 for balanced accuracy, 0.776 for macro-F1, and 0.725 for kappa, exceeding the strongest aggregate comparator by 0.022, 0.023, and 0.028, respectively. Performance remained strong across compact feature sets, although the optimal budget differed by dataset. Conclusion: ARISE provides a transparent, adaptive framework that jointly addresses relevance, stability, redundancy, and multiclass discrimination. Its consistent results across datasets, classifiers, metrics, and feature-set sizes support further evaluation for small-sample molecular classification.
1. Introduction
Small-sample molecular studies require feature selectors that remain stable, nonredundant, and interpretable despite high dimensionality and varied signal geometries. ARISE addresses these needs with a multiclass, stability-aware ensemble combining complementary relevance, robustness, residual redundancy control, and pairwise coverage while evaluating feature selection independently of classifier optimization.
- Motivation: In large p, small n molecular studies, selectors must resist sampling perturbations and avoid redundant or leakage-contaminated panels that may fail interpretation, transfer, or validation.Candidate feature counts can reach tens of thousands, whereas cohorts may contain only tens or hundreds of observations.
- Motivation: Conventional filters emphasize different signal geometries, and no single geometry is uniformly appropriate across tissues, diseases, platforms, class imbalance, and feature budgets.Mean-shift, mutual information, interval-overlap, and Bhattacharyya criteria capture distinct distributional or dependence structures.
- ARISE framework: ARISE unifies complementary percentile-scaled filters, class-balanced subsample robustness, within-class residual-correlation penalties, unresolved-class-pair coverage, and adaptive profile tuning.The selector uses a fixed profile library with nested minimax-regret racing and soft consensus.
- Contributions: The manuscript contributes a unified binary and multiclass construction with exact nested feature paths and theoretical properties including boundedness, pair-coverage submodularity, and class-shift-invariant redundancy control.It also combines relevance evidence with robustness, residual redundancy control, and explicit multiclass pair coverage.
- Scope: The study is a methodological evaluation rather than evidence of clinical deployment, using diverse molecular datasets and fixed classifiers to isolate the effect of feature selection.The classifiers are deliberately held fixed, and learner hyperparameters remain unchanged across selectors.
2. Related work
Related work spans classical filter, wrapper, and embedded feature-selection methods, alongside stability assessment and nested model-selection practices. Recent approaches increasingly combine accuracy, stability, compactness, interpretability, and diverse biomedical data representations through multi-objective, ensemble, graph, single-cell, and multi-omics methods.
- Classical feature selection: Feature selection is commonly organized into filters, wrappers, and embedded approaches, with representative methods including maximum-relevance/minimum-redundancy, support-vector recursive elimination, lasso, and elastic net.These methods can be brittle when their assumptions about effect geometry or correlation do not transfer across heterogeneous settings.
- Stability and validation: Chance-corrected stability and stability selection made resampling consistency an explicit feature-selection objective.ARISE uses subsampling stability as a score component without claiming classical stability selection’s family-wise error control.
- Stability and validation: Nested model selection and preprocessing are necessary because nonnested tuning creates optimistic bias, a concern recognized in microarray prognosis studies.This validation principle is part of the methodological context motivating ARISE.
- Multi-objective and ensemble selection: Recent multi-objective and ensemble methods target complementary combinations of accuracy, stability, compactness, and interpretability.Examples include bias correction, filter–wrapper stacking, rough-hypervolume search, information criteria, bi-level ensembles, rank-revealing subspaces, and expert-augmented selection.
- Biomedical feature discovery: Biomedical feature discovery increasingly covers network, single-cell, multi-view, and multi-omics learning, including regulatory graphs, trajectory-preserving selection, information imbalance, electronic-health-record-assisted omics, incomplete integration, and cross-platform prediction.Foundation models, graph methods, and multiscale association frameworks further broaden representational spaces.
3. Materials and methods … 3.3. Profile relevance, stability, redundancy, and pair coverage
ARISE was evaluated across five molecular datasets using fixed classifiers and metrics, with all pipeline quantities learned within training partitions. Its method combines 15 weighted relevance profiles with balanced subsampling, separate stability assessment, residual-informed redundancy control, and multiclass pair coverage.
- 3.1. Study datasets and data provenance: Each dataset used the same kNN, SVM, and random forest classifiers and reported balanced accuracy, macro-F1, and Cohen’s κ.Across-study aggregation treated the dataset as the unit, while folds, feature budgets, classifiers, and predictions were repeated measurements within datasets.
- 3.2. Notation and component construction: ARISE defines nmin from class sizes, samples 70% of the smallest class up to a minimum rule across B = 8 draws, and recomputes seven relevance components.The components are POS, Fisher, SNR, Welch, Bhattacharyya, rank-based, and MI, followed by re-percentiling across screened features.
- 3.2. Notation and component construction: All preprocessing, component scores, subsampling quantities, profile tuning, feature paths, and classifier parameters are learned from the applicable training partition.Transformations embedded in published source objects are treated as fixed inputs rather than reestimated.
- 3.2. Notation and component construction: Seven relevance components are percentile-normalized within class pairs, averaged equally over pairs, and combined through C = 15 fixed candidate profiles with nonnegative weights.The pairwise AUC contributes only to feature-relevance construction and is not an evaluation metric.
- 3.3. Profile relevance, stability, redundancy, and pair coverage: The combined base score mixes profile-weighted relevance and subsample-averaged robustness, while selected-set Jaccard stability is evaluated separately.The robustness quantity is a mean subsample rank, not a dispersion measure or selection frequency.
- 3.3. Profile relevance, stability, redundancy, and pair coverage: Redundancy is computed after within-class centering, and multiclass coverage uses profile-weighted pairwise utilities with renormalized pairwise-component weights.For binary outcomes or pure-MI profiles, pair coverage is disabled by setting Fc ≡0 and ηc = 0.
- 3.3. Profile relevance, stability, redundancy, and pair coverage: At each path step, ARISE selects from a deterministic local candidate universe using profile-specific pair-coverage and redundancy parameters, producing one ordered path of exact-prefix feature budgets.No hard correlation deletion or mandatory cluster representative is used.
3.4. Two-stage nested profile racing and soft consensus … 3.8. Computational details
ARISE uses nested two-stage profile racing and soft consensus to construct feature paths, then evaluates them with fixed classifiers under repeated nested validation. The framework also defines exploratory across-dataset diagnostics and reproducibility-oriented computational safeguards.
- 3.4. Two-stage nested profile racing and soft consensus: 15 profiles undergo two-stage inner-cross-validation racing, with all profiles screened at budgets 5 and 15 before exactly three advance to expanded evaluation.The profiles include eight mixed designs and seven pure components; selection uses a padded or capped one-standard-error rule.
- 3.4. Two-stage nested profile racing and soft consensus: Soft profile weights are converted into a consensus path after Stage 2, and its prefixes provide selected feature sets for every budget in Q = {1, 5, 10, 15, 20, 25, 30, 35}.Component scores, stability, redundancy, pair coverage, and the greedy path are recomputed on the full outer-training partition.
- 3.5. Comparators and fixed classifiers: The comparator library contains POS, Fisher score, SNR, MI, Bhattacharyya distance, and Welch score, with limited grids tuned using the same inner objective.POS whisker values were {1, 1.5, 2}, and MI bin counts were {2, 3, 4}.
- 3.5. Comparators and fixed classifiers: The fixed classifier set comprises distance-weighted kNN, radial SVM, and random forest, using identical selector-independent hyperparameters and class weighting within each training partition.The SVM uses C = 1 and γ = 1/q; the random forest uses 120 trees and mtry = ⌊√q⌋.
- 3.6. Nested validation, metrics, and aggregation: Validation uses five stratified outer folds repeated fifty times, with one three-fold inner split and training-only fitting of preprocessing, selection, scoring, and models.This yields 250 held-out fold values per budget–classifier condition.
- 3.6. Nested validation, metrics, and aggregation: Balanced accuracy, macro-F1, and Cohen’s kappa jointly assess equal class recall, classwise precision–recall balance, and chance-corrected agreement.Across-dataset means equally weight the five datasets after averaging 250 outer-fold values and then 24 budget–classifier condition means.
- 3.7. Exploratory statistical analysis: Across-dataset inference uses Friedman rank diagnostics and exact two-sided sign tests, with six comparator P values Holm-adjusted within each metric.Because only five heterogeneous datasets were available, the tests are treated as exploratory diagnostics rather than definitive inference.
- 3.8. Computational details: Implementation caches component and pairwise summaries, reuses nested prefixes, parallelizes outer folds, and applies deterministic caps, seeds, split assignments, tie handling, and completeness checks.Profile screening caps are 525, 630, or 700 for mixed profiles and 700 for pure profiles; the shared workspace contains at most 700 features.
4. Theoretical properties
Theoretical results establish bounded scores, monotone submodular multiclass pair coverage, invariance of residualized redundancy measures to within-class shifts, and deterministic nested feature paths. They also formalize conditional concentration, procedural outer-test separation, and computational costs while clarifying limitations of these guarantees.
- Bounded construction: All relevance, stability, aggregate, pair-coverage, and marginal-gain quantities remain in [0, 1] under the stated bounded-input and simplex-weight conditions.The marginal gain is nonnegative because adding a feature cannot reduce coordinatewise maxima.
- Pair coverage: The pair-coverage objective Fc is normalized, monotone, and submodular, including the enabled multiclass implementation and trivially when coverage is disabled.Diminishing returns follows because current pairwise maxima can only increase, and nonnegative averaging preserves these properties.
- Shift invariance: Within-class additive shifts leave the residualized vector and every redundancy correlation unchanged when the within-class scale is nonzero.The transformed class mean, within-class scale, and residual cross-products remain identical.
- Deterministic nested path: Conditional on training data, configuration, seed, candidate-universe order, and enough finite-score features, greedy selection returns a unique no-duplicate path whose smaller paths are exact prefixes of larger paths.Fixed score and feature-index tie-breaking ensures uniqueness and the prefix property.
- Concentration: The concentration result is a formal Monte Carlo statement for conditional subsample means, not biological selection consistency; with B = 8 and up to 700 screened features, its simultaneous bound can be numerically vacuous.It follows from Hoeffding’s inequality and a union bound under conditionally independent balanced subsamples.
- Outer-test separation: Procedural outer-test separation holds because preprocessing, selection, tuning, feature paths, and classifier fits are frozen from outer-training data before test prediction, but embedded pre-study preprocessing is not addressed.The guarantee is conditional on the fixed input data object.
5. Results
Across 210,000 held-out assessments, ARISE achieved the strongest aggregate and dataset-wise rankings across balanced accuracy, macro-F1, and Cohen’s κ. Performance peaked at 30 features overall, while stability varied substantially across datasets and supported flexible feature subsets.
- Evaluation design: 210,000 held-out assessments covered five datasets, seven feature-selection methods, eight feature budgets, three classifiers, and 250 outer-fold evaluations.Balanced accuracy, macro-F1, and Cohen’s κ were obtained for every evaluation.
- Aggregate performance: ARISE achieved the highest equal-dataset mean and best mean rank for balanced accuracy, macro-F1, and Cohen’s κ.SNR was the strongest comparator for balanced accuracy and macro-F1, while MI was closest for Cohen’s κ; ARISE’s advantages were 0.0221, 0.0233, and 0.0284, respectively.
- Dataset-wise performance: ARISE ranked first among the seven displayed methods in all 15 dataset–metric cells.Its advantage over the best remaining comparator ranged from 0.0051 to 0.0261 for balanced accuracy, 0.0044 to 0.0316 for macro-F1, and 0.0054 to 0.0383 for κ.
- Statistical analysis: 0.00100, 0.00107, and 0.00124 were the Friedman asymptotic P-values for balanced accuracy, macro-F1, and Cohen’s κ, respectively.Because only five datasets were independent experimental blocks, the analysis was exploratory rather than confirmatory.
- Feature-budget effects: 0.8289 was the maximum common-budget composite at 30 features, compared with 0.4595 at one feature and 0.8166 at 35 features.At 30 features, the separate means were 0.8501 for balanced accuracy, 0.8359 for macro-F1, and 0.8007 for κ; dataset-specific maxima occurred at 20, 15, 35, 35, and 30 features for D1–D5.
- Selected-set stability: Mean pairwise Jaccard overlap at 35 features ranged from 0.836 for D1 to 0.237 for D5.The lower overlaps in higher-dimensional datasets were interpreted as flexibility among near-equivalent feature subsets, supporting external validation and assay-level replication for biomarker interpretation.
6. Discussion
ARISE achieved the strongest aggregate performance across the evaluated datasets, metrics, and classifiers, while the discussion frames it as a transparent methodological candidate requiring independent evaluation. Its contributions and limitations concern signal integration, dataset context, feature-budget selection, stability interpretation, and translational validation.
- Performance: ARISE ranked first in all 15 dataset–metric cells and had the largest equal-dataset mean across five datasets, three metrics, and three fixed classifiers.Its three-metric composite also exceeded those of SNR and MI, the second- and third-ranked standalone methods.
- Validation and transparency: ARISE is presented as a transparent ordered-panel method and should undergo further independent evaluation using chance-corrected stability, pathway coherence, assay replication, and untouched cohorts.The discussion contrasts this transparency with graph or foundation models and cautions against interpreting selection frequency alone.
- Methodological contribution: Percentile normalization, class-balanced subsampling, residual correlation, pair coverage, and soft consensus jointly reconcile heterogeneous relevance, robustness, redundancy, multiclass, and profile-selection signals.The architecture averages parameters of eligible near-tied profiles instead of selecting a single winner.
- Scope and limitations: The five datasets represent distinct translational contexts, but their evidence remains limited by restricted feature availability, controlled experimental design, and internal resampling.D1 uses only 45 CpGs, D2 is a controlled Nutrimouse experiment, and D3 remains internal resampling rather than untouched-cohort validation.
- Feature budget and stability: Thirty selected features was the best common setting, whereas dataset-specific optima ranged from 15 to 35, supporting budget selection inside training unless panel size is fixed beforehand.Low overlap at larger budgets in D4 and D5, alongside partly combinatorial high overlap in D1, qualifies stability interpretation.
7. Conclusion
ARISE is presented as a coherent multiclass feature-selection framework that ranked first across all evaluated dataset–metric cells. The conclusion also defines methodological boundaries involving benchmark diversity, fixed classifiers, repeated resampling, and dataset-level inference.
- Conclusion: 15 dataset–metric cells were won by ARISE across five datasets, three metrics, and three fixed classifiers.The framework combines complementary relevance evidence, balanced-subsample robustness, residual redundancy control, class-pair coverage, nested profile racing, and convex consensus.
- Conclusion: 30 features produced the best common-budget mean, although dataset-specific optima differed.This supports evaluating feature budgets by dataset rather than assuming one universally optimal size.
- Methodological boundaries: Five heterogeneous benchmark datasets varied in sample size, dimensionality, molecular platform, and class structure within a controlled evaluation setting.Package-level preprocessing was retained to ensure reproducibility with established public data objects.
- Methodological boundaries: 250 held-out folds per condition were generated by the repeated five-fold design, with datasets—not individual folds—as the unit of across-study inference.Fixed classifiers were used intentionally so predictive differences could be attributed more directly to feature selection than to extensive learner-specific hyperparameter optimization.
Data availability
The analyzed datasets are publicly available through the data resources and originating studies identified in table 1 and section 3.1, with object-level provenance supported by reported object names and package versions.
- The datasets are publicly available through the data resources and originating studies identified in table 1 and section 3.1.
- Object names and package versions are reported in the Methods section to support object-level provenance.
Funding
The work was sponsored by United Arab Emirates University under grant 12B086.
- United Arab Emirates University sponsored the work under grant 12B086.
CRediT authorship contribution statement
The CRediT statement assigns methodological, analytical, software, writing, and supervisory contributions across the author team. Contributions also include investigation, data curation, validation, visualization, resources, funding acquisition, and project administration.
- Zardad Khan contributed conceptualization, methodology, software, formal analysis, visualization, supervision, and manuscript drafting and revision.
- Amjad Ali contributed methodology, software, investigation, data curation, validation, formal analysis, and manuscript revision.
- Naz Gul contributed investigation, data curation, formal analysis, validation, visualization, methodology, and manuscript revision.
- Sheema Gul contributed investigation and data curation, while Saeed Aldahmani contributed resources, validation, visualization, manuscript revision, conceptualization, funding acquisition, project administration, and supervision.