Source-linked AI summary
Cross-validation failure: small sample sizes lead to large error bars
Gaël Varoquaux
TL;DR
The paper examines whether cross-validation reliably measures predictive-model accuracy in neuroimaging, where sample sizes are often small. Using simple analyses and discussion of methodological variability, it shows that cross-validation errors are large and commonly underestimated, and considers ways to increase sample sizes despite data heterogeneity.
Problem
Small neuroimaging samples produce large uncertainty in predictive-model accuracy, while reliable validation is important for applications such as biomarkers and methods development.
Method
The paper uses simple analyses to examine cross-validation errors and discusses methodological vibration, data sharing, pooling, and alternative statistical controls.
Results
100 observations typically lead to ±10% errors in prediction accuracy, while fold-based error estimates strongly underestimate the true errors.
Takeaways & Limitations
Larger datasets and pooled or shared data are promising ways to improve predictive neuroimaging, including prediction on difficult problems despite increased variability.
Takeaways & Limitations
Neuroimaging data contain correlations and confounding effects that reduce statistical degrees of freedom and add intrinsic variance to prediction accuracy.
Abstract
from arXiv · showhide
Predictive models ground many state-of-the-art developments in statistical brain image analysis: decoding, MVPA, searchlight, or extraction of biomarkers. The principled approach to establish their validity and usefulness is cross-validation, testing prediction on unseen data. Here, I would like to raise awareness on error bars of cross-validation, which are often underestimated. Simple experiments show that sample sizes of many neuroimaging studies inherently lead to large error bars, eg $\pm$10% for 100 samples. The standard error across folds strongly underestimates them. These large error bars compromise the reliability of conclusions drawn with predictive models, such as biomarkers or methods developments where, unlike with cognitive neuroimaging MVPA approaches, more samples cannot be acquired by repeating the experiment across many subjects. Solutions to increase sample size must be investigated, tackling possible increases in heterogeneity of the data.
1. Introduction
Machine-learning methods have advanced many brain-imaging analyses, but their validity depends on accurate evaluation of generalization to unseen data. Cross-validation provides this test, yet neuroimaging studies can have large accuracy errors because of small samples, threatening conclusions from predictive models.
- Machine-learning methods have advanced decoding, information mapping, individual-difference prediction, encoding models, and reverse inference in brain imaging.
- Model validity is established by generalization: accurate prediction of properties of new data independent of the data used for fitting.Cross-validation splits available data into training and test sets.
- Cross-validation is central to statistical control in decoding, MVPA, searchlight, and computer-aided diagnostic methods.
- Cross-validation errors in measuring prediction accuracy are typically around ±10%, making its error bars a serious concern.
- Simple analyses show that cross-validation errors are inherent to small sample sizes and may undermine predictive-model reliability and publication credibility.The problem is especially severe for methods development and inter-subject diagnostics, while cognitive neuroscience studies often access more samples through multiple trials and subjects.
2. Results: cross-validation errors
Cross-validation estimates of prediction accuracy have large, sample-size-dependent errors in neuroimaging, and common fold-based error estimates can underestimate those errors. Simulations, empirical neuroimaging results, and a public challenge show that these discrepancies persist across validation strategies and data settings.
- Empirical cross-validation errors: At least ±10% confidence bounds occur for cross-validation accuracy estimates on neuroimaging data, regardless of using leave-one-run-out or random splitting.These bounds imply a 5% chance of being 10% above or below the true generalization accuracy.
- External test-set comparison: Public-versus-private test-set discrepancies in a 144-subject MRI challenge imply errors on the order of ±15%, with the single-measurement margin smaller by about a factor of two.The public and private test sets contained 30 and 28 subjects, respectively.
- Sample-size effects: For 100 samples, simulations reproduce the large neuroimaging error bars, while increasing sample size markedly narrows them.Both leave-one-out and more sophisticated cross-validation strategies retain large error bars at small sample sizes.
- Intrinsic sampling noise: ±7% confidence bounds arise even in the best-case binomial model with 100 independent observations and no decoder-training variability.Neuroimaging correlations and confounds reduce effective degrees of freedom and produce larger errors than the binomial model.
- Underestimated error bars: Fold-dispersion standard-error formulas underestimate confidence bounds by a factor of 0.7 in the best simulated case because fold predictions are not independent.Permutation testing provides good statistical control of prediction accuracy.
- Analytical variation: With fewer sessions, cross-validation scores under permuted labels deviate further from the 50% chance level, reaching 57% for six sessions and 71% for four.Using all 12 sessions produced scores ranging from 44% to 52%; subset means also varied notably.
3. Implications for neuroimaging
Large cross-validation uncertainty and analytical flexibility can weaken the reliability and generalizability of predictive neuroimaging findings. Larger, pooled, and cross-site datasets may help, but heterogeneity and acquisition constraints remain important boundaries.
- Large prediction variance combined with publication incentives can weaken scientific progress.The paper links large error bars and selective publication to unreliable conclusions.
- Analytical flexibility can create artificial cross-validated improvements that do not carry over to new data.The paper describes this as overfit and notes that independent test sets are difficult to acquire in neuroimaging.
- Reported accuracy is typically higher in small-sample studies than in studies with many samples.Uncontrolled heterogeneity may contribute to this decrease, although few studies directly compare heterogeneous large cohorts with controlled smaller groups.
- Pooling shared neuroimaging data can increase sample sizes while limiting acquisition costs.The paper cites platforms containing thousands of subjects and notes that pooling may use well-matched studies or broader literature coverage.
- Cross-site prediction can support generalization beyond a single scanner or cohort, but inclusion criteria and site differences can confound predictions and interpretations.The paper reports successful cross-site prediction while warning that heterogeneity may affect clinical relevance.
- Methods development should compare approaches across several datasets when neuroimaging sample sizes are typical.The paper presents this as the only sound approach it identifies for avoiding analytical loopholes in methods development.
4. Conclusion: improving predictive neuroimaging
Small samples make predictive-model accuracy estimates unreliable, while arbitrary analytic choices can produce improvements that fail to generalize. The paper therefore points toward larger datasets and recommendations validated across many datasets.
- Typical neuroimaging sample sizes of 100 observations produce approximately ±10% errors in prediction accuracy.Standard error across cross-validation folds strongly underestimates these errors because folds are far from independent.
- Arbitrary analytic-pipeline choices can improve measured prediction accuracy without improving performance on new data.This makes meaningful methods-development improvements difficult to establish.
- Small-sample predictive modeling is especially problematic in neuroimaging, where acquiring large datasets is difficult.The paper states that better classifiers or cross-validation approaches will not fix the underlying problem.
- A specific algorithm’s reported benefits were not reproducible on other datasets, although some core voxel-clustering ideas were later validated across many datasets.This example illustrates why methods recommendations should be tested beyond a single dataset.
- Larger datasets are presented as promising for neuroimaging, with multivariate models suited to capturing their richness and larger samples improving prediction on difficult problems despite increased variability.
Appendix A.1. The reusable holdout
The reusable holdout improves cross-validation by allowing a holdout set to be reused without overfitting, but it does not remove intrinsic uncertainty in prediction-accuracy measurements.
- The reusable holdout reuses a given holdout set while avoiding overfitting it.Its procedure jitters the prediction-error measure below a threshold and refuses conclusions beyond a confidence-related threshold.
- The technique does not fix intrinsic uncertainty in prediction accuracy; instead, it embeds that uncertainty in the validation procedure.For a given control on generalization performance, the threshold is set proportional to √n.
Appendix A.2. Confidence bounds for varying expected accuracy
Binomial confidence bounds vary with expected accuracy and sample size, providing a conservative lower bound on cross-validation uncertainty. The appendix also illustrates a two-stage split separating validation from decoding before cross-validation.
- Confidence bounds: The appendix considers multiclass decoding, where chance and observed accuracies may be lower, while the mechanisms driving cross-validation estimation errors remain the same.A binomial law still gives a lower bound on the error distribution in these settings.
- Confidence bounds: Binomial distributions show sampling-noise shapes for varying expected accuracies and sample numbers, for both null and observed values.
- Confidence bounds: The 5th and 95th percentiles of the binomial distribution provide conservative lower bounds on confidence bounds for varying expected accuracy and sample size.Experiments indicate that the binomial distribution underestimates errors, so actual confidence bounds are likely higher.
- Data splitting: Figure A2 depicts splitting data first into validation and decoding sets, followed by cross-validation on the decoding set.
Appendix C.1. Dataset simulation
The simulation generates two Gaussian classes in a 100-dimensional feature space, while using a 2D representation only for visualization. It compares standard error estimates with observed cross-validation error bounds.
- Dataset simulation: The simulated data contain two classes, each modeled by a Gaussian with identity covariance in 100 dimensions.
- Dataset simulation: The visualization reduces the simulated data to two features, although the actual experiments use 300 features.
- Error estimates: For a two-sided test, confidence bounds are calculated as 1.96 times the standard error of the mean.
- Error estimates: 30 samples correspond to ±13.8% and ±18.9% bounds, while 100 samples correspond to ±7.4% and ±10.3% bounds.
- Error estimates: 30 samples correspond to ±3.4% and ±15.3% bounds, while 100 samples correspond to ±2.0% and ±8.1% bounds.
- Error estimates: The figure compares confidence limits estimated from standard error across folds with the actual estimation error observed in simulations.The standard 95% confidence limit is substantially lower than the observed 95th percentile of error.
Appendix E. Experiments with the perfect predictor
Experiments with a data-independent perfect predictor isolate cross-validation mismatch from predictive-model instability. The mismatch distributions are assessed under leave-one-out and repeated 20% random-split strategies.
- The perfect predictor uses knowledge of the data-generating process to make the best possible classification decision independently of the data.
- Observed variability is attributable to sampling noise in the test set, with similar errors for leave-one-out and repeated 20% random splits.
- The perfect predictor experiments use a data-generating separation set to produce 75% prediction accuracy.
- Cross-validation mismatch is evaluated as the distribution between measured prediction accuracy and the predictor’s expected error across sample sizes.