Source-linked AI summary
Can parametric statistical methods be trusted for fMRI based group studies?
Anders Eklund, Thomas Nichols, Hans Knutsson
TL;DR
Common task-fMRI group analyses rely on parametric assumptions that have not been comprehensively validated with empirical data. The study evaluates these methods using millions of random analyses of resting-state data and finds conservative voxel-wise inference but invalid cluster-wise inference, whereas permutation testing is valid for both.
Problem
Common fMRI statistical methods have rarely been validated with real data, leaving the empirical validity of widely used group-analysis software insufficiently evaluated.
Method
The study applies millions of random group analyses to resting-state fMRI null data to evaluate SPM, FSL, AFNI, and non-parametric permutation testing.
Results
Parametric methods are conservative for voxel-wise inference but produce cluster-wise familywise error rates up to 60% versus the nominal 5%, while permutation testing is valid for both.
Takeaways & Limitations
The findings call into question the validity of parametric cluster-wise inference in group fMRI studies and support validating neuroimaging statistical methods.
Takeaways & Limitations
The null-data approach is potentially limited because resting-state fMRI may contain consistent trends or transients that compromise its suitability as null data.
Abstract
from arXiv · showhide
The most widely used task fMRI analyses use parametric methods that depend on a variety of assumptions. While individual aspects of these fMRI models have been evaluated, they have not been evaluated in a comprehensive manner with empirical data. In this work, a total of 2 million random task fMRI group analyses have been performed using resting state fMRI data, to compute empirical familywise error rates for the software packages SPM, FSL and AFNI, as well as a standard non-parametric permutation method. While there is some variation, for a nominal familywise error rate of 5% the parametric statistical methods are shown to be conservative for voxel-wise inference and invalid for cluster-wise inference; in particular, cluster size inference with a cluster defining threshold of p = 0.01 generates familywise error rates up to 60%. We conduct a number of follow up analyses and investigations that suggest the cause of the invalid cluster inferences is spatial auto correlation functions that do not follow the assumed Gaussian shape. By comparison, the non-parametric permutation test, which is based on a small number of assumptions, is found to produce valid results for voxel as well as cluster wise inference. Using real task data, we compare the results between one parametric method and the permutation test, and find stark differences in the conclusions drawn between the two using cluster inference. These findings speak to the need of validating the statistical methods being used in the neuroimaging field.
1. INTRODUCTION
Because fMRI statistical methods have rarely been validated with real data, this study evaluates complete group-inference workflows in SPM, FSL, and AFNI using null data and random group comparisons. The results indicate conservative voxel-wise inference but potentially severe false positives for parametric cluster-wise inference.
- Motivation: fMRI statistical methods have rarely been validated with real data, with prior validation relying mainly on simulated datasets because data collection is costly.International data-sharing initiatives made empirical evaluation using real neuroimaging data possible.
- Study objective: The study evaluates SPM, FSL, and AFNI in their entirety for group inference, including each package’s recommended integrated preprocessing steps.The evaluation addresses the unresolved statistical validity of multiple commonly used fMRI software packages.
- Evaluation strategy: Randomly comparing groups of healthy controls tests a true null hypothesis and enables empirical estimation of two-sample t-test false-positive rates.The procedure counts analyses producing any false positives under null data.
- Main result: Parametric methods are conservative for voxel-wise inference with familywise error control, but cluster-wise inference produces false-positive rates up to 60% versus the nominal 5%.Voxel-wise inference tests each voxel independently, whereas cluster-wise inference tests neighboring voxels simultaneously.
2. RESULTS
Across 1,920,000 random group analyses, parametric methods showed inflated cluster-wise familywise error rates but valid, often conservative voxel-wise inference. The study links cluster-wise inaccuracies to violations of spatial assumptions underlying random field theory.
- Experimental design: 1,920,000 random group analyses evaluated empirical false positive rates across SPM, FSL, AFNI, 128 parameter combinations, and three thresholding approaches.The analyses used 1,000 random analyses per parameter combination and tested five software tools.
- Main results: Parametric software’s cluster-wise familywise error rates far exceeded the nominal 5% level, whereas voxel-wise inferences were valid but often conservative.These findings were reported across two-sample and one-sample tests and group sizes of 20 and 40.
- Sources of inaccuracy: Cluster-wise inference depends on constant spatial smoothness and a squared-exponential spatial autocorrelation function in addition to Gaussian random field theory.These assumptions were investigated to explain inaccuracies in parametric cluster inference.
- Sources of inaccuracy: False clusters followed spatial smoothness patterns, with SPM, FSL OLS, and AFNI OLS showing generally greater smoothness in gray than white brain tissue.The tissue-dependent smoothness pattern had previously been observed in voxel-based morphometry data.
- Parametric versus non-parametric inference: 10 - 100 times larger non-parametric cluster p-values were obtained for CDT p = 0.01 and cluster size 400 voxels than the corresponding parametric p-values.For CDT p = 0.001 and cluster size 100 voxels, non-parametric p-values were approximately 1.25 - 10 times larger.
3. DISCUSSION
The discussion concludes that parametric group-fMRI methods are conservative for voxel-wise inference but can produce invalid, spuriously low cluster-wise p-values, especially under liberal cluster-defining thresholds. The authors attribute this primarily to mismatches between empirical spatial autocorrelation and model assumptions and advocate validating existing methods, including non-parametric permutation tests, on real fMRI data.
- Validity of parametric inference: Parametric SPM, FSL, and AFNI analyses can yield erroneous FWE-corrected cluster p-values that inflate statistical significance.The authors state that this calls into question the validity of many published fMRI studies using parametric cluster-wise inference.
- Validity of parametric inference: Parametric methods work well conservatively for voxel-wise inference but fail for cluster-wise inference across typical task-fMRI settings.The study is described as the most comprehensive evaluation of these settings across multiple software tools.
- Spatial autocorrelation: The empirical SACFs exceed squared-exponential expectations at large distances, helping explain why cluster inference performs better at p = 0.001 than at p = 0.01.Low thresholds create larger clusters whose radii make the SACF tail more influential.
- Software-specific effects: AFNI’s old 3dClustSim produced especially high familywise error rates because lower group smoothness and a bug yielded much lower cluster-extent thresholds.The corrected function raises the threshold by 15% at 8 mm smoothness, reducing AFNI familywise error rates particularly at higher smoothing.
- Analysis parameters: Liberal cluster-defining thresholds increase false positives for SPM, FSL, and AFNI, whereas the permutation test is unaffected by this parameter.The cluster-defining threshold is identified as the most important parameter for the parametric methods.
- Implications: The authors recommend validating existing statistical methods with real fMRI data and emphasize non-parametric permutation tests as an available alternative to methods based on questionable assumptions.They argue that many established statistical methods are seldom used in neuroimaging despite the feasibility of evaluating common methods empirically.
5. METHODS · 5.1. Resting state fMRI data
The study used resting-state fMRI from 396 healthy young adults to generate null group comparisons, then evaluated parametric and permutation-based analyses across multiple software packages and preprocessing pipelines. Group inference used voxel-wise and cluster-wise familywise-error correction, with significance defined at FWE-corrected p < 0.05.
- 5.1. Resting state fMRI data: 396 healthy controls were drawn from Beijing and Cambridge datasets, with 198 subjects per dataset and narrow age ranges near 21 years.Beijing scans used TR = 2 seconds and 225 time points per subject.
- 5.1. Resting state fMRI data: Because participants were healthy, similarly aged, and task-free, randomly split subgroups were expected to show no significant brain-activity differences.Analyses were performed separately for the Beijing and Cambridge datasets.
- 5.1.1. Random group generation: Random groups were formed by permuting all 198 subject numbers and assigning successive sets of 20 subjects to each group.Approximately 1.31·10^42 group divisions were possible for n = 198 and k = 40, although analyses were not independent.
- 5.1.2. Code availability: Parametric group analyses used SPM 8, FSL 5.0.7, and AFNI, while BROCCOLI implemented the non-parametric analyses to reduce processing time.FSL’s randomise function was available but was not used.
- 5.1.3. First level analyses: First-level analyses normalized subject data to standard brain space, applied motion correction, and tested smoothing levels of 4, 6, 8, and 10 mm full width at half maximum.Slice timing correction was not performed.
- 5.1.3. First level analyses: SPM, FSL, and AFNI used distinct registration workflows to transform fMRI data, contrasts, and variances into MNI or Talairach/template atlas spaces.SPM used coregistration and segmentation; FSL used FLIRT with BBR; AFNI used align_epi_anat.py and @auto_tlrc.
- 5.1.4. Group analyses: Random-effects group analyses used subject-level results; SPM used OLS, whereas FSL used both default FLAME1 and OLS models.BROCCOLI performed OLS regression in each permutation and recorded maximum voxel or cluster statistics to form empirical null distributions.
- 5.1.4. Group analyses: Voxel-wise and cluster-wise FWE correction used package-specific parametric procedures or empirical permutation distributions, with significance defined as any voxel or cluster having FWE-corrected p < 0.05.SPM and FSL used random field theory; AFNI used Bonferroni correction voxel-wise and 3dClustSim cluster-wise, while non-parametric tests used empirical maxima.
5.2. Task based fMRI data
Task-based fMRI datasets from OpenfMRI were analyzed with FSL to compare cluster-based parametric and non-parametric group-analysis p-values. The analyses used 5 mm smoothing and motion regressors, with datasets spanning several cognitive tasks and cluster-defining thresholds of p = 0.01 or p = 0.001.
- Analysis setup: OpenfMRI task datasets were analyzed only with FSL, using 5 mm smoothing and motion regressors in all cases.The analysis examined how cluster-based p-values differ between parametric and non-parametric group analyses.
- Datasets: The rhyme judgment dataset included 13 subjects performing word-versus-pseudoword rhyming judgments.Data were collected with a 2-second repetition time and 160 time points per subject, using separate word and pseudoword regressors.
- Datasets: The mixed-gambles dataset included 16 subjects deciding whether to accept gain/loss gambles during scanning.The data were collected using a 3 T Siemens Allegra scanner with a 2-second repetition time and 240 volumes per subject.