Source-linked AI summary

FMRI Clustering in AFNI: False Positive Rates Redux

Robert W. Cox, Gang Chen, Daniel R. Glen, Richard C. Reynolds, Paul A. Taylor

arXiv:1702.04845v1q-bio.QMstat.AP

TL;DR

The paper addresses an important, widespread problem in determining FMRI cluster-size thresholds. The authors reanalyze prior results and repeat selected simulations using AFNI, including updated spatial-smoothness modeling and permutation/randomization approaches. New AFNI permutation/randomization methods produced FPRs clustered tightly around 5% across voxelwise p ≤ 0.01.

  • Problem

    The paper addresses an important, widespread problem in determining FMRI cluster-size thresholds.

  • Method

    The authors reanalyze prior results and repeat selected simulations using AFNI, including updated spatial-smoothness modeling and permutation/randomization approaches.

  • Results

    New AFNI permutation/randomization methods produced FPRs clustered tightly around 5% across voxelwise p ≤ 0.01.

  • Takeaways & Limitations

    The results support using permutation/randomization methods to obtain FMRI cluster-level FPRs near the nominal 5% rate.

  • Takeaways & Limitations

    Updated parametric methods greatly reduced FPRs, but they remained above the nominal 5% in many cases.

Abstract

from arXiv · show

Recent reports of inflated false positive rates (FPRs) in FMRI group analysis tools by Eklund et al. (2016) have become a large topic within (and outside) neuroimaging. They concluded that: existing parametric methods for determining statistically significant clusters had greatly inflated FPRs ("up to 70%," mainly due to the faulty assumption that the noise spatial autocorrelation function is Gaussian- shaped and stationary), calling into question potentially "countless" previous results; in contrast, nonparametric methods, such as their approach, accurately reflected nominal 5% FPRs. They also stated that AFNI showed "particularly high" FPRs compared to other software, largely due to a bug in 3dClustSim. We comment on these points using their own results and figures and by repeating some of their simulations. Briefly, while parametric methods show some FPR inflation in those tests (and assumptions of Gaussian-shaped spatial smoothness also appear to be generally incorrect), their emphasis on reporting the single worst result from thousands of simulation cases greatly exaggerated the scale of the problem. Importantly, FPR statistics depend on "task" paradigm and voxelwise p-value threshold; as such, we show how results of their study provide useful suggestions for FMRI study design and analysis, rather than simply a catastrophic downgrading of the field's earlier results. Regarding AFNI (which we maintain), 3dClustSim's bug-effect was greatly overstated - their own results show that AFNI results were not "particularly" worse than others. We describe further updates in AFNI for characterizing spatial smoothness more appropriately (greatly reducing FPRs, though some remain >5%); additionally, we outline two newly implemented permutation/randomization-based approaches producing FPRs clustered much more tightly about 5% for voxelwise p<=0.01.

Introduction

The report argues that Eklund et al.’s extreme summaries overstated typical false-positive performance and evaluates AFNI’s bug, updated parametric modeling, and permutation-based alternatives. It finds that permutation/randomization methods cluster FPRs tightly around 5% for voxelwise p ≤ 0.01, while improved parametric methods greatly reduce but do not always eliminate inflation.

  • Scope and motivation: The report addresses inflated cluster-count FPRs using ENK16’s data and repeated simulations, while cautioning that extreme claims mischaracterized their own results.The discussion concerns whole-brain cluster counts rather than voxelwise p-values.
  • Interpretation: Eklund et al.’s [2] “up to 70% FPR” emphasized the single worst result among more than 3,000 simulations rather than typical performance.Under the same criterion, their permutation method would reach “up to 40% FPR”; medians or ranges would better compare distributions.
  • Past results: Repeating simulations showed that the former 3dClustSim bug was not negligible but did not produce “particularly high” FPRs.The report directly challenges Eklund et al.’s characterization of AFNI’s bug effect.
  • Present methods: The new non-Gaussian spatial autocorrelation approach greatly reduced 3dttest++ parametric FPRs, though many cases remained above the nominal 5%.AFNI evaluated spatial-smoothness assumptions and implemented a new approach for estimating the autocorrelation function.
  • Present methods: FPRs clustered tightly about 5% across all voxelwise p ≤ 0.01 thresholds with AFNI’s new permutation/randomization approach.The method generates cluster-size thresholds from group-level residuals and was demonstrated on simulations following ENK16.

Methods: Simulations

The simulations repeated ENK16’s basic scenarios using 198 Beijing-Zang FCON-1000 resting-state datasets, with updated AFNI preprocessing and comparable group-analysis procedures. They varied smoothing, voxelwise thresholds, and four pseudo-stimulus paradigms to assess false-positive rates.

  • Simulation datasets and processing: Simulations repeated the procedures in [1] using 198 Beijing-Zang datasets from FCON-1000.AFNI preprocessing followed current recommendations, differed somewhat from ENK16, and produced quite comparable results.
  • Group analysis: Group analyses used 1000 random sub-collections of the resting-state datasets and 3dttest++ as described in ENK16.
  • Simulation parameters: The basic ENK16 scenarios varied Gaussian smoothing at FWHM 4, 6, 8, and 10 mm and voxelwise p-value thresholds of 0.01, 0.005, and 0.001.The intermediate p=0.005 threshold was added to ENK16’s 0.01 and 0.001 settings.
  • Simulation paradigms: Null analyses used four pseudo-stimulus timings: 10-s and 30-s ON/OFF blocks, plus regular 2-s and random 1–4-s event-related tasks.Each subject received the same stimulus timing for each case, as in ENK16.

Results

AFNI’s 3dClustSim bug had a modest effect on false-positive rates, while long-tailed spatial autocorrelation and analysis conditions more substantially influenced inflation. Updated mixed-ACF and nonparametric approaches improved calibration, with NN=1 clustering keeping all tested FPRs within the nominal confidence interval.

  • Nonparametric approaches: All NN=1 clustering FPRs fell within the nominal 95% confidence interval of 3.65–6.35% across tested voxelwise thresholds.The results were similar across 96 cases and additional one-sample, paired, and covariate tests.

Discussion and Conclusions · A note on “The Bug” and on bugs in general

The older 3dClustSim bug had a relatively small effect on false-positive rates and did not make AFNI perform worse than comparable tools. Its discovery and public correction instead illustrate reproducibility practices: bugs are inevitable, but clear disclosure and rapid repair help preserve validity and reproducibility.

  • A note on “The Bug” and on bugs in general: The 3dClustSim bug reduced false-positive rates after correction, but it was a minor factor in the overall inflation.The authors describe the bug as a minor feature whose correction lowered FPRs without producing a major change in these tests.
  • A note on “The Bug” and on bugs in general: Before and after the fix, 3dClustSim performed comparably to the other investigated software tools and therefore could not by itself “have a large impact on the interpretation”.The bug’s presence was unfortunate, but the authors argue that substantial changes required newly introduced methods.
  • A note on “The Bug” and on bugs in general: Bugs, inappropriate software settings, and incorrect method implementations can damage the validity and reproducibility of reported results.The authors identify preventing bugs as a central concern for software intended for public use.
  • A note on “The Bug” and on bugs in general: The public focus on the bug was portrayed as grounds for rejecting 15 years of brain studies and up to 40,000 peer-reviewed publications.The authors say this reaction tacitly or explicitly assumed that reported results would be unreproducible.
  • A note on “The Bug” and on bugs in general: Rather than demonstrating “a crisis of reproducibility” in FMRI, publicizing the bug was described as verification that FMRI analysis could be reproduced.The bug’s discovery and dissemination became part of the reproducibility process, despite being frustrating for users and maintainers.
  • A note on “The Bug” and on bugs in general: AFNI’s maintenance philosophy is to correct bugs and update publicly available software as soon as practicable.The project also posts significant changes and maintains a public online list of updates, changes, and bugs to support users and reproducibility.
  • A note on “The Bug” and on bugs in general: Because software bugs are inevitable, clear descriptions and rapid repair are presented as the best tools for limiting their effects.The authors note that even widely used distributions regularly release bug fixes.

The state of clustering

AFNI’s updated non-Gaussian ACF approach improves false-positive-rate control, while nonparametric clustering also appears promising but is limited by computational and modeling constraints. The authors recommend context-dependent statistical thresholding, careful interpretation beyond p-values, and practical AFNI procedures for different group-analysis models.

  • The state of clustering: Updated AFNI modeling of non-Gaussian spatial autocorrelation greatly improves FPR controllability, while new nonparametric clustering produces FPRs near the nominal rate.Permutation/randomization methods require few assumptions, but their practical utility depends on model complexity and implementation constraints.
  • The state of clustering: Permutation testing is not universally advantageous: it can be computationally or methodologically prohibitive for complex models, covariates, missing data, and mixed effects, and may sacrifice power.A fixed number of permutations imposes a lower p-value bound, potentially missing small highly significant clusters that parametric methods or more permutations would detect.
  • Final (for now) thoughts on statistics in FMRI and some recommendations: Equitable clustering makes results less sensitive to optional blurring radius, neighborhood value, and voxelwise p-value choices, while statistical thresholds remain only one part of interpretation.The authors argue that statistical thresholding should inform, not determine, neuroscientific conclusions.
  • Final (for now) thoughts on statistics in FMRI and some recommendations: Restricting clustering to an accurately aligned, slightly inflated gray-matter mask reduced cluster-size thresholds by about 25% without expected FPR inflation.This approach depends on precise nonlinear alignment and must be vetted to ensure all regions of interest and adequately aligned subjects are included.

Appendix A. Additional parsing and plotting of Eklund et al.’s FPR results

Appendix A reanalyzes Eklund et al.’s FPR results by statistical test, sample design, stimulus paradigm, and voxelwise threshold. The results show that FPRs vary systematically with these choices, with especially close agreement near 5% for two-sample event-related tests at p=0.001.

  • Appendix A. Additional parsing and plotting of Eklund et al.’s FPR results: At p=0.001, two-sample event-related tests clustered all methods closely near the nominal 5% FPR.This pattern was observed across software for event-related stimuli.
  • Appendix A. Additional parsing and plotting of Eklund et al.’s FPR results: Although software differences appeared for some parameter combinations, parametric methods performed quite similarly across most subsets.The appendix presents these patterns by one- versus two-sample testing and by block or event-related stimulus design.
  • Appendix A. Additional parsing and plotting of Eklund et al.’s FPR results: For event-related stimuli, two-sample tests produced uniformly lower FPR distributions across software and voxelwise thresholds.They had lower means, lower maxima, and fewer outliers than the corresponding results.
  • Appendix A. Additional parsing and plotting of Eklund et al.’s FPR results: Two-sample testing also reduced FPR for shorter B1 blocks, whereas the reduction was much less apparent for longer B2 blocks.The effect therefore differed across stimulus paradigms and block durations.
  • Appendix A. Additional parsing and plotting of Eklund et al.’s FPR results: FPR patterns associated with voxelwise threshold, statistical test, and stimulus paradigm may inform FMRI study design and analysis.These comparisons are presented as useful features of the parsed results rather than a single overall software ranking.

Appendix B. Studies with 3dttest++ and 6 different cluster-size thresholding methods

Across 1,536 null-simulation estimates, permutation/randomization methods generally controlled FPR near 5%, whereas parametric cluster thresholds varied substantially. Gaussian ACF thresholds were consistently more liberal than mixed-model ACF thresholds, with subject-mean mixed-model estimates performing reasonably under smaller voxelwise p-values and greater smoothing.

  • Results: Permutation/randomization methods generally provided good control near the desired 5% FPR across experimental designs, although some cases were conservative or liberal.The two methods were not dramatically wrong overall; lower voxelwise p-values and two-sample testing were favored, particularly because the resting-state null was not structure-free noise.
  • Results: Parametric thresholding varied much more, with Gaussian ACF thresholds consistently more liberal than mixed-model ACF thresholds despite matched FWHM.The mixed model permits longer-range correlations, explaining the more conservative thresholds relative to the Gaussian model.
  • Results: Subject-mean mixed-model ACF parameters were generally more conservative than parameters estimated from t-test residuals, with reasonable FPRs at p-values 0.001 and 0.002 and smoothing levels of 8 and 10 mm.These findings supported the paper’s recommendations.

FIGURES

The figures examine false positive rates across AFNI software scenarios and statistical testing configurations. They also illustrate spatial-autocorrelation fits, cluster-size thresholds, noise smoothness, and alternative thresholding methods.

  • Figure 1 examines FPRs across AFNI software scenarios using 1000 two-sample tests.
  • Figure 2 summarizes the FPR results examined in [2] across their combined test results.
  • Figure 3 compares the original Gaussian fit with a globally estimated fit.
  • Figure 4 shows cluster-size thresholds from 3dClustSim across 198 estimated ACF cases.
  • Figures 5–7 depict FPRs or noise-smoothness measures for updated clustering and thresholding approaches, including 3dttest++, FWQM, and ETAC.

Supplementary Information for “FMRI Clustering in AFNI: False Positive Rates Redux”† · Processing scripts

The supplementary processing scripts document AFNI’s preprocessing, subject-level analyses, and group-level simulations used to generate the paper’s false-positive-rate results. They also identify the supplementary tables containing the simulation FPR values for the main-text figures.

  • Processing scripts: Anatomical data were nonlinear-warped to the MNI 2009 template once per subject, then reused across simulations to reduce repeated computation.Script_1.warper.csh performs the anatomical-to-standard-space alignment, and its outputs are used by subsequent subject-processing commands.
  • Processing scripts: Functional data were processed separately for block and event-related designs using afni_proc.py-generated pipelines with alignment, normalization, blurring, masking, scaling, and regression.The block-design and event-related scripts use the prior anatomical alignment and specify design-specific stimulus inputs and response models.
  • Processing scripts: The processing pipeline generated quality-control scripts for inspecting alignment, motion estimates, and related subject-level outputs.A single afni_proc.py command creates the full pipeline and associated quality-control scripts.
  • Processing scripts: Group analyses randomly sampled analyzed datasets, then ran 3dttest++ and 3dClustSim to create thresholded activation maps whose counts yielded the reported FPR statistics.These procedures produced the false-positive-rate results plotted in the main text.
  • Processing scripts: The scripts expose tested processing parameters including stimulus file, response model, blur radius, and subject ID for both block and event-related analyses.The block and event scripts each require four arguments corresponding to these analysis inputs.
  • Processing scripts: The supplementary data were used for Figures 1 and 4, while Supplementary Tables 1 and 2 report decimal FPR values for the corresponding simulation results.Table 1 covers software scenarios using 1000 two-sample t-tests with 20 subjects per sample; Table 2 uses 3dttest++ Clustsim thresholds and one-sided NN=1 clustering.
Loading 1702.04845v1…