Source-linked AI summary

Accurate autocorrelation modeling substantially improves fMRI reliability

Wiktor Olszowy, John Aston, Catarina Rua, Guy B. Williams

arXiv:1711.09877v4q-bio.QM

TL;DR

The paper examines whether major fMRI packages accurately model temporal autocorrelation, comparing AFNI, FSL, SPM, and SPM FAST across diverse datasets. It finds better whitening with AFNI and FAST than with FSL and default SPM, with residual noise confounding first-level results, especially for low-frequency designs. The authors discuss implications for group analyses and recommend diagnostic plots, while noting limitations of the null-data choices.

  • Problem

    Temporal autocorrelation in fMRI can produce false positives when inadequately modeled, but comparative evidence across widely varied protocols and packages was limited.

  • Method

    The study compared AFNI, FSL, default SPM, and SPM FAST across 11 datasets spanning different protocols and populations, assessing whitening and first- and second-level results.

  • Results

    AFNI and FAST had the best whitening performance, whereas FSL and default SPM left substantial low-frequency autocorrelated noise that heavily confounded first-level results.

  • Takeaways & Limitations

    More accurate autocorrelation modeling could improve task-fMRI reliability, and diagnostic plots could help investigators detect residual autocorrelated noise.

  • Takeaways & Limitations

    Resting-state and task data were imperfect null data, because intrinsic activity or the true design could confound analyses using assumed designs.

Abstract

from arXiv · show

Given the recent controversies in some neuroimaging statistical methods, we compare the most frequently used functional Magnetic Resonance Imaging (fMRI) analysis packages: AFNI, FSL and SPM, with regard to temporal autocorrelation modeling. This process, sometimes known as pre-whitening, is conducted in virtually all task fMRI studies. We employ eleven datasets containing 980 scans corresponding to different fMRI protocols and subject populations. Though autocorrelation modeling in AFNI is not perfect, its performance is much higher than the performance of autocorrelation modeling in FSL and SPM. The residual autocorrelated noise in FSL and SPM leads to heavily confounded first level results, particularly for low-frequency experimental designs. Our results show superior performance of SPM's alternative pre-whitening: FAST, over SPM's default. The reliability of task fMRI studies would increase with more accurate autocorrelation modeling. Furthermore, reliability could increase if the packages provided diagnostic plots. This way the investigator would be aware of pre-whitening problems.

Methods

The study compared AFNI, FSL, SPM, and SPM FAST using diverse datasets and standardized analysis components, evaluating whitening, first-level specificity-sensitivity, and group-level effects.

  • Data: Eleven datasets covered resting-state and task studies, healthy subjects and patients, and varied TRs, field strengths, and voxel sizes.The datasets included 980 scans across different fMRI protocols and subject populations.
  • Analysis pipeline: Preprocessing, brain masks, MNI registrations, and multiple-comparison corrections were kept consistent across packages to limit confounding influences.The analysis compared residual power spectra, significant-cluster distributions, significant-voxel percentages, and positive rates.
  • Specificity and sensitivity: Null analyses treated resting-state scans, simulated resting-state scans, and task scans with wrong designs as data for assessing positive rates and specificity.For null data, the positive rate was interpreted as the familywise error rate.
  • Statistical analysis: All packages used the same canonical double-gamma HRF and temporal derivative, while physiological recordings were excluded because they were unavailable for most datasets.The canonical response peak was set at 5 seconds and the undershoot at around 15 seconds after stimulus onset.
  • Whitening evaluation: The study evaluated default SPM noise modeling alongside FAST and compared whitening through GLM-residual power spectra, where ideal white residuals produce flat spectra.The power spectra were averaged across brain voxels and subjects after variance normalization.

Discussion

Discussion of the results shows that AFNI and FAST generally outperformed FSL and SPM in whitening and first-level specificity, while residual autocorrelation particularly affected low-frequency designs. The findings also identify analysis-scope limitations and support diagnostic tools and improved pre-whitening methods.

  • First-level consequences: Lower assumed design frequencies increased significant voxels in FSL and SPM for several datasets, consistent with unremoved positive autocorrelation.Autocorrelated processes have increasing variances at lower frequencies, increasing mismatch between the true residual process and assumed model.
  • First-level consequences: AFNI and FAST produced larger differences between true and wrong designs than FSL and SPM, indicating more accurate autocorrelation modeling.AFNI and FAST also had lower familywise error rates than FSL and SPM in the reported analyses.
  • Robustness across protocols: The strongest apparent responses occurred in very-short-TR NKI scans, where TRs were 0.645s and 1.4s and adjacent time-point correlations are higher.The authors link short TRs with greater likelihood of significant activation, while noting that multiple protocols were affected.
  • Robustness across protocols: AFNI and FAST outperformed FSL and SPM across all four smoothing levels, despite smoothing altering gray- and white-matter signal mixtures.The study treated smoothing as a confounder rather than as the target of inference.
  • Limitations and practical implications: Results could be more robust with physiological recordings, while resting-state and simulated null data were imperfect representations for task-fMRI evaluation.Physiological recordings were not always acquired or incorporated, and simulated resting-state results differed substantially from acquired scans.
  • Second-level consequences: Imperfect pre-whitening confounded group analyses, although false positives did not propagate to the group level with AFNI’s 3dMEMA mixed-effects model.The authors suggest 3dMEMA makes little use of standard-error maps and resembles an SPM random-effects model in this respect.
  • Limitations and practical implications: The authors recommend diagnostic plots because popular packages do not provide them, and they support wider use of FAST in SPM and alternative FSL autocorrelation models.Diagnostic plots could reveal residual autocorrelated noise in GLM residuals.
Loading 1711.09877v4…