Source-linked AI summary

FMRI Clustering in AFNI: False Positive Rates Redux

Robert W. Cox, Gang Chen, Daniel R. Glen, Richard C. Reynolds, Paul A. Taylor

arXiv:1702.04845v1q-bio.QMstat.AP

TL;DR

이 논문은 FMRI cluster-size threshold를 정하는 과정에서 발생하는 중요하고 광범위한 문제를 다룬다. 저자들은 AFNI를 사용해 기존 결과를 재분석하고 일부 simulation을 반복했으며, 업데이트된 spatial-smoothness modeling과 permutation/randomization 접근법을 포함했다. 새로운 AFNI permutation/randomization 방법은 voxelwise p ≤ 0.01 전반에서 5%를 중심으로 밀집된 FPR을 산출했다.

  • 문제

    이 논문은 FMRI cluster-size threshold를 정하는 과정에서 발생하는 중요하고 광범위한 문제를 다룬다.

  • 방법

    저자들은 AFNI를 사용해 기존 결과를 재분석하고 일부 simulation을 반복했으며, 업데이트된 spatial-smoothness modeling과 permutation/randomization 접근법을 포함했다.

  • 결과

    새로운 AFNI permutation/randomization 방법은 voxelwise p ≤ 0.01 전반에서 5%를 중심으로 밀집된 FPR을 산출했다.

  • 시사점 및 한계

    결과는 명목상 5%에 가까운 FMRI cluster-level FPR을 얻기 위해 permutation/randomization 방법을 사용할 근거를 제공한다.

  • 시사점 및 한계

    업데이트된 parametric 방법은 FPR을 크게 낮췄지만, 많은 경우 명목상 5%를 여전히 웃돌았다.

Abstract

from arXiv · show

Recent reports of inflated false positive rates (FPRs) in FMRI group analysis tools by Eklund et al. (2016) have become a large topic within (and outside) neuroimaging. They concluded that: existing parametric methods for determining statistically significant clusters had greatly inflated FPRs ("up to 70%," mainly due to the faulty assumption that the noise spatial autocorrelation function is Gaussian- shaped and stationary), calling into question potentially "countless" previous results; in contrast, nonparametric methods, such as their approach, accurately reflected nominal 5% FPRs. They also stated that AFNI showed "particularly high" FPRs compared to other software, largely due to a bug in 3dClustSim. We comment on these points using their own results and figures and by repeating some of their simulations. Briefly, while parametric methods show some FPR inflation in those tests (and assumptions of Gaussian-shaped spatial smoothness also appear to be generally incorrect), their emphasis on reporting the single worst result from thousands of simulation cases greatly exaggerated the scale of the problem. Importantly, FPR statistics depend on "task" paradigm and voxelwise p-value threshold; as such, we show how results of their study provide useful suggestions for FMRI study design and analysis, rather than simply a catastrophic downgrading of the field's earlier results. Regarding AFNI (which we maintain), 3dClustSim's bug-effect was greatly overstated - their own results show that AFNI results were not "particularly" worse than others. We describe further updates in AFNI for characterizing spatial smoothness more appropriately (greatly reducing FPRs, though some remain >5%); additionally, we outline two newly implemented permutation/randomization-based approaches producing FPRs clustered much more tightly about 5% for voxelwise p<=0.01.

서론

이 보고서는 Eklund et al.’s 극단적 요약이 일반적인 false-positive 성능을 과장했다고 주장하고, AFNI의 bug, 개선된 parametric modeling, permutation 기반 대안을 평가한다. permutation/randomization 방법은 voxelwise p ≤ 0.01에서 FPR을 5% 부근에 촘촘히 모으는 반면, 개선된 parametric 방법은 inflation을 크게 줄이지만 항상 제거하지는 못한다.

  • 범위와 동기: 이 보고서는 ENK16’s data와 반복 simulation을 사용해 inflated cluster-count FPR을 다루면서, 극단적 주장이 자체 결과를 잘못 묘사했다고 지적한다.논의의 대상은 voxelwise p-values가 아니라 whole-brain cluster counts다.
  • 해석: Eklund et al.’s [2] “up to 70% FPR”은 일반적인 성능이 아니라 3,000개가 넘는 simulation 중 단 하나의 최악 결과를 강조했다.같은 기준을 적용하면 이들의 permutation method는 “up to 40% FPR”에 도달한다. 분포를 더 잘 비교하려면 median이나 range가 적절하다.
  • 과거 결과: Simulation을 반복한 결과, 과거의 3dClustSim bug는 무시할 수 없었지만 “particularly high” FPR을 만들지는 않았다.이 보고서는 AFNI’s bug effect에 대한 Eklund et al.’s 의 특성을 직접 반박한다.
  • 현재 방법: 새로운 non-Gaussian spatial autocorrelation 접근법은 3dttest++ parametric FPR을 크게 줄였지만, 많은 경우가 여전히 nominal 5%를 초과했다.AFNI는 spatial-smoothness 가정을 평가하고 autocorrelation function 추정을 위한 새로운 접근법을 구현했다.
  • 현재 방법: AFNI의 새로운 permutation/randomization 접근법을 사용하면 모든 voxelwise p ≤ 0.01 threshold에서 FPR이 5% 부근에 촘촘히 모였다.이 방법은 group-level residuals에서 cluster-size threshold를 생성하며, ENK16을 따른 simulation에서 시연되었다.

방법: 시뮬레이션

시뮬레이션은 198개의 Beijing-Zang FCON-1000 resting-state dataset을 사용해 ENK16의 기본 시나리오를 반복했으며, 업데이트된 AFNI preprocessing과 이에 상응하는 group-analysis 절차를 적용했다. False-positive rate를 평가하기 위해 smoothing, voxelwise threshold, 네 가지 pseudo-stimulus paradigm을 변화시켰다.

  • 시뮬레이션 dataset 및 processing: 시뮬레이션은 FCON-1000 의 198개 Beijing-Zang dataset을 사용해 [1]의 절차를 반복했다.AFNI preprocessing은 현재 권고를 따랐고 ENK16과는 다소 달랐지만, 결과는 상당히 유사했다.
  • Group analysis: Group analysis는 resting-state dataset에서 무작위로 추출한 1000개의 sub-collection과 ENK16에 기술된 3dttest++를 사용했다.
  • 시뮬레이션 parameter: 기본 ENK16 시나리오는 FWHM 4, 6, 8, and 10 mm의 Gaussian smoothing과 0.01, 0.005, 0.001의 voxelwise p-value threshold를 변화시켰다.중간값인 p=0.005 threshold는 ENK16의 0.01 및 0.001 설정에 추가되었다.
  • 시뮬레이션 paradigm: Null analysis는 네 가지 pseudo-stimulus timing을 사용했다. 즉 10-s 및 30-s ON/OFF block과 regular 2-s 및 random 1–4-s event-related task다.ENK16과 마찬가지로 각 case에서 모든 피험자에게 동일한 stimulus timing을 적용했다.

결과

AFNI의 3dClustSim 버그는 false-positive rate에 미치는 영향이 제한적이었으며, long-tailed spatial autocorrelation과 분석 조건이 inflation에 더 큰 영향을 미쳤다. 업데이트된 mixed-ACF와 nonparametric 접근법은 calibration을 개선했으며, NN=1 clustering에서는 검정한 모든 FPR이 nominal confidence interval 안에 머물렀다.

  • Nonparametric 접근법: NN=1 clustering의 모든 FPR은 검정한 voxelwise threshold 전반에서 nominal 95% confidence interval인 3.65–6.35% 안에 들어왔다.96개 case와 추가로 수행한 one-sample, paired, covariate test에서도 결과는 유사했다.

논의 및 결론 · “The Bug”와 일반적인 버그에 관한 주의

이전 3dClustSim 버그는 false-positive rate에 미친 영향이 비교적 작았으며 AFNI가 유사한 도구보다 나쁜 성능을 보이게 하지 않았다. 이 버그의 발견과 공개적인 수정은 재현성 실천을 보여 준다. 버그는 불가피하지만, 명확한 공개와 신속한 수정을 통해 타당성과 재현성을 보존할 수 있다.

  • “The Bug”와 일반적인 버그에 관한 주의: 3dClustSim 버그는 수정 후 false-positive rate를 낮췄지만, 전체적인 inflation에서 차지한 비중은 작은 요인이었다.저자들은 이 버그를 작은 요인으로 설명하며, 수정으로 FPR은 낮아졌지만 이 검정들에서 큰 변화가 발생하지는 않았다고 말한다.
  • “The Bug”와 일반적인 버그에 관한 주의: 수정 전후에 3dClustSim은 조사된 다른 software tool들과 비슷한 성능을 보였으므로, 그 자체로 “해석에 큰 영향을 미칠” 수는 없었다.버그의 존재는 유감스러웠지만, 저자들은 상당한 변화를 일으키려면 새롭게 도입된 방법이 필요했다고 주장한다.
  • “The Bug”와 일반적인 버그에 관한 주의: 버그, 부적절한 software 설정, 잘못된 method 구현은 보고된 결과의 타당성과 재현성을 훼손할 수 있다.저자들은 공개 사용을 목적으로 하는 software에서 버그를 방지하는 일을 핵심적인 관심사로 꼽는다.
  • “The Bug”와 일반적인 버그에 관한 주의: 버그에 대한 대중의 관심은 15년간의 뇌 연구와 최대 40,000편의 peer-reviewed publication을 폐기해야 할 근거인 것처럼 묘사되었다.저자들은 이러한 반응이 보고된 결과가 재현되지 않을 것이라는 점을 암묵적 또는 명시적으로 전제했다고 말한다.
  • “The Bug”와 일반적인 버그에 관한 주의: 버그의 공개는 FMRI에서 “재현성의 위기”를 입증한 것이 아니라, FMRI 분석이 재현될 수 있음을 검증한 것으로 묘사되었다.버그의 발견과 전파는 사용자와 유지관리자에게 좌절감을 주었지만 재현성 과정의 일부가 되었다.
  • “The Bug”와 일반적인 버그에 관한 주의: AFNI의 유지관리 철학은 가능한 한 조속히 버그를 수정하고 공개적으로 이용 가능한 software를 업데이트하는 것이다.또한 이 프로젝트는 주요 변경 사항을 게시하고, 사용자와 재현성을 지원하기 위해 업데이트, 변경 사항, 버그의 공개 online 목록을 유지한다.
  • “The Bug”와 일반적인 버그에 관한 주의: software 버그는 불가피하므로, 명확한 설명과 신속한 수리가 그 영향을 제한하는 최선의 수단으로 제시된다.저자들은 널리 사용되는 배포판도 정기적으로 버그 수정을 배포한다고 지적한다.

클러스터링의 현황

AFNI의 업데이트된 non-Gaussian ACF 접근법은 false-positive rate 제어를 개선하며, nonparametric clustering도 유망해 보이지만 계산 및 모델링 제약을 받는다. 저자들은 맥락 의존적 statistical thresholding, p-value를 넘어선 신중한 해석, 다양한 group-analysis model에 적용할 수 있는 실용적인 AFNI 절차를 권고한다.

  • 클러스터링의 현황: 업데이트된 AFNI의 non-Gaussian spatial autocorrelation modeling은 FPR 제어 가능성을 크게 개선하며, 새로운 nonparametric clustering은 nominal rate에 가까운 FPR을 산출한다.Permutation/randomization method는 가정이 적지만, 실제 활용도는 model complexity와 implementation 제약에 좌우된다.
  • 클러스터링의 현황: Permutation testing은 보편적으로 유리하지 않다. 복잡한 model, covariate, missing data, mixed effect에서는 계산상 또는 방법론적으로 지나치게 부담이 클 수 있으며, power를 희생할 수도 있다.고정된 permutation 횟수는 p-value의 하한을 부과하므로, parametric method나 더 많은 permutation이라면 검출했을 작지만 highly significant한 cluster를 놓칠 수 있다.
  • FMRI 통계에 대한 최종적인(현재로서는) 고찰과 몇 가지 권고: Equitable clustering은 선택적인 blurring radius, neighborhood value, voxelwise p-value에 대한 결과의 민감도를 낮추지만, statistical threshold는 여전히 해석의 한 부분일 뿐이다.저자들은 statistical thresholding이 neuroscientific conclusion을 결정하는 것이 아니라 이에 정보를 제공해야 한다고 주장한다.
  • FMRI 통계에 대한 최종적인(현재로서는) 고찰과 몇 가지 권고: 정확하게 정렬된, 약간 확장된 gray-matter mask로 clustering을 제한하면 예상되는 FPR 증가 없이 cluster-size threshold를 약 25% 낮출 수 있다.이 접근법은 정밀한 nonlinear alignment에 의존하며, 모든 관심 영역과 충분히 정렬된 피험자가 포함되었는지 검증해야 한다.

부록 A. Eklund et al.’s FPR 결과에 대한 추가 파싱 및 플로팅

부록 A에서는 통계 검정, 표본 설계, 자극 패러다임, voxelwise threshold에 따라 Eklund et al.’s FPR 결과를 재분석한다. 결과는 FPR이 이러한 선택에 따라 체계적으로 달라지며, 특히 p=0.001에서 two-sample event-related tests의 경우 5%에 매우 근접하게 일치함을 보여준다.

  • 부록 A. Eklund et al.’s FPR 결과에 대한 추가 파싱 및 플로팅: p=0.001에서 two-sample event-related tests는 모든 방법의 결과를 명목상 5% FPR에 가깝게 밀집시켰다.이 패턴은 event-related stimuli에 대해 소프트웨어 전반에서 관찰되었다.
  • 부록 A. Eklund et al.’s FPR 결과에 대한 추가 파싱 및 플로팅: 일부 parameter 조합에서는 소프트웨어 간 차이가 나타났지만, 대부분의 subset에서 parametric methods는 상당히 유사하게 수행되었다.부록에서는 이러한 패턴을 one-sample과 two-sample testing, 그리고 block 또는 event-related stimulus design별로 제시한다.
  • 부록 A. Eklund et al.’s FPR 결과에 대한 추가 파싱 및 플로팅: event-related stimuli에서는 two-sample tests가 소프트웨어와 voxelwise thresholds 전반에서 일관되게 더 낮은 FPR distributions를 보였다.해당 결과는 대응되는 결과보다 mean과 maximum이 낮고 outlier도 적었다.
  • 부록 A. Eklund et al.’s FPR 결과에 대한 추가 파싱 및 플로팅: two-sample testing은 더 짧은 B1 blocks에서도 FPR을 낮췄지만, 더 긴 B2 blocks에서는 그 감소가 훨씬 덜 뚜렷했다.따라서 그 효과는 stimulus paradigms와 block durations에 따라 달랐다.
  • 부록 A. Eklund et al.’s FPR 결과에 대한 추가 파싱 및 플로팅: voxelwise threshold, statistical test, stimulus paradigm과 연관된 FPR 패턴은 FMRI study design and analysis에 참고가 될 수 있다.이러한 비교는 단일한 전체 소프트웨어 순위가 아니라, 파싱된 결과가 제공하는 유용한 특징으로 제시된다.

부록 B. 3dttest++ 및 6가지 서로 다른 cluster-size thresholding 방법을 사용한 연구

1,536개의 null-simulation 추정치에서 permutation/randomization 방법은 대체로 FPR을 5%에 가깝게 통제한 반면, parametric cluster threshold는 크게 달라졌다. Gaussian ACF threshold는 mixed-model ACF threshold보다 일관되게 더 관대했으며, subject-mean mixed-model 추정치는 더 작은 voxelwise p-value와 더 강한 smoothing에서 대체로 양호한 성능을 보였다.

  • 결과: Permutation/randomization 방법은 실험 설계 전반에서 목표한 5% FPR에 가까운 양호한 통제를 대체로 제공했지만, 일부 경우에는 보수적이거나 관대했다.두 방법은 전반적으로 극단적으로 부정확하지 않았다. 특히 resting-state null이 구조가 없는 noise가 아니었기 때문에 더 낮은 voxelwise p-value와 two-sample testing이 선호되었다.
  • 결과: Parametric thresholding은 훨씬 더 크게 달라졌으며, 동일한 FWHM에서도 Gaussian ACF threshold가 mixed-model ACF threshold보다 일관되게 더 관대했다.Mixed model은 더 장거리의 correlation을 허용하므로 Gaussian model에 비해 threshold가 더 보수적이었다.
  • 결과: Subject-mean mixed-model ACF parameter는 t-test residual에서 추정한 parameter보다 대체로 더 보수적이었으며, p-value 0.001과 0.002 및 smoothing level 8과 10 mm에서 FPR이 양호했다.이러한 결과는 논문의 권고를 뒷받침했다.

그림

그림은 AFNI software 시나리오와 statistical testing 구성 전반의 false positive rates를 검토한다. 또한 spatial-autocorrelation fit, cluster-size threshold, noise smoothness, 대안적 thresholding 방법을 보여준다.

  • Figure 1은 1000회의 two-sample test를 사용해 AFNI software 시나리오 전반의 FPR을 검토한다.
  • Figure 2는 [2]에서 검토한 FPR 결과를 통합된 test 결과 전반에 걸쳐 요약한다.
  • Figure 3은 원래의 Gaussian fit과 전역적으로 추정한 fit을 비교한다.
  • Figure 4는 198개의 추정된 ACF 사례에 대한 3dClustSim의 cluster-size threshold를 보여준다.
  • Figures 5–7은 3dttest++, FWQM, ETAC을 포함한 업데이트된 clustering 및 thresholding 접근법의 FPR 또는 noise-smoothness measure를 묘사한다.

“FMRI Clustering in AFNI: False Positive Rates Redux”†을 위한 보충 정보 · 처리 스크립트

보충 처리 스크립트는 논문의 false-positive-rate 결과를 생성하는 데 사용된 AFNI의 전처리, 피험자 수준 분석, 집단 수준 simulation을 문서화한다. 또한 본문 그림의 simulation FPR 값을 포함한 보충 표를 식별한다.

  • 처리 스크립트: 해부학적 데이터는 피험자당 한 번 MNI 2009 template으로 nonlinear-warped한 후, 반복 계산을 줄이기 위해 simulation 전반에서 재사용했다.Script_1.warper.csh는 해부학적 공간을 표준 공간에 정렬하며, 그 출력은 이후 피험자 처리 명령에서 사용된다.
  • 처리 스크립트: 기능적 데이터는 alignment, normalization, blurring, masking, scaling, regression을 포함하는 afni_proc.py 생성 pipeline을 사용해 block 및 event-related designs별로 처리했다.block-design 및 event-related 스크립트는 앞서 수행한 해부학적 정렬을 사용하고, 설계별 stimulus input과 response model을 지정한다.
  • 처리 스크립트: 처리 pipeline은 alignment, motion estimate 및 관련 피험자 수준 출력을 검사하기 위한 quality-control scripts를 생성했다.단일 afni_proc.py 명령이 전체 pipeline과 관련 quality-control scripts를 생성한다.
  • 처리 스크립트: 집단 분석은 분석된 dataset을 무작위로 sampling한 다음 3dttest++와 3dClustSim을 실행해 thresholded activation map을 생성했으며, 그 개수로 보고된 FPR statistics를 산출했다.이 절차로 본문에 도시된 false-positive-rate 결과를 생성했다.
  • 처리 스크립트: 스크립트는 block 및 event-related 분석 모두에 대해 stimulus file, response model, blur radius, subject ID를 포함한 테스트된 처리 parameter를 노출한다.block 및 event 스크립트는 각각 이러한 분석 input에 대응하는 네 개의 argument를 요구한다.
  • 처리 스크립트: 보충 데이터는 Figures 1 and 4에 사용되었으며, Supplementary Tables 1 and 2는 해당 simulation 결과의 decimal FPR values를 보고한다.Table 1은 sample당 20명의 피험자를 사용한 1000 two-sample t-tests의 software scenario를 다루며, Table 2는 3dttest++ Clustsim thresholds와 one-sided NN=1 clustering을 사용한다.
Loading 1702.04845v1…