Source-linked AI summary

The reliability of a deep learning model in clinical out-of-distribution MRI data: a multicohort study

Gustav Mårtensson, Daniel Ferreira, Tobias Granberg, Lena Cavallin, Ketil Oppedal, Alessandro Padovani, Irena Rektorova, Laura Bonanni, Matteo Pardini, Milica Kramberger, John-Paul Taylor, Jakub Hort, Jón Snædal, Jaime Kulisevsky, Frederic Blanc, Angelo Antonini, Patrizia Mecocci, Bruno Vellas, Magda Tsolaki, Iwona Kłoszewska, Hilkka Soininen, Simon Lovestone, Andrew Simmons, Dag Aarsland, Eric Westman

arXiv:1911.00515v1physics.med-phcs.CVcs.LGeess.IVq-bio.QMstat.ML

TL;DR

Clinical MRI deployment is difficult to assess because research-trained deep learning models may encounter different scanners, protocols, and disease populations. This study trains AVRA on controlled combinations of 3117 neuroradiologist-rated MRI scans and evaluates it across clinical out-of-distribution cohorts. Performance generalized well to similar protocols, declined on visibly different clinical images, and improved when training included more heterogeneous data.

  • Problem

    Research-cohort MRI datasets often use harmonized protocols and restrictive populations, leaving limited evidence about model generalization to heterogeneous clinical out-of-distribution data.

  • Method

    The study trained multiple AVRA convolutional neural networks on different combinations of research and clinical cohorts and evaluated them on external clinics and systematically partitioned cohorts.

  • Results

    Performance generalized well to similar protocols but dropped in clinical cohorts with visibly different image contrasts; broader scanner and protocol diversity during training improved out-of-distribution performance.

  • Takeaways & Limitations

    Reliable clinical evaluation should include multiple external cohorts, and heterogeneous training data can improve robustness on unseen MRI data.

  • Takeaways & Limitations

    The study evaluates one network architecture, and its kappa-based assessment can be noisy because continuous predictions must be rounded.

Abstract

from arXiv · show

Deep learning (DL) methods have in recent years yielded impressive results in medical imaging, with the potential to function as clinical aid to radiologists. However, DL models in medical imaging are often trained on public research cohorts with images acquired with a single scanner or with strict protocol harmonization, which is not representative of a clinical setting. The aim of this study was to investigate how well a DL model performs in unseen clinical data sets---collected with different scanners, protocols and disease populations---and whether more heterogeneous training data improves generalization. In total, 3117 MRI scans of brains from multiple dementia research cohorts and memory clinics, that had been visually rated by a neuroradiologist according to Scheltens' scale of medial temporal atrophy (MTA), were included in this study. By training multiple versions of a convolutional neural network on different subsets of this data to predict MTA ratings, we assessed the impact of including images from a wider distribution during training had on performance in external memory clinic data. Our results showed that our model generalized well to data sets acquired with similar protocols as the training data, but substantially worse in clinical cohorts with visibly different tissue contrasts in the images. This implies that future DL studies investigating performance in out-of-distribution (OOD) MRI data need to assess multiple external cohorts for reliable results. Further, by including data from a wider range of scanners and protocols the performance improved in OOD data, which suggests that more heterogeneous training data makes the model generalize better. To conclude, this is the most comprehensive study to date investigating the domain shift in deep learning on MRI data, and we advocate rigorous evaluation of DL models on clinical data prior to being certified for deployment.

1. Introduction

Deep learning models show promise as clinical aids, but research-cohort training data often differs from clinical MRI in scanners, protocols, image quality, and disease populations. This study therefore investigates AVRA’s performance on out-of-distribution clinical neuroimaging data and whether more heterogeneous training improves generalization.

  • Clinical deployment requires models to work across different scanners, protocol parameters, and image quality.
  • Public neuroimaging datasets commonly use harmonized acquisition protocols and restrictive cohort criteria that differ from heterogeneous clinical populations.
  • Research-cohort training may reduce generalization to new scanners, protocols, and more heterogeneous clinical populations.
  • Prior work found that a clinical deep learning model performed poorly when applied to images from a new scanning device.
  • The study systematically evaluates AVRA on clinical out-of-distribution cohorts and tests whether broader training-cohort combinations improve generalization.

2. Material and methods

The study uses 3117 neuroradiologist-rated T1-weighted MRI images from research and clinical cohorts to test AVRA across acquisition and population shifts. Multiple cohort combinations, controlled partitions, external clinics, and repeat scans support evaluation of generalization and rating consistency.

  • The study analyzes 3117 MRI images from five cohorts differing in scanners, protocols, disease populations, and clinical or research settings.
  • A neuroradiologist rated all 3117 T1-weighted brain images using Scheltens’ medial temporal atrophy scale, independently of diagnosis, age, and sex.
  • AVRA is a recurrent convolutional neural network that processes unprocessed MRI volumes slice-by-slice and predicts a continuous atrophy score.
  • The experiment trains multiple AVRA versions on different cohort combinations while focusing on the MTA scale and retaining the previously described architecture and hyperparameters.
  • E-DLB was partitioned into non-overlapping training and test subsets to simulate deployment without local labels, retraining with 25% additional data, or retraining with 50% additional data.
  • Performance was evaluated on two external clinics with single scanners but visibly different image intensities, using Cohen’s linearly weighted kappa to assess rating agreement.

3. Results

AVRA generalized less reliably to clinical cohorts when trained only on research data, while adding clinical training data generally improved agreement and accuracy. Predictions also differed systematically across unseen centers, and some scanner-specific variability remained.

  • Training only on ADNI produced a general performance drop in clinical cohorts, especially E-DLBC1, compared with ADNI testing.Adding AddNeuroMed helped little, whereas including clinical MemClin had a positive impact.
  • Including clinical cohorts in training generally improved rating agreements and accuracies on clinical test sets, although improvements were not consistent.
  • Models trained with and without wider-range clinical images made systematically different predictions in unseen E-DLBC1 and E-DLBC2 centers, most notably E-DLBC1.Neither center contributed images to either training set; the broader training set included more varied protocols.
  • Test-retest ratings showed small within-model intra-subject variability for most subjects, with the clearest differences arising in images acquired with Siemens Trio 3T.The direction of differences varied across subjects, suggesting possible protocol-specific rating effects.
  • Table 3 compares Cohen’s κw and MSE across test sets and training-cohort combinations while keeping training size fixed at N = 1568.Overlapping train-test combinations receive no agreement score, and the greatest agreement for each test set is bolded.

4. Discussion

The study shows that CNN performance generally drops on clinical MRI data outside homogeneous research distributions, while more heterogeneous training improves robustness without substantially harming within-distribution performance. Results also show that external evaluation, scanner effects, disease-population differences, and model limitations all matter when considering clinical deployment.

  • CNN performance generally dropped on clinical data when models were trained on homogeneous research cohorts.
  • Training on images from a wider range of scanners and protocols increased robustness on unseen out-of-distribution data without substantially damaging within-distribution performance.The improvement persisted when training-set size and label distribution were fixed.
  • A single external center was insufficient to assess generalization because agreement was substantially lower in E-DLBC1 than in E-DLBC2 after research-cohort training.The contrasting external-center results indicate that domain shift can vary across clinical cohorts.
  • Agreement was higher in DLB and PDD populations than in the AD population, potentially because the ADNI-trained model systematically underestimated higher MTA values.The authors suggest disease-unspecific visual rating scales may support cross-population generalization.
  • Test-retest ratings were generally consistent, although images from the Siemens Trio 3T showed notable inter-scanner differences and the heterogeneous model had slightly higher variability.Within-scanner and within-field-strength variability was practically absent, supporting potential usefulness for longitudinal studies collected harmonizedly.
  • The findings are limited by evaluation of a single network architecture, possible overfitting to the training protocol, metric sensitivities, and no assessment of preprocessing or intensity normalization.

5. Conclusion

AVRA generalized well to cohorts with protocols similar to training data but performed substantially worse on some clinical data with different tissue contrasts. More heterogeneous training data improved out-of-distribution performance, supporting rigorous clinical evaluation before deployment.

  • AVRA trained on homogeneous research data generalized well to cohorts with similar protocols but worse to clinical data.
  • Performance dropped to an unacceptably low level on images from one specific memory clinic.
  • Including data from a wider range of scanners and protocols improved performance on out-of-distribution data.
  • The study advocates rigorous testing of deep learning models on out-of-distribution data before clinical deployment.

Appendix A. Supplementary data

The supplementary material documents cohort definitions, MRI scanners and protocols, and accuracy analyses across training and test-set combinations.

  • Supplementary accuracy results complement MSE and Cohen’s κw, while full scanning parameters are provided in Tables A.2–A.6.
  • Table A.1 holds the training size at N = 1568 and fixes the training distribution across cohort combinations.
  • A ✓ marks cohorts included in training, and agreement is omitted when training and test images overlap.
  • Cohort labels distinguish memory-clinic data, complete or pathology-selected E-DLB subsets, random E-DLB samples, and center-specific splits.
  • Tables A.2–A.6 list scanner and MRI parameters for ADNI, AddNeuroMed, MemClin, E-DLB, and test-retest cohorts.
Loading 1911.00515v1…