Source-linked AI summary

Uncovering convolutional neural network decisions for diagnosing multiple sclerosis on conventional MRI using layer-wise relevance propagation

Fabian Eitel, Emily Soehler, Judith Bellmann-Strobl, Alexander U. Brandt, Klemens Ruprecht, René M. Giess, Joseph Kuchling, Susanna Asseyer, Martin Weygandt, John-Dylan Haynes, Michael Scheel, Friedemann Paul, Kerstin Ritter

arXiv:1904.08771v1cs.CV

TL;DR

Deep-learning MRI classifiers can be difficult to interpret, limiting their clinical validation. The paper combines CNNs, transfer learning, and LRP heatmaps to diagnose MS and reveal the image features behind individual decisions. The framework achieved strong classification performance and highlighted lesions, lesion locations, and non-lesional brain regions relevant to MS.

  • Problem

    Deep-learning diagnostic decisions are often non-transparent, hindering clinical integration, error tracking, and knowledge discovery.

  • Method

    The study pre-trained a CNN on ADNI MRI data, fine-tuned it to classify MS, and used LRP heatmaps to attribute each decision to MRI voxels.

  • Results

    A pre-trained CNN achieved balanced accuracy of 87.04% and identified lesions, lesion locations, non-lesional white matter, and gray-matter regions such as the thalamus.

  • Takeaways & Limitations

    LRP helped explain individual CNN decisions and assess whether the model learned MS-relevant information from conventional MRI.

  • Takeaways & Limitations

    The MS sample comprised 147 subjects, which the authors state is generally too small to learn robust representations and generalize to other datasets.

Abstract

from arXiv · show

Machine learning-based imaging diagnostics has recently reached or even superseded the level of clinical experts in several clinical domains. However, classification decisions of a trained machine learning system are typically non-transparent, a major hindrance for clinical integration, error tracking or knowledge discovery. In this study, we present a transparent deep learning framework relying on convolutional neural networks (CNNs) and layer-wise relevance propagation (LRP) for diagnosing multiple sclerosis (MS). MS is commonly diagnosed utilizing a combination of clinical presentation and conventional magnetic resonance imaging (MRI), specifically the occurrence and presentation of white matter lesions in T2-weighted images. We hypothesized that using LRP in a naive predictive model would enable us to uncover relevant image features that a trained CNN uses for decision-making. Since imaging markers in MS are well-established this would enable us to validate the respective CNN model. First, we pre-trained a CNN on MRI data from the Alzheimer's Disease Neuroimaging Initiative (n = 921), afterwards specializing the CNN to discriminate between MS patients and healthy controls (n = 147). Using LRP, we then produced a heatmap for each subject in the holdout set depicting the voxel-wise relevance for a particular classification decision. The resulting CNN model resulted in a balanced accuracy of 87.04% and an area under the curve of 96.08% in a receiver operating characteristic curve. The subsequent LRP visualization revealed that the CNN model focuses indeed on individual lesions, but also incorporates additional information such as lesion location, non-lesional white matter or gray matter areas such as the thalamus, which are established conventional and advanced MRI markers in MS. We conclude that LRP and the proposed framework have the capability to make diagnostic decisions of...

1. Introduction

MS diagnosis relies on clinical presentation and conventional T2-weighted MRI, while deep-learning decisions remain difficult to interpret. This study investigates whether LRP can reveal MRI features used by a CNN and help validate its decisions.

  • MS diagnosis commonly combines clinical presentation with white matter lesions visible in conventional T2-weighted MRI.
  • Conventional T2-weighted MRI is sensitive to MS-relevant white matter lesions but relatively nonspecific to underlying disease processes.
  • Earlier MS machine-learning approaches combined standard algorithms with hand-crafted MRI features such as lesion characteristics and local intensity patterns.
  • Deep-learning models are criticized as black boxes because their large parameter spaces and nonlinear interactions obscure classification decisions.
  • The study presents a transparent CNN framework using LRP heatmaps to investigate MS diagnostic decisions and support model verification.

2.1. Subjects

The study retrospectively analyzed 76 patients with clinically definite MS and 71 healthy controls from the VIMS study.

  • The cohort included 76 patients with clinically definite MS and 71 healthy controls.Patients were diagnosed according to the 2010 McDonald criteria and had to be aged 18–69 with an MRI scan.

2.2. MRI acquisition and preprocessing

MRI data included high-resolution MPRAGE and FLAIR scans acquired on the same 3 T scanner, followed by registration, lesion processing, and tissue segmentation.

  • MRI acquisition used volumetric high-resolution T1-weighted MPRAGE and FLAIR sequences on the same 3 T scanner.
  • MPRAGE images were registered to MNI space, while FLAIR images were coregistered to MPRAGE and transformed using its normalization parameters.
  • Lesion areas in MRI data were replaced by local average intensities in normal-appearing white matter for lesion-filled analyses.
  • White matter maps were obtained using the SPM 12 tissue segmentation algorithm.

2.3. ADNI data for pre-training

ADNI MRI data provided a larger, multisite pre-training cohort spanning Alzheimer’s disease and cognitively normal subjects.

  • The pre-training dataset comprised 921 MRI scans from 276 ADNI subjects, including Alzheimer’s disease patients and cognitively normal subjects.Subjects contributed one to three time points, and scans were acquired with 1.5 T scanners at multiple sites.

2.4. Classification and visualization analyses

The study trains CNNs to distinguish MS patients from healthy controls and uses LRP to explain voxel-level classification relevance. Transfer learning and atlas-based analyses support comparison of model decisions with established MS imaging features.

  • Classification and visualization analyses: The CNN models classified MS patients versus healthy controls, with variants using ADNI pre-training, MS-only training, and lesion-filled FLAIR data.The study also evaluated whether classification could rely on normal-appearing brain matter.
  • Transfer learning: Because the MS dataset was small, the CNN was pre-trained to separate Alzheimer’s disease patients from healthy controls before fine-tuning on MS data.The ADNI data included 921 MRI scans from 276 subjects, and validation and testing used disjoint subject sets.
  • Interpretability rationale: The study contrasts LRP with sensitivity analysis, which can highlight voxels affecting classification without indicating class-directed relevance.The authors motivate LRP as an approach intended to address this interpretability limitation.
  • Layer-wise relevance propagation: LRP generates subject-specific heatmaps by propagating the classification score through the network rather than using gradients.Relevance is assigned using activations and learned weights, with epsilon set to 0.001 to avoid division by zero.
  • Visualization analyses: The analysis compared individual and average heatmaps and quantified regional relevance using gray- and white-matter brain atlases.The atlases were the Neuromorphometrics atlas and the JHU DTI-based white-matter atlas.

3. Results

The transfer-learned CNN achieved strong holdout classification performance, while LRP maps identified spatially specific relevance in lesions and other brain regions. Relevance patterns differed between MS patients and healthy controls and became more delineated with pre-training.

  • Classification performance: 87.04% balanced accuracy and 96.08% AUC were achieved by the CNN pre-trained on ADNI and fine-tuned on MS data.The model also achieved 93.08% sensitivity and outperformed the other classifiers in AUC.
  • Classification performance: 88.46% balanced accuracy and 94.62% AUC were achieved by the SVM using T2 lesion load, whereas the FLAIR-image SVM reached 66.92% AUC.The MS-only CNN achieved 71.23% balanced accuracy and 85.46% AUC before pre-training.
  • Normal-appearing brain matter: Lesion-filled FLAIR data still supported reasonable CNN performance, indicating that relevance was not restricted to visible lesions.The lesion-filled model achieved 70.15% balanced accuracy and 90.92% AUC.
  • Individual heatmaps: LRP heatmaps for high-scoring MS patients highlighted periventricular lesions and the corpus callosum, while some frontal and other regions carried negative relevance.The maps distinguish regions supporting or opposing the model’s MS classification.
  • Average heatmaps: Average heatmaps showed stronger positive relevance in posterior periventricular white matter for correctly classified MS patients than for healthy controls.Healthy controls generally had a negative whole-brain relevance sum, whereas MS patients had a positive sum.
  • Transfer-learning effects: Transfer learning produced more distinct positive and negative relevance clusters than the CNN trained only on MS data.The randomly initialized model had sparse tiny relevance values, while the pre-trained model showed more delineated clusters.

4. Discussion

The CNN learned MS-relevant information beyond simply detecting hyperintense lesions, while transfer learning improved performance and produced more focused LRP heatmaps. However, the limited sample size constrains robust representation learning and generalization.

  • The CNN assigned 9.71% of total relevance to lesion areas, despite lesion coverage comprising only 0.44% of the training data.
  • LRP showed that the CNN distinguished lesion locations, assigning positive relevance primarily to posterior periventricular regions and strongest relevance to posterior corona radiata, corpus callosum, and thalamic radiation.
  • The model also used normal-appearing white and gray matter, including the thalamus, while lesion-filled data shifted relevance away from typical lesion-bearing regions.
  • Transfer learning across diseases and MRI sequences increased balanced accuracy by about 16 percentage points and yielded more focused heatmaps centered on posterior periventricular lesion areas.
  • The framework's heatmaps can help assess whether CNNs trained on small samples learned meaningful disease-related features or data biases and artifacts.
  • The main limitation was the limited sample size, with n = 147 considered too low for robust representations and generalization to other datasets.

5. Conclusion

The study concludes that CNNs can learn MS-relevant information from typical-sized neuroimaging datasets, especially when combined with transfer learning and LRP. The models used lesions, lesion location, and normal-appearing brain regions as information sources.

  • CNN models learned MS-relevant information from a typical-sized neuroimaging dataset.
  • Pre-training substantially increased prediction performance, including across diseases and MRI sequences, while LRP helped explain individual decisions and assess learned features.
  • The CNNs primarily used hyperintense lesions but also incorporated lesion location and normal-appearing brain areas.

6. Funding

The study received support from the German Research Foundation, the Manfred and Ursula-Müller Stiftung, and Charité – Universitätsmedizin Berlin.

  • Funding came from the German Research Foundation, the Manfred and Ursula-Müller Stiftung, and Charité – Universitätsmedizin Berlin.
Loading 1904.08771v1…