Source-linked AI summary

Predicting Cognitive Decline with Deep Learning of Brain Metabolism and Amyloid Imaging

Hongyoon Choi, Kyong Hwan Jin

arXiv:1704.06033v1cs.CVcs.AIstat.ML

TL;DR

The paper addresses the need to identify MCI patients likely to experience cognitive decline. It develops a deep CNN using baseline FDG and florbetapir PET from AD and normal subjects, then applies the trained network to MCI patients. Prediction accuracy for conversion to AD was 84.2%, exceeded conventional feature-based quantification, and network scores strongly correlated with longitudinal cognitive changes.

  • Problem

    MCI patients show different rates of cognitive decline, creating a need to identify those likely to benefit from treatment.

  • Method

    A deep CNN uses baseline FDG and AV-45 PET from AD and normal subjects to predict cognitive decline in MCI patients without feature extraction or complicated image processing.

  • Results

    84.2% prediction accuracy for conversion to AD in MCI patients outperformed conventional feature-based quantification methods.

  • Takeaways & Limitations

    CNN output scores strongly correlated with longitudinal cognitive measurements, supporting deep learning as a tool for predicting disease outcome from brain images.

  • Takeaways & Limitations

    The two brain hemispheres may cause classification error because of their spatial relationship in the images.

Abstract

from arXiv · show

For effective treatment of Alzheimer disease (AD), it is important to identify subjects who are most likely to exhibit rapid cognitive decline. Herein, we developed a novel framework based on a deep convolutional neural network which can predict future cognitive decline in mild cognitive impairment (MCI) patients using flurodeoxyglucose and florbetapir positron emission tomography (PET). The architecture of the network only relies on baseline PET studies of AD and normal subjects as the training dataset. Feature extraction and complicated image preprocessing including nonlinear warping are unnecessary for our approach. Accuracy of prediction (84.2%) for conversion to AD in MCI patients outperformed conventional feature-based quantification approaches. ROC analyses revealed that performance of CNN-based approach was significantly higher than that of the conventional quantification methods (p < 0.05). Output scores of the network were strongly correlated with the longitudinal change in cognitive measurements. These results show the feasibility of deep learning as a tool for predicting disease outcome using brain images.

Abbreviations

The paper defines abbreviations for imaging, diagnostic, cognitive, and evaluation terms used throughout the study.

  • MCI means mild cognitive impairment, while ADNI refers to the Alzheimer’s Disease Neuroimaging Initiative.
  • FDG denotes 18F-fluorodeoxyglucose, and CNN denotes convolutional neural network.
  • CDR-SB, ADAS-Cog, FAQ, and MMSE name the cognitive and functional assessment measures used in the study.
  • ROC denotes receiver operating characteristic, the analysis used to assess predictive performance.

Introduction

Because MCI is heterogeneous and patients decline at different rates, the study developed a CNN method using FDG and AV-45 PET to predict cognitive decline with minimal image processing.

  • MCI patients show different rates of cognitive decline, and some never convert to AD, motivating prediction of future outcomes.
  • Identifying MCI patients likely to benefit from treatment is described as an important clinical objective.
  • Brain metabolism and amyloid load measured by FDG and AV-45 PET have been investigated as biomarkers of conversion to AD.
  • Visual analysis lacks quantitative and objective data, whereas quantitative analyses commonly require complicated processing.
  • The study developed a deep CNN method applied to FDG and AV-45 PET images for predicting cognitive decline in MCI patients.
  • The automated method was designed to discriminate cognitive-outcome groups with minimized image processing and combine FDG and AV-45 information quantitatively.
  • The resulting CNN-based biomarker was strongly correlated with future cognitive decline.

Methods

The study used ADNI subjects with baseline FDG and AV-45 PET, three-year clinical follow-up for MCI participants, and minimally processed multimodal images for CNN training and testing.

  • MCI patients had baseline FDG and AV-45 PET scans plus three-year follow-up clinical evaluation, enabling converter and nonconverter grouping.
  • The dataset included 182 normal controls, 139 AD patients, and 171 MCI subjects.
  • Cognitive function was assessed with CDR-SB, ADAS-Cog, FAQ, and MMSE, with measurements repeated at one and three years.
  • PET images were co-registered, averaged across time frames, standardized to 1.5 x 1.5 x 1.5 mm voxels, and smoothed by scanner.
  • Preprocessing omitted nonlinear spatial warping, so brain size and shape were not changed; the resulting images were used for CNN training and testing.

Study design

The framework trained a 3D CNN on baseline multimodal PET from AD and normal subjects, then applied it to MCI patients to classify converters and nonconverters and generate ConvScore.

  • The study aimed to develop a deep CNN method that predicts cognitive decline and selects MCI subjects who eventually convert to AD.
  • The CNN was trained on AD and normal-control PET data, then directly applied to MCI converter-versus-nonconverter classification.
  • The network’s final output was defined as ConvScore, indicating how close baseline images are to AD and serving as a potential predictive biomarker.
  • FDG and AV-45 PET were used as two channels in a 3D CNN to exploit diverse multimodal features.
  • The architecture used three-dimensional convolution and ReLU activation, producing a final two-node output for AD and normal control.
  • The study used 10-fold cross-validation and assessed MCI-conversion prediction accuracy and ROC performance.

Network architecture

The network uses three convolutional layers, followed by ReLU, max pooling, and a fully connected softmax layer, to classify PET images. It is trained with supervised learning on AD and NC images and produces ConvScore from the AD output node.

  • Three convolutional layers are followed by ReLU, max pooling, and one fully connected layer.The architecture uses 3-D convolutional filters and reduces feature-map size through pooling and stride operations.
  • Feature maps increase from 2 to 64, 128, and 512 across the network layers.
  • The first convolution receives a 160×160×160×2 input volume, representing spatial dimensions and two input feature maps.
  • Supervised learning assigns output labels of one for AD and two for NC and minimizes a loss over the training dataset using stochastic gradient descent.
  • The network is trained on AD and NC imaging with ten-fold cross-validation, data augmentation by left-right flipping, and 50 training epochs.
  • ConvScore is defined as the quantitative value of the AD output node, indicating how closely an input PET study resembles AD or an MCI converter.

Prediction of cognitive decline in MCI subjects

A CNN trained on AD and NC PET images was applied to MCI subjects to predict conversion to AD. ConvScore was also evaluated against longitudinal cognitive changes.

  • A network trained with AD and NC PET images classified MCI subjects as predicted converters or nonconverters.
  • MCI conversion was predicted when the network’s output probability exceeded 0.5, with sensitivity, specificity, accuracy, and ROC analysis measured.
  • ConvScore was correlated with longitudinal changes in CDR-SB, ADAS-Cog, FAQ, and MMSE measurements.One-year and three-year follow-up measurements were compared with baseline studies.

Feature volume of interests based analysis

The conventional comparison method uses VOI-based quantification of FDG and AV-45 PET after image processing, spatial normalization, segmentation, and regional uptake extraction.

  • VOI-based analyses were performed for both FDG and AV-45 PET images.
  • FDG PET volumes underwent nonlinear spatial normalization before regional uptake was calculated.
  • FDG quantification used uptake from bilateral angular, temporal, and posterior cingulate cortices relative to reference regions.The reference regions included the pons and cerebellar vermis.
  • For AV-45 PET, FreeSurfer-segmented cortical regions provided mean uptake values from frontal, cingulate, parietal, and temporal areas.
  • The overall cortical mean uptake was expressed relative to uptake in the whole cerebellum.

Statistics

CNN and conventional VOI-based approaches were compared using diagnostic and prediction accuracy, ROC AUC, and correlations between ConvScore and longitudinal cognitive measurements.

  • CNN and feature VOI-based approaches were compared for diagnostic and prediction accuracy using McNemar’s nonparametric test.
  • ROC analyses measured AUC for ConvScore and feature VOI-based parameters, with correlated curves compared using DeLong’s nonparametric test.
  • Pearson’s correlation assessed the relationship between ConvScore and longitudinal changes in cognitive measurements.
  • Statistical significance was defined as P-value < 0.05.

Results

The CNN-based approach accurately classified AD versus normal controls and predicted MCI conversion, outperforming feature-VOI analyses. Its ConvScore also correlated significantly with longitudinal cognitive changes and distinguished MCI converters from nonconverters.

  • 492 subjects were included: 139 with AD, 171 with MCI, and 182 normal controls.
  • AD classification: 96.0% accuracy, 93.5% sensitivity, and 97.8% specificity were achieved for AD versus normal-control classification.These measures were significantly higher than those of VOI-based analyses.
  • ROC analysis: AUC was 0.98 for ConvScore, 0.91 for FDG, and 0.84 for AV-45 in AD classification.ConvScore was significantly higher than both feature-VOI analyses, with p < 0.001 for each comparison.
  • ROC analysis: AUC was 0.89 for ConvScore, 0.82 for FDG, and 0.83 for AV-45 in predicting MCI conversion.The comparisons were significant versus FDG at p < 0.01 and versus AV-45 at p < 0.05.
  • Longitudinal cognitive outcomes: ConvScore significantly correlated with changes in CDR-SB, ADAS-Cog, FAQ, and MMSE at both 1-year and 3-year follow-up.At 3 years, correlations were r=0.63 for CDR-SB, r=0.24 for ADAS-Cog, r=0.67 for FAQ, and r=-0.61 for MMSE; changes were steeper than at 1 year.
  • Longitudinal cognitive outcomes: ConvScore was significantly higher in MCI converters than nonconverters and was proposed as a quantitative biomarker for predicting conversion.

Discussion

The deep CNN approach predicted MCI cognitive decline from minimally processed multimodal PET images and produced ConvScore, a quantitative biomarker linked to longitudinal cognition. It outperformed conventional feature-based methods, while augmentation and limited data introduced important scope considerations.

  • The network automatically learned image features and enabled prediction with minimal image processing, without spatial transformation or manual feature extraction.The approach used spatially unnormalized baseline images and automatically performed feature extraction through multiple network layers.
  • 84.2% accuracy differentiated MCI converters from nonconverters, outperforming conventional feature-VOI methods.
  • CNN prediction accuracy was significantly higher than conventional quantification methods and other feature-selection machine-learning algorithms.The comparisons involved different imaging modalities and clinical variables in some prior studies.
  • ConvScore combined metabolism and amyloid information from multimodal PET images and correlated strongly with longitudinal cognitive measurements.Low glucose metabolism and high cortical amyloid deposition at baseline were associated with subsequent cognitive decline.
  • ConvScore may help identify prodromal patients likely to benefit from early intervention in clinical trials.The authors connect its correlation with cognitive change to subject selection for treatment studies.
  • Left-right image flipping improved network performance but could introduce classification errors because the cerebral hemispheres have partly different functions.The authors identify larger imaging cohorts and deeper architectures as future avenues for improvement.
Loading 1704.06033v1…