Source-linked AI summary

Artificial Intelligence, speech and language processing approaches to monitoring Alzheimer's Disease: a systematic review

Sofia de la Fuente Garcia, Craig Ritchie, Saturnino Luz

arXiv:2010.06047v1cs.AIcs.CLeess.AS

TL;DR

Speech and language technology is promising for monitoring Alzheimer’s disease, but heterogeneous, non-standardised studies limit consensus and clinical translation. This systematic review synthesises the field and finds generally high reported performance, while noting that clinical implementation remains limited.

  • Problem

    Heterogeneous, small, inconsistent, and non-standardised studies limit consensus and translation of speech and language technology for Alzheimer’s disease.

  • Method

    The paper presents a systematic review of interactive artificial intelligence, natural language processing, and speech-technology studies targeting Alzheimer’s detection and progression monitoring.

  • Results

    Almost all reviewed studies report relatively high performance, with speech and language technology at least equally discriminative as neuropsychological assessment methods.

  • Takeaways & Limitations

    Interactive artificial intelligence has potential to gradually introduce speech- and language-based changes into Alzheimer’s clinical practice.

  • Takeaways & Limitations

    Automation remains challenging because transcription, feature generation, and feature reduction must balance informative speech content, delivery, and structure.

Abstract

from arXiv · show

Language is a valuable source of clinical information in Alzheimer's Disease, as it declines concurrently with neurodegeneration. Consequently, speech and language data have been extensively studied in connection with its diagnosis. This paper summarises current findings on the use of artificial intelligence, speech and language processing to predict cognitive decline in the context of Alzheimer's Disease, detailing current research procedures, highlighting their limitations and suggesting strategies to address them. We conducted a systematic review of original research between 2000 and 2019, registered in PROSPERO (reference CRD42018116606). An interdisciplinary search covered six databases on engineering (ACM and IEEE), psychology (PsycINFO), medicine (PubMed and Embase) and Web of Science. Bibliographies of relevant papers were screened until December 2019. From 3,654 search results 51 articles were selected against the eligibility criteria. Four tables summarise their findings: study details (aim, population, interventions, comparisons, methods and outcomes), data details (size, type, modalities, annotation, balance, availability and language of study), methodology (pre-processing, feature generation, machine learning, evaluation and results) and clinical applicability (research implications, clinical potential, risk of bias and strengths/limitations). While promising results are reported across nearly all 51 studies, very few have been implemented in clinical research or practice. We concluded that the main limitations of the field are poor standardisation, limited comparability of results, and a degree of disconnect between study aims and clinical applications. Attempts to close these gaps should support translation of future research into clinical practice.

Introduction … Risk of bias in individual studies

This systematic review synthesized AI-based speech and language research for monitoring Alzheimer’s disease, using broad eligibility and search procedures to address a heterogeneous evidence base. It extracted methodological and clinical-applicability information while assessing risks such as overfitting and poor reporting.

  • Introduction: The introduction frames Alzheimer’s disease as a progressive neurodegenerative disorder in which early language impairment makes speech and language valuable clinical information sources.Speech offers non-invasive, trackable data, and speech problems may occur across disease stages.
  • Introduction: The review aimed to summarize AI approaches for speech-based Alzheimer’s monitoring, including their aims, findings, methods, and readiness for clinical evaluation.The stated objectives were to support future research, guidelines, and eventual clinical implementation.
  • Methods: The review protocol was registered in PROSPERO under reference CRD42018116606 before eligibility, searching, selection, extraction, synthesis, and bias-assessment procedures were conducted.The protocol covered the review’s planned methodological stages.
  • Elegibility criteria: Eligible studies used automatic machine learning, computational linguistics, speech technology, or related AI methods for Alzheimer’s screening, detection, or prediction.The review included subjective cognitive impairment, mild cognitive impairment, Alzheimer’s disease, and related dementia terminology, while excluding traditional-statistics-only and language-unrelated neuroimaging studies.
  • Elegibility criteria: The review covered peer-reviewed original research published from 2000 through 2019, excluding redundant publications from the same research group.Conference abstracts and systematic reviews were excluded, and overlapping records were resolved by selecting the most comprehensive and up-to-date article.
  • Information Sources; Search Strategy: Searches conducted between October and December 2019 covered ACM, Embase, IEEE, PsycINFO, PubMed, and Web of Science, with additional forward citation tracking and reference screening.AI terms were omitted from the search queries because preliminary testing suggested they constrained retrieval too much.
  • Study records selection: Two reviewers independently screened records in title-and-abstract and full-text phases, resolving disagreements through discussion or third-author adjudication.Relevant emerging titles were added during screening, and redundant reports were handled during full-text assessment.
  • Data collection process; Data items (extraction tool): A purpose-built extraction tool used four tables covering study characteristics, data details, methodology, and clinical applicability, with independent extraction and third-author mediation when needed.The tables included population, modalities, annotation, preprocessing, features, machine-learning tasks, evaluation, results, clinical potential, risk of bias, and limitations.

Data synthesis … Population (table 5: SPICMO)

The review narratively synthesizes diagnostic and prognostic evidence because meta-analysis was beyond scope, while organizing study aims, populations, data resources, methods, and clinical implications. Across the included literature, most studies classify cognitive impairment using heterogeneous populations and datasets, limiting comparability and clinical translation.

  • Data synthesis; Confidence in cumulative evidence: The review uses narrative synthesis and comparative outcome reporting where possible, rather than conducting a meta-analysis or evaluating treatment implementation.The synthesis follows the tables’ structure and assesses diagnostic and prognostic tools, not intervention confidence.
  • Background on AI, Cognitive tests and Databases; AI, machine learning, and speech technologies; Data extraction: The review defines terminology, summarizes AI and machine-learning concepts, and introduces common performance measures, cognitive assessments, and databases to support table interpretation.Reported measures include accuracy, sensitivity, specificity, positive predictive value, area under the receiver operating characteristic curve, and F scores.
  • Cognitive tests; Population (table 5: SPICMO): Cognitive assessments support clinical group assignment and may provide classifier baselines, but MMSE can show ceiling effects when assessing pre-clinical Alzheimer’s disease.Common diagnostic batteries include MMSE, MoCA, HDS-R, and Clinical Dementia Rating, while speech tasks include semantic fluency and Cookie Theft descriptions.
  • Databases: The Pitt Corpus is the most commonly used monologue dataset, whereas the Carolina Conversations Collection is the only available dialogue database described.Other resources include Hungarian spontaneous speech, the Gothenburg MCI database, and the potentially unavailable IVA dataset.
  • Databases: The IVA dataset comprises structured interviews recorded with an Intelligent Virtual Agent, but its potential availability is unknown.This limits its usefulness for reproducibility and broader dataset access.
  • Results: 3,654 records were identified, 306 duplicates were removed, and 3,128 papers were excluded during title-and-abstract screening before 220 entered the next phase.The search combined six digital databases with bibliography, citation, and research-portal screening.
  • Discussion; Study aim and design (table 5: SPICMO): 41 out of 51 papers address impairment presence or absence, while seven examine three or four disease stages, two examine longitudinal change, and one describes discourse patterns.The review suggests that future work could use longitudinal data to build models of cognitive change.
  • Study aim and design (table 5: SPICMO): Most studies distinguish healthy participants from cognitively impaired groups using binary classifiers, while only one predicts MCI-to-AD conversion and one predicts progression from healthy cognition.Longitudinal data were used in the two progression studies, and one study instead examined discourse patterns.

Interventions (table 5: SPICMO) … Pre-processing (table 7: Methodology)

The reviewed studies mainly used health-assessment-linked, constrained speech tasks and heterogeneous diagnostic groups, datasets, modalities, languages, and preprocessing procedures. Limited reporting of class balance, data availability, and preprocessing reduces comparability, reliability, and clinical relevance.

  • Interventions (table 5: SPICMO): Speech generation tasks almost invariably followed general clinical or cognitive health assessments, but some studies did not specify criteria for assigning comparison groups.Tasks included verbal fluency, story recall, picture description, consultations, and cognitive examinations such as MMSE.
  • Interventions (table 5: SPICMO): Most interventions were constrained laboratory tasks that improve standardisation and cognitive-load control, whereas spontaneous and conversational speech enable naturalistic longitudinal capture.Naturalistic data may mitigate effects such as an “off day” or poor sleep during controlled cross-sectional testing.
  • Comparison groups (table 5: SPICMO): Most studies compared diagnosed AD or related-condition cohorts with healthy controls, providing little insight into pre-clinical disease stages.The review homogenised diagnostic nomenclature into HC, SCI, MCI, AD, and CI, while some studies used alternative or subdivided categories.
  • Outcomes of interest (table 5: SPICMO): Binary-classification performance varied widely with data, recording conditions, and modelling variables, while Mirzaei et al. reported 62% for 3-way HC, MCI, AD classification.Class imbalance can make accuracy misleading, motivating reporting of sensitivity, specificity, F scores, unweighted average recall, contingency tables, and ROC curves.
  • Size of dataset or subset (table 6: Data Details): Only 6 studies had at least 100 participants or speech samples in every experimental group, and participant counts were not always distinguished from speech-sample counts.The Pitt Corpus was the largest dataset; one study reported 473 speech samples from 264 participants.
  • Other modalities; Data annotation; Data balance; Data availability; Language (table 6: Data Details): Datasets varied in modalities, diagnostic labels, balance, availability, and language: 39% (20) of studies presented class balance, while 77% (39) failed to report data availability.English accounted for 41% of studies, alongside many other languages; inconsistent terminology and limited sharing hinder clinical relevance and worldwide implementation.
  • Pre-processing (table 7: Methodology): Preprocessing was poorly documented: text commonly involved manual or ASR transcription, acoustic processing mainly used VAD when reported, and complete accounts were uncommon.The review recommends more thorough reporting because preprocessing is crucial to reliability and replicability.

Feature generation (table 7: Methodology) · ML task/method (table 7: Methodology) · Evaluation techniques (table 7: Methodology)

The reviewed studies used heterogeneous text-based, acoustic, and occasionally multimodal features, mostly with supervised conventional classifiers and limited cognitive-score prediction. Evaluation commonly relied on cross-validation, but inconsistent feature selection, missing baselines, and inadequate metrics reduced comparability and clinical interpretability.

  • Feature generation (table 7: Methodology): Most studies used a single data type, although some combined text-based and acoustic features or added modalities such as images and gait measurements.The review describes most published research as specific to one type of data or another.
  • Feature generation (table 7: Methodology): Text-based features commonly included lexical and syntactic indices, while acoustic features emphasized prosodic timing, pauses, and spectral measures such as MFCCs.Some studies also extracted dialogue features including turn-taking, vocalisations, speech rate, and dysfluencies.
  • Feature generation (table 7: Methodology): 30% of studies did not report feature selection, while another 30% of those reporting it used statistical-index filter approaches.Other approaches included wrappers, recursive feature elimination, information gain, PCA, greedy search, and cross-validation.
  • Feature generation (table 7: Methodology): Ad hoc feature procedures created striking heterogeneity that limited comparability, whereas standardized sets such as ComPare, eGeMAPS, and emobase offered documented, replicable alternatives.The review notes that eGeMAPS was developed for affective speech and underlying physiological processes.
  • ML task/method (table 7: Methodology): Most studies used supervised classification; conventional SVM, NB, RF, and k-NN classifiers predominated, while neural methods were uncommon, likely because datasets were small.One study used cluster analysis to investigate distinctive discourse patterns.
  • Evaluation techniques (table 7: Methodology): 43% of studies lacked a baseline, and accuracy was the most common metric despite being unsuitable for imbalanced datasets.AUC and EER, which summarize false-alarm and false-negative rates, appeared in less than half of reviewed studies.
  • Evaluation techniques (table 7: Methodology): Cross-validation was reported in all but five papers and was considered appropriate for small datasets, but procedures varied and hyper-parameter optimization was often insufficiently reported.Hold-out evaluation was used by two papers without reported cross-validation, and one paper reported neither procedure.

Results overview (table 7: Methodology)

Classifier performance ranged from 50% or lower to over 90% accuracy, but results are difficult to summarise and require caution because of methodological biases. Further research is needed to explain whether performance differences reflect data characteristics or classifier choice.

  • Performance range: 50% or lower to over 90% accuracy was observed across evaluated classifiers.Performance varied with the metric, data type, and classification algorithm.
  • Methodological limitations: Performance figures require caution because dataset size, imbalances, and non-standardised ad hoc feature generation may introduce bias.These factors limit interpretation and comparability of results.
  • Research needs: Further research should determine why some classifiers perform worse than random while others approach perfect performance.Differences may reflect easier data, such as better quality or clearer diagnoses, or classifier-specific effects.

Research implications (table 8: Clinical applicability)

The review identifies methodological, longitudinal, multimodal, and implementation-related novelty, while emphasizing that replicability and especially generalisability remain essential for translating approaches into clinical practice. Most studies have low generalisability, commonly because of content and transcription dependence.

  • Novelty: Novelty includes ensemble and cascaded classifiers, custom ASR systems, longitudinal data, and multimodal inputs such as MRI, eye-tracking, and gait.Most studies use standard classifiers, while only a few combine different data sources or implement human-robot interaction and telephone-based systems.
  • Novelty: Multimodal data are identified as having the greatest clinical potential because Alzheimer’s Disease may require comprehensive models for successful screening.Only a few studies combine different sources of data, including MRI, eye-tracking, and gait.
  • Replicability: Replicability must be confirmed before clinical translation, although low replicability is not considered a key field-wide problem because all such papers are conference proceedings.The review nevertheless states that preprocessing and feature-generation descriptions must be improved.
  • Generalisability: Transcription-free acoustic methods may increase generalisability by reducing ASR and manual-transcription constraints, supporting language independence and user privacy.The review suggests that such methods could facilitate use with non-English languages and help protect speech-content privacy.
  • Generalisability: 20 studies had low generalisability, 17 moderate, and 14 high, with content dependence followed by ASR or other transcription dependence as the most common limitations.Content dependence makes approaches difficult to apply to other tasks or datasets, especially when they rely heavily on word content such as n-grams.

Clinical potential (table 8: Clinical applicability)

Clinical translation remains limited despite substantial potential: most studies lack external validation or real-world clinical integration, and remote use and interpretability are rarely demonstrated. The review identifies real clinical deployment, disease progression and risk prediction, broader language coverage, and practical screening as key needs.

  • External validation: 84% of reviewed studies present neither external validation procedures nor a system design involving external validation.Data are generally collected detached from clinical practice and analysed later for reporting.
  • Potential application: 78% of reviewed papers present a method potentially applicable as a diagnosis support system for MCI or AD.Other studies address disease progression through SCI participants, within-subject change, or discrimination among HC, MCI, and AD stages.
  • Global health: Research in languages other than English commonly includes acoustic features, while smaller processing units such as phonemes tend to be more generalisable across languages.The review notes that most research is conducted in English, reported as 41%, and links non-English work to broader applicability.
  • Remote application: 67% of studies do not mention remote use, while only 25% suggest it and four studies, reported as 2%, actually experiment with remote applications.These include multimodal human-robot interaction, infrastructure-free screening, and telephone-based approaches.
  • Model interpretability: Only four reviewed papers explicitly mention interpretability or model interpretation.Other studies use inherently interpretable methods, including linear, logistic, generalised linear or additive models, and decision trees.
  • Research needs: The field needs more attempts to use models in real clinical practice and greater focus on disease progression and risk prediction.The review notes that research papers have suggested clinical implementation for twenty years, but few published studies have realised it.

Risk of bias (table 8: Clinical applicability)

The reviewed studies show substantial risks of bias from feature imbalance, inappropriate metrics, insufficient contextualization, possible overfitting, and limited or unbalanced datasets. Larger, balanced datasets and stricter validation are needed to improve methodological rigour.

  • Feature balance: Only 13 studies (25%) balanced class, age, gender, and education features, while five balanced class but not the other features.Strategies addressing class imbalance included subsampling, stratified cross-validation, and careful evaluation methods.
  • Suitable metrics: 18 studies (35%) using imbalanced datasets reported accuracy only, although accuracy is not robust for such datasets.Accuracy should be avoided or complemented with other performance measures when classes are imbalanced.
  • Contextualized results: Only 61% of reviewed studies provided contextualized results through quantitative comparisons with related work or a baseline.Contextualized results enable comparison with related studies and, ideally, a baseline.
  • Overfitting: Although 78% of studies reported cross-validation, 90% did not report a held-out set, indicating a high risk of overfitting.Cross-validation should tune hyperparameters, whereas held-out data should test models on strictly unseen data.
  • Sample size: 13 studies used datasets with ds ⩽50, 24 used ds ⩽100, and 14 used ds > 100; some medium and larger datasets also attempted 3-way or 4-way classification.Group sizes were further reduced in seven medium-sized and one larger study attempting multiway classification, underscoring the need for larger balanced datasets.

Strengths/Limitations (table 8: Clinical applicability)

The review identifies five qualities intended to improve the clinical translatability of AI research for Alzheimer’s disease. Only two studies met all five criteria, while spontaneous conversational speech and greater automation remain priorities.

  • Speech data: Spontaneous, ideally conversational speech is preferred because it is more representative of real-life language and suitable for continuous, longitudinal collection.The authors acknowledge trade-offs between naturalness and standardisation, and between conversational realism and dialogue-related confounds.
  • Transcription: 35% of reviewed papers use a transcription-free approach, whereas the rest rely on manual or ASR transcriptions.The authors consider transcription-free methods more relevant to clinical application because ASR remains constrained and manual transcription is burdensome.
  • Overall applicability: Only two studies meet all five criteria, indicating that the field must further pursue natural speech and automation despite trade-offs involving clinically informative content.The authors specifically identify transcription, feature generation, and feature-set reduction as challenging automation gaps.

Overall Conclusions

AI, speech, and language processing for Alzheimer’s disease detection and progression monitoring is promising, but translation into clinical practice remains limited. The review identifies poor methodological rigour and standardisation as barriers and recommends more clinically feasible, standardised, and ethically supported research.

  • Overall Conclusions: The review found AI, speech, and language processing to be a promising field for extracting digital biomarkers for Alzheimer’s disease detection and progression monitoring.It describes this as the first systematic review of interactive AI methods applied to these purposes.
  • Overall Conclusions: Despite nearly 20 years of research, no reviewed study had achieved actual translation into clinical practice.The authors speculate that limited interdisciplinary cooperation between AI/ML experts and clinicians may contribute to slow uptake.
  • Overall Conclusions: Many studies reduced features outside cross-validation, while barely any reported hold-out procedures or experiments on entirely separate datasets.The authors identify separate-dataset validation as the ideal scenario for robust model validation.
  • Overall Conclusions: Small, variably qualified datasets limit rigorous evaluation and the creation of adequate subsets while preserving experimental-group size and integrity.The review argues that standards for data and methodology could increase study-evaluation strictness.
  • Overall Conclusions: Future research should standardise feature sets, prioritise feasible remote monitoring through passively collected natural conversations, and combine digital biomarkers with established biomarkers.The review specifically recommends eGeMAPS as one standardised feature set and highlights ethical and regulatory hurdles surrounding recorded personal data.

Supplementary material · Keys to table interpretability

The supplementary material defines conventions for interpreting the review tables, including publication selection, rounding, feature terminology, and balance notation. It distinguishes dataset-wide, within-class, and between-class feature balance while specifying how gender, age, education, and class balance are reported.

  • Supplementary material: Only the later publication was included when a conference paper was subsequently extended and published elsewhere.
  • Keys to table interpretability: Percentages were rounded to the nearest decimal place, and education was expressed in average years like age unless otherwise specified.
  • Keys to table interpretability: The review classified acoustic and paralinguistic features together as “acoustic” and omitted unspecified classifier or model types from table designations.
  • Keys to table interpretability: Dataset Feature Balance indicates whether a feature is evenly distributed across the dataset as a whole, such as gender across all participants.
  • Keys to table interpretability: Within-Classes Feature Balance indicates whether a feature is evenly distributed within each experimental class, using separate balance labels for groups such as HC and AD.
  • Keys to table interpretability: Between-Classes Feature Balance indicates whether a feature is evenly distributed across experimental classes, with age and education assessed through group averages.
  • Keys to table interpretability: Age and education are reported as BCB: A, E or BCB: no-A, no-E, while No-BCB means no features are balanced and unspecified BCB means all three are balanced.
  • Keys to table interpretability: Gender balance is denoted by GB, WCGB, or BCB: G, with corresponding no-balance forms; class balance is denoted by CB or no-CB.

Abbreviations and Acronyms

The review standardizes diagnostic labels by mapping cognitively normal groups to HC, dementia groups to AD, pre-clinical memory loss to SCI, and unspecified symptomatic impairment to CI.

  • Diagnostic abbreviations: HC denotes cognitively normal or normal elderly participants, AD denotes dementia groups, SCI denotes pre-clinical Subjective Memory Loss, and CI denotes symptomatic groups without an official diagnosis term.These mappings are defined specifically for this review.

Methods

The Methods section defines abbreviations for speech and language processing, data representation, machine learning, and evaluation terminology used in the review.

  • Terminology: Speech and language processing terms include ASR, or Automatic Speech Recognition, and ADR, or Active Data Representation.The glossary also defines CNN as Convolutional Neural Network and CV as Cross-validation.
  • Terminology: Data representation and statistical learning terms include DR, DT, GC, GNB, Information Gain, k-Nearest Neighbour, and Linear Discriminant Analysis.The passage expands DR as Data representation, DT as Decision Trees, and GNB as Gaussian Naive Bayes.
  • Terminology: Feature selection, regression, language modeling, and neural-network terms include LASSO, Logistic Regression, Latent Semantic Analysis, LSTM, and Multi-layer Perceptron.The glossary expands LASSO as Least Absolute Shrinkage and Selection Operator, LSTM as Long-short-term-memory, and RNN as Recurrent Neural Network.
  • Terminology: Evaluation terminology includes Accuracy, Area Under the Curve, Classification Error Rate, Equal Error Rate, Precision, Recall, Specificity, and Sensitivity.The glossary also defines receiving operating characteristic, True Positives, False Alarms, and Unweighted Average Recall.

2. SPICMO (PICOS) table

The SPICMO table reorganises the clinical-review framework around study aim, population, intervention, comparisons, methodology and outcomes, adding methods to the conventional PICOS structure. It standardises comparison-group terminology and summarises study designs, approaches and reported detection performance.

  • Study components: The table records interventions such as cognitive or clinical assessments, recorded speech tasks and written tasks, alongside study aims including target-group detection or discrimination between impairment stages.Design distinctions include text versus speech and narrative versus monologue.
  • Table structure: The comparison framework standardises heterogeneous labels into Healthy Controls, Subjective Cognitive Impairment, Mild Cognitive Impairment, Alzheimer’s Disease and unspecified Cognitive Impairment.Examples equated by the review include normal controls with HC, healthy elderly with HC, subjective memory complaints with SCI, and dementia with AD.
  • Study components: Methodology entries cover acoustic or natural-language feature generation, feature selection or extraction, and the machine-learning task used.Feature-reduction methods include filtering, wrapping, PCA, LSA and ADR when reported.
  • Table structure: SPICMO shifts the conventional PICOS column order so that study aim precedes population, intervention, comparisons, methodology and outcomes.The review adds a dedicated methods column because methodology is considered essential to the review.

Clinical applicability

Clinical applicability was assessed through external validation, potential applications, generalisability, global-health and remote-use considerations, alongside replicability and methodological limitations. Reviewed approaches included diagnosis support and disease-progression monitoring, but assessment also identified concerns involving class balance, metric choice, contextualisation, overfitting and transcription requirements.

  • Replicability: Replicability was rated low, partial or full according to procedural detail and the availability of data or data identifiers.Low replicability meant that both data were unavailable and the method description was incomplete or unsatisfactory.
  • Clinical potential: Clinical potential was classified by whether external validation had been attempted or realistic clinical testing was embedded in the design.Potential applications included early screening, disease-progression monitoring and support for diagnosing MCI and AD.
  • Methodological limitations: Clinical assessment examined class balance, suitable metrics, contextualised results and overfitting, including whether cross-validation and/or hold-out procedures were used.Reported examples included accuracy-only metrics, cross-validation without a hold-out set, unclear hold-out status, and one case reporting accuracy and AUC.
  • Methodological limitations: Transcription-free analysis was considered clinically relevant because manual transcription is time-consuming and automatic speech recognition performs poorly on impaired speech and requires language-specific training.Text analysis usually requires transcripts, adding an extra methodological step.
  • Potential applications: Diagnosis support was the most repeatedly identified potential application, while some acoustic-feature studies targeted disease progression.The table also recorded multilingual or language-specific settings and suggested remote applicability for some approaches.
Loading 2010.06047v1…