Source-linked AI summary

A Review of Challenges and Opportunities in Machine Learning for Health

Marzyeh Ghassemi, Tristan Naumann, Peter Schulam, Andrew L. Beam, Irene Y. Chen, Rajesh Ranganath

arXiv:1806.00388v4cs.LGcs.CYstat.ML

TL;DR

Healthcare’s expanding EHR data create opportunities for machine learning, but clinical learning faces challenges from missing observations, unreliable or evolving labels, and leakage. This review provides a practical primer on these technical issues and opportunities for clinically useful systems, while emphasizing collaboration with clinical experts and careful outcome design. Its supported conclusion is that machine learning research in healthcare should address non-stationarity, interpretability, and meaningful representations within clinically relevant settings.

  • Problem

    Healthcare EHRs offer clinically meaningful data, but clinical learning is complicated by missing observations, poorly specified outcomes, and data-generation processes tied to care.

  • Method

    The article reviews healthcare-specific technical challenges and organizes opportunities for automation, clinical support, and expanded clinical capacities.

  • Results

    The review identifies adapting to shifts in data sources and mechanisms, improving interpretability, and learning meaningful representations as key research opportunities.

  • Takeaways & Limitations

    Clinical collaboration and early engagement with experts may support models that are clinically useful and operationally feasible.

  • Takeaways & Limitations

    Clinical actions may be unreliable labels, and information leakage can produce high predictive performance without clinical utility.

Abstract

from arXiv · show

Modern electronic health records (EHRs) provide data to answer clinically meaningful questions. The growing data in EHRs makes healthcare ripe for the use of machine learning. However, learning in a clinical setting presents unique challenges that complicate the use of common machine learning methodologies. For example, diseases in EHRs are poorly labeled, conditions can encompass multiple underlying endotypes, and healthy individuals are underrepresented. This article serves as a primer to illuminate these challenges and highlights opportunities for members of the machine learning community to contribute to healthcare.

Introduction

Healthcare depends fundamentally on clinical data, making machine learning’s ability to extract information especially relevant. Yet applying machine learning directly to healthcare remains difficult because clinical data are generated primarily to support care and involve distinct technical challenges.

  • Clinical data support patient-specific treatment decisions and healthcare improvement.
  • Machine learning’s data-driven advances and healthcare’s reliance on data motivate machine learning research for healthcare.
  • Healthcare machine learning has advanced in diagnosis, pathology, autism subtyping, and large-scale phenotyping.
  • Direct application remains fraught because healthcare data are generated and managed primarily to support care rather than machine learning.
  • This review focuses on EHRs, emphasizing broad opportunities and careful considerations beyond narrower prior reviews.
  • The article covers technical challenges, organizes clinical opportunities into automation, support, and capacity expansion, and highlights shifts, interpretability, and representations.

Unique Technical Challenges in Healthcare Tasks

Healthcare machine learning must address causality, missingness, and outcome definition across modeling frameworks and learning targets. Causal questions are particularly difficult because observational healthcare data reflect unknown action-selection policies and can produce misleading relationships.

  • Causality, missingness, and outcome definition require careful consideration across supervised, unsupervised, classification, and regression settings.
  • Causal healthcare questions ask what would happen under treatment interventions, requiring formal causal models beyond classical machine learning.
  • Observational data complicate causal learning because actions may reflect an unknown agent policy.
  • Simpson’s paradox can reverse observed relationships when additional variables, such as level of care, enter the model.
  • Causal models can evaluate treatments and help avoid harmful predictions driven by treatment policies in training data.

Models in Health Must Consider Missingness

Healthcare data commonly contain missing observations, and measurement processes depend on prior clinical observations. Models must therefore account for missingness mechanisms and potential social or practice-related biases to avoid biased or unfair results.

  • Models in Health Must Consider Missingness: Many healthcare observations are missing even when datasets include all important variables.
  • Models in Health Must Consider Missingness: Missingness mechanisms are classified as MCAR, MAR, or MNAR according to how measurement probability relates to observed or unobserved values.
  • Models in Health Must Consider Missingness: Missingness sources should be examined before deployment because the presence of a laboratory measurement can itself convey patient-state information.
  • Models in Health Must Consider Missingness: Including missingness indicators can improve prediction, whereas ignoring measurement processes can misestimate feature importance and reduce robustness to practice changes.
  • Models in Health Must Consider Missingness: Missingness may reflect differences in access, practice, or recording, producing unfair performance across populations.

Make Careful Choices in Defining Outcomes

Reliable outcomes are foundational for defining healthcare machine learning tasks, but EHR labels may be heterogeneous, clinically evolving, or contaminated by information leakage. Outcome design must therefore integrate multiple sources, assess clinical relevance, and preserve the intended prediction setting.

  • Make Careful Choices in Defining Outcomes: Reliable outcomes create gold-standard labels for supervised learning and well-defined cohorts for clustering.
  • Make Careful Choices in Defining Outcomes: EHR labels should combine heterogeneous sources because structured diagnostic codes may be imprecise or unreliable.
  • Make Careful Choices in Defining Outcomes: Phenotyping pools different data types to obtain more reliable labels, including information extracted from clinical notes.
  • Make Careful Choices in Defining Outcomes: Predictive performance depends on the criteria underlying medical definitions, which evolve with scientific understanding.
  • Make Careful Choices in Defining Outcomes: Clinical actions may be poor labels when observed treatments differ from optimal care or reflect choices other clinicians would make.
  • Make Careful Choices in Defining Outcomes: Label leakage can yield high predictive performance without clinical utility when features contain information about the target outcome.

Addressing a Hierarchy of Healthcare Opportunities

The paper frames healthcare opportunities around automating clinical tasks, providing clinical support, and expanding clinical capacities, while emphasizing early clinical-stakeholder involvement.

  • Healthcare opportunities are organized into automating clinical tasks, providing clinical support, and expanding clinical capacities.
  • Clinical stakeholders should be engaged early because deployment details can change a technical solution’s intent.

Clinical Task Automation: Automating clinical tasks during diagnosis and treatment

Clinical task automation targets well-defined work currently performed by clinicians, including medical image evaluation and routine process management. These applications can be evaluated against existing standards while optimizing rather than replacing clinical staff.

  • Clinical Task Automation: Well-defined clinician tasks with known inputs and outputs are low-hanging opportunities requiring relatively little domain adaptation and investment.Performance can be measured against existing standards.
  • Clinical Task Automation: Automation should optimize clinical workflows rather than replace staff, whose roles may evolve as these techniques improve.
  • Automating clinical tasks during diagnosis and treatment: Medical image evaluation is a natural machine-learning opportunity because clinicians map fixed image inputs to outputs such as diagnoses.Reported applications include diabetic retinopathy, skin-lesion, lymph-node-metastasis, and hip-fracture detection.
  • Automating clinical tasks during diagnosis and treatment: Routine process automation can reduce staff burden by algorithmically prioritizing emergency triage and summarizing disparate medical-record information.

Clinical Support and Augmentation: Optimizing clinical decision and practice support

Clinical support and augmentation address information loss, treatment variation, fragmented records, limited evidence, and opportunities for continuous monitoring and individualized treatment. These applications require attention to clinical value, collaboration, and healthcare-specific data challenges.

  • Clinical Support and Augmentation: Clinical support requires identifying pain points with staff and jointly specifying input data, output targets, and evaluation functions.Support addresses work constrained by limited time and resources, which can lead to information loss and errors.
  • Standardizing clinical processes: Standardized order sets and default dosages can support more consistent medication decisions, although static protocols are easier to automate than changing clinical guidelines.
  • Integrating fragmented records: Machine learning can integrate fragmented records to identify patterns such as domestic abuse up to 30 months before healthcare-system recognition.
  • Expanding Clinical Capacities: Digitized records create opportunities for new healthcare capacities, whose impact should be assessed through both innovation and clinical value with substantial clinical collaboration.
  • Expanding the coverage of evidence: Only 10–20% of treatments were estimated to be backed by a randomized controlled trial, motivating natural experiments to investigate more clinical questions.Trial populations may not represent heterogeneous patients, and individualized care can produce unique treatment pathways.
  • Moving towards continuous behavioral monitoring: Wearable data can support continuous, non-invasive classifications and alerts, but label leakage, soft labels, confounding, and missingness require careful consideration.For chronic conditions, apparent early detection may identify an existing treatment because patients are often already treated.
  • Precision medicine for early individualized treatment: Longitudinal data can support individualized treatment effects for syndromes with multiple possible causes, including acute kidney injury.The approach relates to N=1 crossover studies and uses repeated measurements to distinguish causes over time.

Opportunities for New Research in Machine Learning

The paper identifies research opportunities in machine learning for healthcare that address data non-stationarity, model interpretability, and appropriate representation discovery.

  • Opportunities for New Research in Machine Learning: Promising research directions include handling data non-stationarity, improving model interpretability, and discovering appropriate representations.

Accommodating Data and Practice Non-stationarity in Learning and Deployment

Healthcare models must account for non-stationarity within institutions, across data sources, and through feedback loops. Changing populations, practices, search behaviors, and data infrastructures can degrade performance or propagate bias.

  • Shift over time: Changing patient populations and treatment procedures can degrade predictive performance as the statistical properties of the target change.Models should become robust to these changes or acknowledge mis-calibration for the new population.
  • Shift over time: Google Flu Trends persistently overestimated flu after user search behaviors shifted, motivating models that continuously update.
  • Shift over sources: Models learned at one hospital may not generalize to another because practices, populations, equipment, and EHR feature mappings differ.Infrastructure for testing across multiple sites and normalization across data-collection practices remains an opportunity.
  • Feedback loops: Models trained on clinical practice can amplify healthcare biases when deployed predictions influence future training data.The paper connects this risk to feedback loops observed in predictive policing and points to algorithmic fairness research.

Creating interpretable models and recommendations

Healthcare machine learning should move beyond opaque predictions toward systems that clinicians can interpret, justify, and use collaboratively. Key opportunities include clinically meaningful representations, multimodal integration, and models that accommodate healthcare data’s temporal and structural complexity.

  • Creating interpretable models and recommendations: Black-box models create deployment challenges because clinicians must understand and justify treatment deviations under clinical and legal requirements.The paper notes that models cannot be deployed in clinical settings at low cost.
  • Creating interpretable models and recommendations: Interpretability can be pursued through feature minimization, regularization, model classes with post-hoc analyses, or posterior distributions over decision lists.Decision lists can present relative risks in a form clinicians can use.
  • Moving from interpretation to justification: Justifiability should trace the predictive path through individual outputs, training data, or learning algorithms rather than merely highlight features.Influence functions and locally interpretable results are presented as examples, with justification also relevant to security against adversarial attacks.
  • Adding interaction to machine learning and evaluation: Collaborative systems can combine the complementary strengths of physicians and learning systems while accounting for clinicians’ caregiving and empathy roles.
  • Learning meaningful representations for the domain: Healthcare representation learning should integrate multiple sources and modalities while capturing domain-appropriate structure, hierarchy, similarity, and conditional relationships.These representations should support diverse tasks, multiple valid outputs, domain knowledge, and zero-shot learning in unseen categories.

Conclusion

The paper offers a practical guide for researchers entering healthcare machine learning and encourages early collaboration with clinical experts. Such collaboration may support models that are clinically useful and operationally feasible.

  • Conclusion: Researchers should engage clinical experts early when identifying and tackling important healthcare machine-learning problems.The paper presents collaboration as a route that may lead to clinically useful and operationally feasible models.
Loading 1806.00388v4…