Source-linked AI summary

A Comprehensive Survey of Foundation Models in Medicine

Wasif Khan, Seowung Leem, Kyle B. See, Joshua K. Wong, Shaoting Zhang, Ruogu Fang

arXiv:2406.10729v3cs.LGcs.AIcs.CV

TL;DR

Healthcare foundation-model research lacks a comprehensive taxonomy spanning its diverse model families, modalities, and applications. This survey synthesizes FM evolution, architectures, specialized medical models, healthcare applications, challenges, and open research issues, finding that domain-specific models often outperform general-purpose counterparts while deployment remains constrained by data, accuracy, and interpretability challenges.

  • Problem

    Existing surveys do not comprehensively cover healthcare foundation models across language, vision, protein, audio, graph, and other medical modalities.

  • Method

    The survey synthesizes FM evolution, learning strategies, flagship models, healthcare applications, taxonomies, opportunities, challenges, and open research questions.

  • Results

    Specialized healthcare models tend to outperform general-purpose counterparts across clinical NLP and medical image segmentation, while medical FMs perform well across tasks including biomedical QA and protein structure prediction.

  • Takeaways & Limitations

    Responsible healthcare FM deployment requires attention to domain-specific data, accuracy, interpretability, and the practical constraints of medical use.

Abstract

from arXiv · show

Foundation models (FMs) are large-scale deep learning models trained on massive datasets, often using self-supervised learning techniques. These models serve as a versatile base for a wide range of downstream tasks, including those in medicine and healthcare. FMs have demonstrated remarkable success across multiple healthcare domains. However, existing surveys in this field do not comprehensively cover all areas where FMs have made significant strides. In this survey, we present a comprehensive review of FMs in medicine, focusing on their evolution, learning strategies, flagship models, applications, and associated challenges. We examine how prominent FMs, such as the BERT and GPT families, are transforming various aspects of healthcare, including clinical large language models, medical image analysis, and omics research. Additionally, we provide a detailed taxonomy of FM-enabled healthcare applications, spanning clinical natural language processing, medical computer vision, graph learning, and other biology- and omics- related tasks. Despite the transformative potentials of FMs, they also pose unique challenges. This survey delves into these challenges and highlights open research questions and lessons learned to guide researchers and practitioners. Our goal is to provide valuable insights into the capabilities of FMs in health, facilitating responsible deployment and mitigating associated risks.

I. INTRODUCTION

Foundation models evolved from transformer-based architectures and self-supervised learning into versatile systems for healthcare, where domain-specific data and task diversity motivate dedicated medical models. The survey organizes this landscape around medical FM development, applications, taxonomies, opportunities, challenges, and open questions.

  • I. INTRODUCTION: Transformers introduced parallel self-attention and enabled foundation models that generalized across text, vision, video, speech, tabular data, protein sequences, and other domains.Vision Transformers extended the architecture beyond NLP, while pretrained models can be fine-tuned for healthcare applications.
  • I. INTRODUCTION: BERT became an early influential foundation model, with bidirectional sentence encoding and scalable architectures that inspired RoBERTa, BART, T5, and later large language models.These descendants became prevalent in state-of-the-art NLP model development.
  • I. INTRODUCTION: Healthcare-specific foundation models address limited labeled data and difficult cross-task generalization that constrain specialized task models and direct adoption of general-purpose systems.The survey identifies complex medical data structures and limited public medical data as additional adoption challenges.
  • I. INTRODUCTION: The survey covers medical FM history, applications, taxonomies, opportunities, challenges, open research questions, and lessons learned across healthcare.Its application coverage includes clinical NLP, medical computer vision, graph learning, biology, and omics.
  • I. INTRODUCTION: Existing surveys largely emphasize general-purpose FMs or limited model families, leaving healthcare taxonomies spanning language, vision, protein, audio, and graph models insufficiently covered.This survey positions its healthcare focus as a response to that coverage gap.

E. Contribution

The survey contributes a broad account of healthcare foundation models, their architectures, applications, and domain-specific adaptations. It emphasizes that medical performance often requires specialized data and models rather than direct transfer from general-purpose systems.

  • E. Contribution: The survey reviews healthcare FMs across clinical NLP, medical computer vision, biology and omics, and related healthcare applications while outlining challenges and future directions.It analyzes recent work published through early 2024.
  • E. Contribution: The survey compares existing survey coverage and presents flagship architectures including BERT, GPT, and Stable Diffusion as foundations for medical FM development.It notes that representative general-purpose models use encoder, decoder, or encoder-decoder architectures.
  • E. Contribution: Directly applying general-purpose FMs to healthcare is difficult because biomedical language, fine-grained medical images, and three-dimensional imaging differ from their training domains.The survey cites uncertainty for general-domain language models, data demands for CLIP, and SAM’s limitations on 3D volumetric images.
  • E. Contribution: Medical-specific models such as BioMegatron, GatorTronGPT, and MedSAM tend to outperform general-purpose counterparts across clinical NLP and medical image segmentation tasks.MedCLIP also shows strong zero-shot image-report pairing capabilities for medical diagnosis and interaction.
  • E. Contribution: Healthcare text-to-image FMs support applications including text-conditional MRI synthesis and chest X-ray generation, with MedXChat reported to improve image and report generation across tasks.The survey also links high-quality synthetic medical images to diagnostic assistance and reduced repeated scans.

E. Omics

Foundation models extend omics analysis beyond conventional task-specific models by learning from large biological datasets and supporting diverse single-cell and molecular tasks. The survey situates these models alongside clinical language, imaging, and information-extraction applications.

  • E. Omics: Omics foundation models analyze complex genomics, proteinomics, metabolomics, transcriptomics, lipidomics, and microbiomics data, addressing limited model generalizability.Reported applications include cell-type classification and imputing biologically relevant single-cell RNA-sequencing values.
  • E. Omics: Geneformer maps gene networks with a context-aware transformer trained on 29.9 million single-cell transcriptomes from publicly available datasets.Its architecture contains six encoder units with self-attention and feedforward layers.
  • E. Omics: The survey places omics among five medical FM categories alongside clinical language, vision, video/audio, and multimodal models.This taxonomy reflects the diversity of healthcare data types and model architectures.

4) Information retrieval (bedside decision support):

Foundation models support bedside decision support through information retrieval, medical image understanding, segmentation, and emerging diffusion-based approaches. These applications aim to improve access to relevant clinical information while reducing dependence on extensive annotations.

  • Information retrieval: Foundation models can support clinical decision-making by retrieving comprehensive, pertinent historical information and optimizing clinical decision-support alerts.Their use is described as a way to improve healthcare delivery and learning health systems.
  • Classification: CheXzero performs chest X-ray pathology classification without explicit annotations, achieving accuracy comparable to radiologists.The example illustrates how self-supervised learning can reduce reliance on manually labeled data.
  • Segmentation: BrainSegFounder reduces labeled-data requirements while achieving high segmentation accuracy and outperforming prior models on benchmark datasets.It is a 3D medical foundation model based on SwinUNETR for neuroimage segmentation.
  • Segmentation: UniverSeg was trained on over 22,000 scans from 53 open-access datasets spanning 16 modalities and generalized across tasks with minimal labeled data.The model was proposed to address limited training data, complex anatomy, and generalization challenges.
  • Segmentation: Diffusion models are presented as promising for volumetric, dental, tumor, and vessel segmentation across multiple medical imaging settings.Reported approaches include diffusion models combined with U-shaped networks or adversarial learning.

4) Localization:

Foundation models enable localization-oriented radiology assistance, image reconstruction, enhancement, modality synthesis, and real-time procedural support. Their potential spans report drafting, anatomically focused descriptions, domain-shift-resilient reconstruction, and multimodal surgical assistance.

  • Localization: Foundation models can draft radiology reports, contextualize findings with patient history, support interactive queries, and improve report completeness through region-guided generation.Large-scale implementation may require integrating patient history, laboratory results, and previous imaging studies.
  • Reconstruction: Diffusion-based MRI reconstruction can address subject motion and domain shifts in multi-contrast brain MRI while supporting accelerated reconstruction.AdaDiff uses adaptive diffusion, an adversarial mapper, and a two-phase reconstruction approach.
  • Reconstruction: Foundation models can enhance healthcare accessibility by improving image quality in facilities lacking advanced hardware, scanning capabilities, or technical expertise.The cited reconstruction work frames diffusion modeling as promising for MRI reconstruction.
  • Image enhancement: Super-resolution methods reconstruct high-resolution medical images from lower-resolution inputs to provide clinicians with more precise visual information.The stated objective is to enhance fine details and image clarity.
  • Procedural support: Multimodal foundation models could assist surgical teams in real time by annotating video streams and alerting them to skipped procedural steps.The proposed assistance builds on augmented reality and multimodality for contextually relevant intervention support.

9) Vision-based conversation:

Vision-based conversational systems combine language with images or videos for healthcare interaction and analysis. The broader survey also connects these applications to graph learning, brain-network analysis, knowledge discovery, drug discovery, and radiomics.

  • Vision-based conversation: Healthcare vision-language conversational agents can interpret textual and visual medical data, including patient-shared images and videos.Examples include Valley, Qilin-Med-VL, VideoChat, VAST, visualGPT, NExT-GPT, and CLIPSyntel.
  • Vision-based conversation: Some healthcare chatbots analyze X-rays and OCT images to detect conditions including pneumonia, drusen, choroidal neovascularization, and diabetic macular edema.These systems extend conversational interaction beyond textual patient inputs.
  • Graph learning: Large language models face graph-learning challenges involving mathematical precision, logical reasoning, spatial perception, and temporal information.Proposed responses include graph-specific frameworks, ChatGPT and Toolformer integration, and multimodal data.
  • Graph learning: Foundation models are positioned as alternatives to traditional graph neural networks for extracting patterns from brain data and complex neural interactions.The survey relates this potential to neurological and psychiatric disorders studied through brain connectomes.
  • Knowledge discovery: Medical knowledge discovery uses foundation models to identify correlations, biomarkers, and therapeutic targets relevant to drug development and personalized medicine.The survey also describes foundation models as successful in knowledge and drug discovery.
  • Drug discovery: Generative drug design and molecular foundation models have improved molecular quality and drug-property prediction while potentially reducing development resources.The cited work highlights possible reductions in cost, time, manpower, and hardware requirements.

3) Proteomics:

Foundation models support proteomics and protein design through sequence generation, property prediction, and structure-oriented modeling, while also extending to medical video and audio analysis.

  • Protein design: FMs support protein design by generating sequences, predicting properties, and designing protein structures from text or learned biological patterns.ProteinDT used contrastive learning and generative modeling, while ProtGPT2, ProGen, GMAI, and diffusion-based models address complementary design tasks.
  • Protein design: 441K text–protein pairs enabled ProteinDT to outperform state-of-the-art methods on protein property prediction benchmarks.
  • Medical video: FMs can extract spatial and temporal information from medical video, including endoscopic footage pretrained on 5 million frames from 33K videos.
  • Speech and audio: Audio FMs support spoken dialogue and can detect smaller temporal variations in depressed mood through zero-shot personalization.

IV. DISCUSSIONS

The survey organizes healthcare FMs into five application categories and reports broad task performance, while emphasizing specialized models, adaptation opportunities, and clinician-centered trustworthiness.

  • A. Taxonomies and Comparative Results: Healthcare FMs span CLLMs, omics, vision-based, video/audio, and multimodal models, reflecting diverse medical data types and tasks.
  • A. Taxonomies and Comparative Results: Specialized medical FMs tend to outperform general-purpose models in clinical NLP and medical image segmentation, while MedCLIP shows strong zero-shot image–report pairing.
  • B. Opportunities: Medical FMs can adapt to rare diseases with few-shot or zero-shot learning, including fine-tuning common-disease imaging models on limited rare-condition data.
  • B. Opportunities: Diverse pretraining data and multimodal inputs can improve robustness, while self-supervised learning and limited fine-tuning can reduce exposure to sensitive patient information.
  • B. Opportunities: Clinician intervention, customization, and references to support underlying reasoning are presented as mechanisms for more trustworthy FM-assisted healthcare.

C. Challenges to FM in Healthcare

Healthcare foundation models face challenges involving computational cost, accuracy, interpretability, validation, verification, bias, data availability, and deployment. The survey identifies domain tailoring, self-supervised learning, multimodality, interpretability, practical model size, and careful validation as key research priorities.

  • Challenges: Healthcare FMs remain difficult to deploy because they require costly computation, may hallucinate, and can be hard to interpret, validate, verify, and regulate.These challenges affect training, inference, monitoring, maintenance, clinical reliability, and legal or ethical accountability.
  • Tailoring FMs for Healthcare: General-purpose models often require healthcare-specific tailoring because medical text, images, and other data differ substantially from general-domain training data.Specialized variants such as BioBERT and GatorTron are fine-tuned on biomedical data for medical NLP tasks.
  • Self-Supervised Learning and High-Quality Medical Data: Self-supervised learning reduces dependence on labeled data, but high-quality annotated medical datasets remain necessary for fine-tuning and reliable healthcare applications.Privacy concerns and data complexity make medical-data acquisition and sharing difficult, motivating privacy-preserving approaches such as federated learning.
  • Multi-modality and Artificial General Intelligence: Multimodal models integrating text, images, and omics data are presented as better suited to complex healthcare decision-making than single-modality systems.The survey cites MedSAM and MedCLIP as examples of multimodal approaches used for diagnostic image classification and report generation.
  • Model Interpretability for Clinical Adoption: Clinical adoption requires human-centered interpretability because strong task performance does not ensure transparent or trustworthy decisions.The survey notes that models can exhibit unexpected behaviors or sensitivity to irrelevant visual features, undermining clinical trust.
  • Model Size and Practical Use: Larger models can improve performance but impose deployment barriers, with MedSAM and BiomedCLIP requiring substantial multi-GPU memory.Personalized-medicine applications also require careful validation so genomic predictions are reliable and actionable.

V. CONCLUSION

The survey reviews foundation models across healthcare, covering general-purpose and specialized models, their applications, and their contributions to healthcare delivery. It also addresses associated challenges and provides insights for responsible development and deployment.

  • V. CONCLUSION: The survey synthesizes general-purpose and healthcare-specific foundation models across nearly all major healthcare applications.Its coverage spans clinical decision support and other healthcare domains described throughout the survey.
  • V. CONCLUSION: The survey frames foundation models as contributing to advances across multiple aspects of healthcare delivery while emphasizing associated challenges.

APPENDIX A BACKGROUND

This background section introduces Transformers, attention, self-supervised learning, reinforcement learning with human feedback, and their roles in foundation-model development. It explains how these mechanisms support parallel sequence processing, representation learning, decision optimization, and model adaptation.

  • A. Transformer: Transformers process sequences through attention in an encoder-decoder architecture, replacing sequential recurrence with parallel computation.The encoder forms intermediate representations and the decoder generates outputs autoregressively.
  • B. Attention: Self-attention weights token representations according to token similarity, while positional encoding preserves relative sequence positions.Transformer attention uses query, key, and value representations to compute token relationships.
  • C. Self-supervised learning (SSL): Self-supervised learning creates training signals from unlabeled data, reducing labeling requirements and supporting downstream generalization.The survey distinguishes contrastive, generative, and adversarial forms of self-supervised learning.
  • D. Reinforcement learning: Human-feedback reinforcement learning supplements predefined rewards by using trainer feedback to guide an agent’s actions.The underlying reinforcement-learning formulation models states, actions, transitions, rewards, and discounting.

E. CLIP

CLIP addresses the labeling and task-specific generalization limits of traditional vision models by learning aligned image-text representations from natural-language supervision. Its dual-encoder design supports zero-shot transfer, although foundation models generally rely on transformer architectures and varied encoder or decoder configurations.

  • E. CLIP: Traditional vision models depend heavily on labeled data and often generalize poorly beyond their training tasks, motivating CLIP’s language-supervised approach.
  • E. CLIP: CLIP uses natural-language supervision over 400 million image-text pairs to learn aligned visual and textual representations.Separate image and text encoders transform both modalities into representations used for cross-modal alignment.
  • E. CLIP: CLIP supports zero-shot classification with natural-language instructions without direct optimization for each benchmark.Its scalable language-based supervision avoids conventional task-specific annotation requirements.
  • E. CLIP: Foundation models commonly use transformer-based encoder-only, decoder-only, or encoder-decoder architectures for downstream tasks.BERT, ViT, and CLIP are encoder-only examples, whereas GPT models are decoder-only.

B. Learning algorithms

The survey describes self-supervised learning, prompting, and feedback-based learning as strategies for adapting foundation models to downstream tasks. It also reviews flagship architectures and efficiency-oriented variants, including BERT, GPT-related models, CLIP, and Stable Diffusion.

  • Learning strategies: Self-supervised learning enables foundation models to learn contextual representations from unlabeled data and generalize through downstream fine-tuning.This reduces dependence on expensive, potentially biased annotations and narrowly applicable supervised learning.
  • Learning strategies: Prompting directs foundation models toward specified tasks by guiding response generation and allocating model resources accordingly.CLIP uses textual prompts for supervision, while MaskCLIP aligns image regions with textual descriptions.
  • BERT: BERT uses masked language modeling during pretraining and self-attention during fine-tuning to support downstream tasks with limited labeled data.RoBERTa extends this approach with longer training, larger corpora, and dynamic masking.
  • BERT: DistilBERT reduces BERT’s size by nearly 40% and increases speed by 60% while maintaining comparable downstream performance.It uses knowledge distillation, with a smaller student model trained to reproduce a larger teacher model.
  • BERT: ALBERT is 1.7 times faster, uses 18 times fewer parameters, and achieves better performance than BERT, whereas BART is computationally expensive with 10% more parameters.

B. GPT

The survey reviews GPT-related and industrial-scale pretrained models, alongside multimodal and generative systems such as LLaVA, DALL-E, and Stable Diffusion. These models extend transformer-based learning across language, vision-language understanding, image generation, and healthcare-oriented applications.

  • GPT: GPT uses unsupervised pretraining followed by supervised fine-tuning for natural language understanding tasks.Its pretraining uses the BooksCorpus dataset, after which learned parameters are adapted with manually annotated data.
  • GPT: DALL-E and DALL-E 2 generate images from text or image inputs using compressed image representations, CLIP-based embeddings, and generative modeling.DALL-E was trained on 250 million online text-image pairs, while DALL-E 2 feeds CLIP-induced image embeddings into a diffusion model.
  • GPT: LLaVA connects a CLIP visual encoder with Vicuna and uses approximately 158,000 image-text instruction pairs for multimodal understanding, captioning, and reasoning.
  • GPT: Stable Diffusion separates compression from generation by operating in a lower-dimensional representation space, improving diffusion-model computational efficiency.
  • Industrial pretrained models: Industrial-scale models extend transformer pretraining to dialogue, reasoning, multilingual tasks, coding, enterprise services, and healthcare applications.LaMDA contains 137B parameters, GLaM contains 1.2 trillion parameters, and PaLM contains 540 billion parameters.
  • Industrial pretrained models: GLaM achieves better performance than GPT-3 on 29 NLP tasks across nearly all zero-shot, one-shot, and few-shot settings while consuming one-third of GPT-3’s energy.
Loading 2406.10729v3…