Source-linked AI summary
Large AI Models in Health Informatics: Applications, Challenges, and the Future
Jianing Qiu, Lin Li, Jiankai Sun, Jiachuan Peng, Peilun Shi, Ruiyang Zhang, Yinzhao Dong, Kyle Lam, Frank P. -W. Lo, Bo Xiao, Wu Yuan, Ningli Wang, Dong Xu, Benny Lo
TL;DR
Health informatics faces expanding multimodal data and costly expert annotation, motivating large AI models as a new approach to biomedical learning. This review surveys their background, applications, challenges, and future directions, reporting advances in biomedical modeling while emphasizing augmentation of professionals rather than replacement and the need for safeguards.
Problem
Expanding multimodal health data and costly expert annotation create challenges for developing broadly capable biomedical and clinical AI models.
Method
The article comprehensively reviews large AI models, their health-informatics applications, challenges, risks, and future directions.
Results
Large AI models support advances across biomedical tasks, including protein and RNA modeling, while scaling and self-supervised learning address some data limitations.
Takeaways & Limitations
The authors conclude that large AI models will augment rather than replace medical professionals, with human-AI cooperation and domain collaboration becoming important.
Takeaways & Limitations
Public health-informatics datasets are generally smaller than general-domain datasets and may be insufficient to unlock the full potential of large AI models.
Abstract
from arXiv · showhide
Large AI models, or foundation models, are models recently emerging with massive scales both parameter-wise and data-wise, the magnitudes of which can reach beyond billions. Once pretrained, large AI models demonstrate impressive performance in various downstream tasks. A prime example is ChatGPT, whose capability has compelled people's imagination about the far-reaching influence that large AI models can have and their potential to transform different domains of our lives. In health informatics, the advent of large AI models has brought new paradigms for the design of methodologies. The scale of multi-modal data in the biomedical and health domain has been ever-expanding especially since the community embraced the era of deep learning, which provides the ground to develop, validate, and advance large AI models for breakthroughs in health-related areas. This article presents a comprehensive review of large AI models, from background to their applications. We identify seven key sectors in which large AI models are applicable and might have substantial influence, including 1) bioinformatics; 2) medical diagnosis; 3) medical imaging; 4) medical informatics; 5) medical education; 6) public health; and 7) medical robotics. We examine their challenges, followed by a critical discussion about potential future directions and pitfalls of large AI models in transforming the field of health informatics.
I. INTRODUCTION
Large AI models introduce a health-informatics paradigm based on large-scale models, training, and generalization across heterogeneous biomedical data. The review surveys their opportunities, applications, challenges, and future directions.
- Core features: Large AI models are characterized by increased model size, large-scale training, multimodal processing, and broad downstream-task performance.Examples include billions of parameters, trillions of language tokens, and billions of vision images, with zero-, one-, and few-shot capabilities.
- Training paradigm: LAMs can reduce reliance on large expert-annotated datasets by using self-supervision and reinforcement learning during training.Expert annotation is expensive and time-consuming, making high-quality clinical-data curation difficult.
- Health-data opportunity: Health informatics contains expanding multimodal data from wearables, EHRs, medical imaging, and genomic sequencing to support large models.These models are expected to represent complex health-related data and generalize to unseen scenarios involving medical decision-making.
- Multimodality: LAMs learn heterogeneous health data through large capacity, unified multimodal input modeling, and improved multimodal learning techniques.The multimodal nature of biomedical and health data provides a basis for developing and evaluating these models.
- Review scope: The review focuses mainly on foundation models while also retrospectively covering seminal non-foundational LAMs that advance biomedicine and health informatics.It identifies a paradigm shift toward large-scale model size, large-scale pre-training, and broad generalization.
- Review scope: The article surveys LAM applications, challenges, limitations, risks, and future directions across biomedical and health informatics.Because the field is rapidly progressing and page-limited, the review cannot cover many works.
II. BACKGROUND OF LARGE AI MODELS
Large AI model development combines large-scale pre-training, adaptable architectures, and broad downstream capabilities across language and vision. Scaling improves performance in many settings, but its optimality and resource demands remain unsettled.
- Large Language Models: LLMs are pre-trained on large unlabelled language datasets, scaled to billions of parameters, and adapted to downstream tasks.Common paradigms include weakly supervised or self-supervised objectives followed by fine-tuning or adaptation.
- Large Language Models: LLMs demonstrate zero-shot, one-shot, and few-shot extraction, summarization, translation, generation, and reasoning capabilities.Prompt engineering techniques such as Chain-of-Thought prompting can further strengthen reasoning.
- Large Language Models: Continuous growth has produced models with hundreds of billions of parameters and trillions of tokens, without agreement that further scaling is optimal.The community also lacks a verified universal scaling law.
- Large Language Models: RLHF uses human preference data to train a reward model that is optimized with reinforcement-learning algorithms to align model behavior with human intent.This approach balances annotation cost, efficacy, and behavioral alignment.
- Large Vision Models: Weakly supervised, self-supervised, or unsupervised pre-training emerged partly because annotation quality and dataset accessibility are compromised.These approaches include autoregressive, generative, and contrastive learning.
- Large Vision Models: Vision models include ViT and CNN architectural families, with hybrid architectures also combining both approaches.Scaling vision transformers can achieve state-of-the-art benchmark accuracy, while redesigned CNNs can also achieve state-of-the-art accuracy.
- Large Vision Models: Vision-model scaling is paired with scaling rules, larger pre-training datasets, and efficient parallelism because simply increasing depth may be suboptimal.Scaling choices can significantly affect final performance.
- Large Vision Models: SAM uses a 632M-parameter ViT-H image encoder, prompt encoder, and transformer mask decoder to predict masks from points, boxes, or text.It demonstrates zero-shot generalization to unseen objects and images.
C. Large Multi-modal Models
Large multimodal models align or fuse representations from different modalities, especially language and vision, using large-scale paired data. Their scope is expanding toward unified frameworks spanning many modalities.
- Data: Training vision-language models can require hundreds of millions of image-text pairs, and LAION-5B provides 5.85B public samples.Earlier datasets of this scale were often closed source.
- Architecture: LVLMs commonly process text and images through separate encoders before aligning or fusing their representations.Alignment may use contrastive learning, while fusion may use another encoder over extracted features.
- Generation: Text-to-image generation uses autoregressive or diffusion models with different generation processes.Autoregressive models predict the next sequence item, whereas diffusion models progressively perturb images with noise before generation.
- Beyond vision-language: Multimodal systems can use language models to instruct vision models for vision-language tasks.This extends multimodal integration beyond jointly learning aligned representations.
- Beyond vision-language: Some models unify modalities beyond language and vision, with ImageBind combining six modalities and Meta-Transformer unifying twelve.
III. APPLICATIONS OF LARGE AI MODELS IN HEALTH INFORMATICS
LAM applications span seven biomedical and health sectors, with protein and RNA foundation models illustrating how self-supervised learning and scaling address limited labeled data. These gains reduce prediction time and cost but do not eliminate the need for experimental validation.
- Application scope: The review identifies bioinformatics, medical diagnosis, medical imaging, medical informatics, medical education, public health, and medical robotics as key application sectors.It compares current LAMs with prior state-of-the-art methods in typical tasks across these sectors.
- Bioinformatics: Protein language models learn global relations and long-range dependencies from unaligned, unlabelled sequences through self-supervision.They support protein-sequence generation, secondary-structure prediction, and structure prediction without always requiring MSAs or templates.
- Bioinformatics: 0.68 vs 0.38 for TM-score on CASP14: ESMfold improved over AlphaFold2 while providing considerably faster inference without MSAs and templates.
- Bioinformatics: RNA-FM learned from 23 million unlabeled ncRNA sequences and achieved a 3-5% increase across Precision, Recall, and F1-score versus UFold.It supports RNA secondary-structure and 3D-closeness prediction and enabled direct 3D RNA structure prediction.
- Bioinformatics: LAMs have reduced molecule-structure prediction time and cost by a large margin, while conventional experiments remain complementary.Experimental molecular properties can further improve predictions, especially for rare data such as orphan proteins.
- Bioinformatics: Protein-structure models remain data-driven and can struggle with unseen data, including missense mutations, while prediction-quality assessment remains unclear.Unverified structures cannot be safely applied to uses such as drug discovery without protocols and quality metrics.
B. Medical Diagnosis
Large AI models are being explored for medical diagnosis and imaging, combining large-scale pretraining with specialized diagnostic systems. Their promise is substantial, but imaging applications remain constrained by resolution-related information loss and unsettled scaling practices.
- Medical Diagnosis: LAMs may support medical diagnosis and decision-making as their safety and factual grounding improve.
- Medical Diagnosis: CheXzero achieved radiologist-level performance on multiple chest X-ray pathologies absent from its self-supervised training.
- Medical Diagnosis: ChatCAD combines specialized diagnostic networks with iterative ChatGPT prompting to support computer-aided medical-image diagnosis.
- Medical Diagnosis: 8.5 million ECGs supported HeartBEiT pretraining, producing accurate cardiac diagnosis, improved explainability, and reduced annotated-data requirements for fine-tuning.
- Medical Imaging: SAM often fails zero-shot segmentation on MRI and OCT, but adaptation and fine-tuning can surpass current state-of-the-art accuracy.
- Medical Imaging: Medical vision models may lose critical lesion information through image downsampling, while optimal model-data scaling remains inconclusive.
D. Medical Informatics
Medical informatics research applies large language models to biomedical and clinical text, with scaling and domain-specific pretraining improving medical-language performance. These models may assist clinical data processing, documentation, patient support, and trial matching.
- Medical Informatics: Biomedical LLMs were developed from abundant EHRs and public medical texts to perform biomedical text-mining tasks.
- Medical Informatics: 8.9 billion parameters and 82 billion pretraining words enabled GatorTron to improve performance across medical language tasks, especially question answering and inference.
- Medical Informatics: Biomedical pretraining from scratch can outperform continued training from general-domain corpora, while LoRA enables efficient adaptation of large models.
- Medical Informatics: Biomedical LLMs may help clinicians process medical data efficiently and reduce time spent documenting EHRs.
- Medical Informatics: LLMs may generate discharge summaries, assist prior-authorizations, personalize medical assistance, and match patients to clinical trials using demographics and medical history.
E. Medical Education
Large AI models may influence medical education by supporting learners, educators, and training delivery. Their applications include tutoring, content creation, personalized materials, remote instruction, grading, and standardized routine training.
- Medical Education: GPT-4 and Med PaLM 2 passed the USMLE with scores above 86%, indicating capabilities in bioethics, clinical reasoning, and medical management.
- Medical Education: GPT-4 can act as a Socratic tutor, while other LLM applications support medical questions, bioinformatics analysis, and paraphrasing for students with dyslexia.
- Medical Education: Medical education applications also raise concerns about illegitimate uses such as plagiarism, motivating awareness and detection efforts.
- Medical Education: LAMs may create teaching and examination content, diversify presentation formats, personalize course materials, and deliver remote education.
- Medical Education: Domain-knowledgeable LAMs may supervise nurse and other routine medical training, offering responsive interactions and potentially more consistent delivery across trainers.
G. Medical Robotics
Large AI models are being integrated into medical robotics to enhance vision, interaction, autonomy, and user control. Proposed applications span surgical assistance, rehabilitation, companionship, navigation, and language-guided manipulation.
- Medical Robotics: LAMs may enhance medical robotic vision, interaction, and autonomy across surgical, wearable, companion, and assistive robots.
- Medical Robotics: Endo-FM could enhance robotic surgery through endoscopic video classification, segmentation, and detection, while workflow analysis may support complication and outcome prediction.
- Medical Robotics: Generative LAMs may simulate surgical procedures for practice and improve robots’ recognition of patient emotions or navigation for visually impaired users.
- Medical Robotics: LAMs may enable robots to recognize emotions, gestures, and speech and respond to high-level language commands, improving rehabilitation interaction and elderly companionship.
- Medical Robotics: Language-guided LAMs could shift robotic pipelines from engineer-in-the-loop to user-in-the-loop control and help less-programming-proficient surgeons adapt manipulations.
- Challenges: The review identifies ongoing challenges and potential risks in developing and deploying LAMs for biomedical, clinical, and healthcare applications.
1) Data:
Health-informatics LAM development is constrained by limited datasets, costly expert curation and acquisition, privacy restrictions, expensive training, and reliability concerns.
- Data:: Large-scale health datasets are smaller than those in general domains and may be insufficient to unlock LAMs’ full biomedical potential.Curation requires clinical expertise and quality assurance; some modalities require costly specialized devices, while consent, legal, and privacy issues restrict use.
- Computation:: Training and fine-tuning contemporary LAMs can exceed most researchers’ and organizations’ budgets.Training a 65B-parameter LLaMA model took about 21 days on 2048 A100 GPUs using 1.4T tokens.
- Reliability:: Clinical deployment requires higher reliability because LLMs can hallucinate, respond sensitively to prompts, and remain vulnerable to distribution shifts and adversarial examples.The paper recommends caution to reduce potential danger from over-reliance in healthcare practice.
- Reliability:: Up-to-date information is critical in many clinical and health scenarios, although LAMs are trained offline.
- Privacy:: LAMs can memorize training data and expose sensitive information through direct prompts or membership-inference attacks.Fine-tuning to refuse sensitive prompts can be bypassed through jailbreaking.
4) Privacy:
LAM privacy and social risks include leakage of user-provided information, learned demographic and language biases, harmful outputs, limited transparency, and poor interpretability.
- Privacy:: Information submitted to LLM-integrated applications may be stored and leaked through chat-history bugs or indirect prompt injection.
- Fairness:: LAMs can reproduce healthcare and data-collection biases involving race, gender, politics, and language.LLMs perform better in particular languages, such as English, than in languages spoken by fewer people.
- Safety:: Aligned LAMs may still produce hate speech, endorse harmful views, or facilitate disinformation and criminal activities.
- Transparency:: Incomplete disclosure of models, training data, and technical details prevents independent reproduction, improvement, and auditing.The transparency threat is especially serious when medical data are private and resulting models cannot be open sourced.
- Interpretability:: Dense hidden layers make LAM behavior difficult to interpret, predict, and explain.The paper describes outputs and reasoning changes that can appear mysterious or meaningless.
7) Transparency:
The paper frames future LAM work around improving capability while addressing responsibility, including scalability, multimodal world models, hidden-capability discovery, and governance.
- B. Responsibility: Responsible deployment requires regulation covering data rights, liability, and approval for critical healthcare services.
- A. Capability: Future health-informatics LAMs should improve existing abilities or add new ones, including versatile medical task solving and higher diagnostic accuracy.
- A. Capability: Scaling datasets and models efficiently remains important and unresolved for improving LAM capability.
- A. Capability: Pretraining across varied tasks and modalities could combine biomedical knowledge, medical corpora, imaging, and physiological signals in one world model.Such a model could complement information missing from an input during diagnosis.
- A. Capability: Probing existing pretrained LAMs can reveal hidden capabilities and improve tasks without further large-scale training.Prompt engineering is presented as an effective approach, including prompts that strengthen reasoning ability.
B. Responsibility
Responsible LAMs require complementary development and deployment strategies addressing reliability, fairness, transparency, and broader social risks through technical, educational, evaluative, and regulatory measures.
- B. Responsibility: Responsible LAM development should address reliability, fairness, transparency, and related risks through learning and deployment strategies.
- B. Responsibility: Bias mitigation can filter biased data, increase underrepresented populations’ data, and include diverse distributions to improve robustness.Efficient inspection of large-scale datasets remains challenging.
- B. Responsibility: Responsible use requires educating users, studying human-LAM partnerships, and evaluating how prompts and model responses are handled.
- B. Responsibility: Comprehensive verification frameworks, improved benchmarking, and rules governing development, deployment, and use are still needed.The paper calls for collaboration among academia, industry, and government.
- B. Responsibility: The authors expect LAMs to augment rather than replace medical professionals, with human-AI cooperation becoming pervasive.