Source-linked AI summary
Artificial Intelligence-Based Methods for Fusion of Electronic Health Records and Imaging Data
Farida Mohsen, Hazrat Ali, Nady El Hajj, Zubair Shah
TL;DR
Clinical interpretation benefits from combining complementary EHR and medical imaging information, but the literature needed systematic synthesis. This scoping review searched four databases and analyzed 34 studies to characterize AI fusion strategies, applications, algorithms, and datasets. Early fusion was most common, and multimodal models generally outperformed single-modality models for the same tasks.
Problem
Clinical applications use complementary EHR and imaging information, motivating synthesis of AI methods that fuse these modalities across tasks.
Method
The review searched Embase, PubMed, Scopus, and Google Scholar and analyzed studies using AI to fuse EHR with medical imaging data.
Results
The review analyzed 34 studies and found that multimodal models generally outperformed single-modality models for the same disease diagnosis or prediction task.
Takeaways & Limitations
The review supports attempting EHR–imaging fusion when multimodal data are obtainable, while noting that many studies use relatively simple strategies.
Takeaways & Limitations
The review was limited to English-language studies published from 2015–2022 that fused EHR with medical imaging, and publication bias may overestimate multimodal benefits.
Abstract
from arXiv · showhide
Healthcare data are inherently multimodal, including electronic health records (EHR), medical images, and multi-omics data. Combining these multimodal data sources contributes to a better understanding of human health and provides optimal personalized healthcare. Advances in artificial intelligence (AI) technologies, particularly machine learning (ML), enable the fusion of these different data modalities to provide multimodal insights. To this end, in this scoping review, we focus on synthesizing and analyzing the literature that uses AI techniques to fuse multimodal medical data for different clinical applications. More specifically, we focus on studies that only fused EHR with medical imaging data to develop various AI methods for clinical applications. We present a comprehensive analysis of the various fusion strategies, the diseases and clinical outcomes for which multimodal fusion was used, the ML algorithms used to perform multimodal fusion for each clinical application, and the available multimodal medical datasets. We followed the PRISMA-ScR guidelines. We searched Embase, PubMed, Scopus, and Google Scholar to retrieve relevant studies. We extracted data from 34 studies that fulfilled the inclusion criteria. In our analysis, a typical workflow was observed: feeding raw data, fusing different data modalities by applying conventional machine learning (ML) or deep learning (DL) algorithms, and finally, evaluating the multimodal fusion through clinical outcome predictions. Specifically, early fusion was the most used technique in most applications for multimodal learning (22 out of 34 studies). We found that multimodality fusion models outperformed traditional single-modality models for the same task. Disease diagnosis and prediction were the most common clinical outcomes (reported in 20 and 10 studies, respectively) from a clinical outcome perspective.
Introduction
Healthcare data span multiple modalities whose complementary information can support clinical inference, while AI and ML enable their fusion for medical applications. This scoping review analyzes AI-based fusion of EHR and medical imaging data and aims to characterize the strategies, applications, and resources in the literature.
- Data fusion combines modalities that provide different viewpoints on a common phenomenon to solve an inference problem.Fusion techniques aim to exploit cooperative and complementary features across modalities.
- Clinical data provide important context for interpreting medical images across radiology, dermatology, ophthalmology, and pathology.Missing pertinent clinical and laboratory data can reduce radiologists’ diagnostic accuracy.
- AI and ML models support fusion of multimodal data with high dimensionality, varied statistical properties, and different missing-value patterns.
- Prior studies reported improved performance over single-source or single-modality models when combining imaging with laboratory, demographic, or clinical data.Examples included Alzheimer’s disease, breast cancer, diabetic retinopathy, COVID-19, and glaucoma applications.
- This scoping review focuses on published studies using AI models to fuse medical images with EHR data for clinical applications.The review distinguishes its scope from reviews focused on imaging-only, omics, IoMT, or multimodal EHR data.
1. Fusion Strategies:
The review asks which fusion strategies researchers use to combine medical imaging data with EHR and which strategy is most frequently used.
- The review examines fusion strategies used to combine medical imaging data with EHR and identifies the most-used method.
2. Diseases:
The review considers the diseases, clinical outcomes, datasets, and broader research context associated with multimodal fusion of EHR and medical imaging data. It emphasizes the need for more multimodal medical data resources.
- The review asks which types of diseases use fusion methods.
- It examines clinical outcomes addressed by different fusion strategies and the ML algorithms used for each outcome.
- It identifies publicly accessible medical multimodal datasets as a resource of interest.
- The review aims to clarify how models align EHR and imaging modalities for clinical tasks and identify shortages of multimodal data resources.
Preliminaries
The review defines its multimodal setting as the combination of medical imaging and EHR data, then classifies fusion by when modality features are combined during model development.
- Preliminaries: The review analyzes EHR and medical imaging as its two primary data modalities.
- Preliminaries: Medical imaging includes clinical N-dimensional data such as X-ray, MRI, fMRI, sMRI, PET, CT, and ultrasound.
- Preliminaries: EHR data include structured codes, laboratory, demographic, family-history, vital-sign, medication, and unstructured report or clinical-note data.
- Preliminaries: Multiple EHR or imaging submodalities are treated as one EHR or imaging modality when studies combine the two broader modalities.
- Fusion strategies: Early fusion joins modality features at the input level before a single ML algorithm trains on them.Features may be original or extracted manually, statistically, computationally, or with neural networks.
- Fusion strategies: Late fusion trains separate modality-specific ML models and combines their predictions through voting, weighted averaging, or a meta-classifier.This is also called decision-level fusion.
- Fusion strategies: Joint fusion combines learned intermediate neural-network features with other modalities while propagating final-model loss back to feature extractors.Iterative weight updates improve learned feature representations during training.
Methods
The review used PRISMA-ScR-guided searches and eligibility criteria to identify studies fusing EHR with medical imaging through AI. Researchers screened records, extracted study characteristics, synthesized findings narratively, and did not assess study quality.
- The review followed PRISMA-ScR guidelines and searched Scopus, PubMed, Embase, and Google Scholar.
- The search targeted AI-based multimodal fusion of medical imaging and EHR data.
- Eligible studies fused EHR with clinical imaging using AI, while studies using single modalities, multimodal omics, nonmedical data, or non-AI models were excluded.
- Two authors conducted title, abstract, and full-text screening, resolving disagreements through discussion or third-author consultation.
- The team piloted a data-extraction form and synthesized studies narratively across fusion strategies, diseases, outcomes, algorithms, data sources, and evaluation.
- The review did not perform quality assessments of included studies.
Results
The review included 34 studies spanning diverse diseases and publication settings, with early fusion the dominant strategy. Multimodal models generally outperformed single-modality comparators, and neurological disorders were most represented.
- Study selection: 34 studies met the inclusion criteria after screening 1158 retrieved records and 44 full texts.
- Study characteristics: 23 studies (∼68%) were journal articles, 22 (∼65%) were published during 2020–2022, and the USA contributed 10 studies (∼30%).
- Fusion strategies: 22 studies (∼65%) used early fusion, typically extracting imaging features before combining them with EHR features.
- Model comparison: 13 of 14 evaluated early-fusion studies reported better performance than both imaging-only and clinical-only models.
- Fusion strategies: 10 studies used joint fusion, combining neural-network-derived imaging and EHR representations through concatenation or downstream neural networks.
- Fusion strategies: 2 studies used late fusion, and one reported that late fusion outperformed early, joint, and single-modality models.
- Diseases: The included studies covered seven disease categories, with neurological disorders predominating in 16 studies.
Clinical outcomes and machine learning models
The reviewed studies applied multimodal fusion of medical imaging and EHR data primarily to diagnosis and prediction. Early, joint, and late fusion strategies were combined with conventional ML and DL models across neurological, psychiatric, cardiovascular, cancer, and other diseases.
- Diagnosis: 20 studies (∼59%) addressed diagnosis, spanning neurological, psychiatric, cardiovascular, cancer, and other diseases.Neurological diagnoses accounted for 9 studies, including 4 on Alzheimer’s disease and 4 on mild cognitive impairment.
- Diagnosis: 13 diagnostic studies used early fusion, concatenating imaging and EHR-derived features before classification with models including SVM, discriminant analysis, neural networks, and other algorithms.Applications included Alzheimer’s disease, mild cognitive impairment, demyelinating diseases, bipolar disorder, schizophrenia, and cancer.
- Diagnosis: 5 diagnostic studies used joint fusion, with deep architectures combining learned imaging representations and processed clinical or EHR features.Examples included Bayesian deep multisource learning for glaucoma and multimodal networks for cardiomegaly and myocardial infarction detection.
- Diagnosis: 2 diagnostic studies used late fusion, and the late-fusion approach outperformed individual image-only and tabular-only models for pulmonary embolism diagnosis.One study combined MRI and cognitive-test models using majority voting; another combined CT and EHR models.
- Early Prediction: 14 studies (∼41.2%) addressed prediction, including disease, treatment outcome, mortality, and overall survival prediction.Disease prediction was reported in 10 studies, treatment outcome prediction in 2, and mortality and overall survival prediction in 1 study each.
- Early Prediction: 6 disease-prediction studies used early fusion, 4 used joint fusion, and both mortality and overall-survival studies used early fusion.Multimodal fusion outperformed single-modality performance in treatment-outcome, mortality, and overall-survival prediction studies.
Datasets
The review included diverse imaging and EHR data sources, with MRI, PET, and structured EHR data used most often. Private datasets predominated, while ADNI was the most frequently used public dataset.
- Imaging and EHR modalities: 13 of 34 studies used MRI images and 8 used PET images, mostly for Alzheimer’s disease diagnosis and prediction.Other imaging modalities included CT, fMRI, structural MRI, diffusion MRI, DTI, ultrasound, X-ray, and fundus images.
- Imaging and EHR modalities: Structured EHR data were the most common EHR modality, appearing in 32 studies.The included studies also used unstructured patient data and multiple types of clinical information.
- Dataset availability: 21 studies (∼59%) used private data sources, whereas 13 used publicly accessible datasets.ADNI was the most used public dataset, appearing in 7 of the 13 studies using public data.
Evaluation metrics
Evaluation metrics varied with the clinical task, with accuracy, AUC, sensitivity, specificity, F1-measure, and precision commonly used for diagnosis and prediction.
- Common metrics: Accuracy, AUC, sensitivity, specificity, F1-measure, and precision were the main metrics used to evaluate diagnosis and prediction tasks.Metric selection was mainly dependent on the clinical task.
Discussion
The review finds that EHR–imaging fusion generally improves disease diagnosis and prediction, while implementation choices depend on modality relationships, dataset size, and available data resources. Early fusion was most common, but limited public data, simple non-imaging features, and heterogeneous reporting constrain broader conclusions.
- Multimodal EHR–imaging models generally outperformed single-modality models for the same disease diagnosis or prediction task.
- Fusion implementation: Early fusion was the most commonly used strategy, combining vectorized imaging representations with 1D EHR features before prediction.CNN-derived imaging features often performed better than manually or software-derived features; CNN-based early fusion requires multiple models to be trained.
- Fusion implementation: Joint fusion was the second most common approach and is preferred mainly for large datasets because neural-network implementations can be limiting with small datasets.Joint models iteratively update feature representations through loss propagation to learn cross-modal correlations.
- Fusion implementation: When modalities are complementary, early or joint fusion was preferred; late fusion was considered suitable for independent modalities and smaller datasets.The review gives Alzheimer’s disease diagnosis as an example using early or joint fusion, while MRI pixels and MMSE scores for mild cognitive impairment illustrate late fusion.
- Applications: Alzheimer’s disease diagnosis and prediction were the most common multimodal applications, and fusion techniques consistently improved Alzheimer’s disease diagnosis.
- Prospects and limitations: The evidence base is constrained by limited benchmarking data, relatively simple non-imaging EHR features, publication bias, and heterogeneous tasks and metrics that prevent consistent direct comparison.The review also notes that unavailable multimodal public data may limit conclusive clinical claims for global populations.
- Future directions: Integrating EHR, imaging, and multi-omics data may provide a more holistic view for personalized medicine, while federated learning may support secure multisite multimodal data collection.
Conclusion
The review synthesizes research on AI-based fusion of EHR and medical imaging, including strategies, clinical tasks, models, diseases, and public datasets. It reports growing interest and effective but often simple fusion approaches, while noting that newer studies may fall outside its strategy definitions.
- The review surveys multimodal medical ML studies combining EHR with medical imaging across fusion strategies, clinical tasks, models, diseases, and publicly accessible datasets.
- Most studies use relatively simple fusion strategies that have been effective but might not fully exploit the rich information in EHR and imaging modalities.
- The field’s continued development may produce studies outside the review’s definitions of fusion strategies or using combined strategies.
- The authors expect multimodal medical data analysis to support clinical decision-making as the field develops.
Author contributions statement
The author contributions statement assigns roles spanning conceptualization, project administration, data curation, synthesis, writing, editing, and supervision.
- F.M., H.A., and Z.S. contributed to conceptualization, while F.M. and H.A. administered the project.
- F.M. curated the data, performed data synthesis, and contributed to the original draft; H.A. and N.E. contributed to review and editing.
- Z.S. and H.A. supervised the study, and all authors read and approved the final manuscript.