Source-linked AI summary
Advancing Multimodal Medical Capabilities of Gemini
Lin Yang, Shawn Xu, Andrew Sellergren, Timo Kohlberger, Yuchen Zhou, Ira Ktena, Atilla Kiraly, Faruk Ahmed, Farhad Hormozdiari, Tiam Jaroensri, Eric Wang, Ellery Wulczyn, Fayaz Jamil, Theo Guidroz, Chuck Lau, Siyuan Qiao, Yun Liu, Akshay Goel, Kendall Park, Arnav Agharwal, Nick George, Yang Wang, Ryutaro Tanno, David G. T. Barrett, Wei-Hung Weng, S. Sara Mahdavi, Khaled Saab, Tao Tu, Sreenivasa Raju Kalidindi, Mozziyar Etemadi, Jorge Cuadros, Gregory Sorensen, Yossi Matias, Katherine Chou, Greg Corrado, Joelle Barral, Shravya Shetty, David Fleet, S. M. Ali Eslami, Daniel Tse, Shruthi Prabhakara, Cory McLean, Dave Steiner, Rory Pilgrim, Christopher Kelly, Shekoofeh Azizi, Daniel Golden
TL;DR
General-purpose multimodal models lack specialized medical knowledge for many clinical tasks. This work fine-tunes Gemini into Med-Gemini variants for medical imaging and genomics, achieving strong cross-task results while leaving important clinical-safety and performance limitations unresolved.
Problem
Specialized medical data and clinically grounded evaluations remain insufficiently addressed by general-purpose multimodal models.
Method
The study fine-tunes Gemini 1.5 into modality-specific Med-Gemini models for 2D and 3D medical images and genomic risk prediction.
Results
Med-Gemini shows best-in-class or promising performance across report generation, VQA, classification, and genomic risk prediction tasks.
Takeaways & Limitations
The results demonstrate the potential of Gemini-based models for a broad range of medical tasks spanning imaging and genomics.
Takeaways & Limitations
The genetic-risk AUC values are upper bounds because the GWAS-derived PRS features were created within UK Biobank.
Abstract
from arXiv · showhide
Many clinical tasks require an understanding of specialized data, such as medical images and genomics, which is not typically found in general-purpose large multimodal models. Building upon Gemini's multimodal models, we develop several models within the new Med-Gemini family that inherit core capabilities of Gemini and are optimized for medical use via fine-tuning with 2D and 3D radiology, histopathology, ophthalmology, dermatology and genomic data. Med-Gemini-2D sets a new standard for AI-based chest X-ray (CXR) report generation based on expert evaluation, exceeding previous best results across two separate datasets by an absolute margin of 1% and 12%, where 57% and 96% of AI reports on normal cases, and 43% and 65% on abnormal cases, are evaluated as "equivalent or better" than the original radiologists' reports. We demonstrate the first ever large multimodal model-based report generation for 3D computed tomography (CT) volumes using Med-Gemini-3D, with 53% of AI reports considered clinically acceptable, although additional research is needed to meet expert radiologist reporting quality. Beyond report generation, Med-Gemini-2D surpasses the previous best performance in CXR visual question answering (VQA) and performs well in CXR classification and radiology VQA, exceeding SoTA or baselines on 17 of 20 tasks. In histopathology, ophthalmology, and dermatology image classification, Med-Gemini-2D surpasses baselines across 18 out of 20 tasks and approaches task-specific model performance. Beyond imaging, Med-Gemini-Polygenic outperforms the standard linear polygenic risk score-based approach for disease risk prediction and generalizes to genetically correlated diseases for which it has never been trained. Although further development and evaluation are necessary in the safety-critical medical domain, our results highlight the potential of Med-Gemini across a wide range of medical tasks.
1. Introduction
Med-Gemini extends Gemini toward diverse clinical applications because multimodal medical capabilities remain underexplored and require clinically grounded evaluation. The resulting family spans medical imaging, genomics, classification, VQA, report generation, and risk prediction, with strong performance across several tasks.
- Motivation: Multimodal medical AI must address specialized data and diverse clinical use cases beyond narrow single-input, single-output tasks.Medical data include imaging, genomics, electronic health records, wearables, biosensors, and biobanks.
- Evaluation: The study improves selected open benchmarks by correcting labels, expanding question-answer coverage, and removing train-test contamination from dataset splits.These improvements were intended to strengthen the quality and clinical relevance of evaluation.
- Contribution: Med-Gemini is a family of Gemini-derived models for medical image classification, VQA, report generation, and genomic risk prediction.The family covers 2D and 3D medical images, genomics, histopathology, ophthalmology, and dermatology.
- Evaluation: The evaluation covers clinically relevant benchmarks across 22 datasets, five task types, six medical image modalities, and eight out-of-distribution datasets.Expert human evaluation is used where clinical judgment is critical, including chest X-ray and CT report generation and open radiology VQA.
- Results: Med-Gemini achieves best-in-class or promising performance in chest X-ray and CT report generation, chest X-ray classification, disease-risk prediction, and several visual question-answering tasks.It also approaches models trained with orders of magnitude more examples on dermatology, histopathology, and ophthalmology classification.
2. Datasets
The study combines public and private medical datasets spanning radiology, pathology, ophthalmology, dermatology, VQA, and genomics. Dataset preparation emphasizes de-identification, patient-level separation where possible, corrected labels, and contamination-resistant evaluation splits.
- Dataset scope: The datasets include de-identified medical images, reports, question-answer pairs, genomic information, and image-caption pairs from public and private sources.They cover radiology, pathology, ophthalmology, dermatology, and genetic risk prediction.
- Radiology: MIMIC-CXR provides 377,110 chest X-ray images from 65,379 patients with corresponding free-text reports.Its images support classification, report generation, and VQA evaluation, while 237,912 training images were used for fine-tuning.
- Ophthalmology: EyePACS converts diabetic-retinopathy lesion labels into captions describing lesion presence or the absence of diabetic-retinopathy-related lesions.The constructed dataset uses 12,976 images with lesions and 3,000 healthy-eye images.
- Data preprocessing: The evaluation uses image-disjoint train, validation, and test splits for all three image types in the revised VQA-Rad setup.The balanced split also roughly equalizes closed versus open questions and question types across anatomical regions.
3. Modeling Methodology
Med-Gemini is fine-tuned from Gemini 1.5 using modality-specific encoders for 2D images, 3D data, and genomics. The pipeline combines multimodal medical data with captioning or VQA objectives and a subsequent instruction-tuning phase.
- Model architecture: Gemini 1.5 was selected because its long context and video encoding support integration of multi-slice images, text, audio, and variable-resolution inputs.Its context window supports up to 1 million tokens.
- Multimodal fine-tuning: Three custom vision encoders were trained separately for 2D modalities, 3D modalities, and genomics.Separate encoders performed better than one encoder across all data formats, and tuning both vision and language components improved visual understanding.
- Model variants: Med-Gemini-2D handles conventional 2D medical images, Med-Gemini-3D processes 3D data such as CT, and Med-Gemini-Polygenic encodes genomic features.Med-Gemini-3D builds on Med-Gemini-2D and synthesizes information across CT slices.
- Genomics: The genomic model was trained to predict eight broad health outcomes from polygenic risk scores projected into 2D representations.The outcomes include coronary artery disease, stroke, type 2 diabetes, glaucoma, chronic obstructive pulmonary disease, rheumatoid arthritis, major depression, and all-cause mortality.
- Instruction tuning: Instruction tuning subsequently refined Med-Gemini’s ability to interpret medical images and signals while following nuanced instructions.This phase used curated multimodal instruction-response pairs.
4. Evaluation and Results
Med-Gemini-2D generally improves medical image classification, VQA, and report-generation performance over Gemini and strong baselines, while Med-Gemini-3D demonstrates end-to-end CT report generation. Results also show domain-specific limitations, including calibration problems, domain-shift sensitivity, and low-quality CT reports.
- Medical image classification: Med-Gemini-2D outperformed Gemini Ultra on most chest X-ray classification labels in-distribution, but performance varied across out-of-distribution datasets.It excelled on some CheXpert tasks while lagging on fracture detection in ChestX-ray14, indicating room to improve under domain shift.
- Medical image classification: Med-Gemini-2D produced robust skin-lesion embeddings and achieved performance close to the specialized Derm Foundation model.Fine-tuning improved the model’s understanding of the skin-lesion embedding space, with F1 and accuracy approaching the specialized model.
- Medical image classification: Med-Gemini-2D consistently outperformed Gemini Ultra on ophthalmology lesion-classification tasks, including specificity of 96.4% versus 39.0% for hard exudates.It nevertheless underperformed a strongly supervised model on anomaly detection, whose training used approximately 200× more labeled data.
- Medical image classification: Question-answering classification remained miscalibrated for microaneurysms and neovascularization tasks, with further work needed to improve calibration.The authors associate this issue with training-data distribution and data-mixing ratios.
- Visual question answering: Med-Gemini-2D outperformed many previous results and Gemini across VQA subsets, including expert-evaluated accuracy of 71.9 on the CXR-only ELIXR split.It also achieved 78.6% accuracy on MIMIC-CXR Yes/No questions and 84.8 accuracy on closed-ended Slake questions.
- Report generation for chest X-rays: Med-Gemini achieved a RadGraph F1-score of 24.4%, improving the previous top-performing model by 3.9% for chest X-ray report generation.RadGraph evaluates both report findings and their relationships to image features.
- Report generation for head/neck CT volumes: Med-Gemini-3D generated reports directly from CT volumes, but only 17% of AI reports were judged equivalent to or better than the original radiologist reports.Correct clinical management would have resulted from 45% of normal-study reports and 57% of abnormal-study reports; errors included missed findings and hallucinations.
- Genomic risk prediction: Med-Gemini-Polygenic achieved higher AUCs than linear PRS benchmarks for all in-distribution outcomes except glaucoma.Its performance often exceeded linear probes when both polygenic risk scores and demographic information were used, although reported AUCs are an upper bound because GWAS features came from UK Biobank data.
5. Qualitative Results
Med-Gemini demonstrates multimodal dialogue, interpretation, and report-generation capabilities across radiology, histopathology, and other medical imaging domains. Expert reviews also identify limitations in phrasing, accuracy, detail, completeness, and volumetric-image reporting.
- Multimodal dialogue: Med-Gemini provides accurate and reasonable multimodal dialogue across chest X-ray, CT, fundus, dermatology, and pathology images.Expert review nevertheless highlights room for improvement in phrasing, accuracy, appropriate detail, and completeness.
- Medical knowledge dialogue: Med-Gemini can answer simple treatment and symptom questions despite fine-tuning focused on image interpretation, but real-world medical guidance is more complex.The pleural-effusion example gives a simple treatment explanation, while the authors frame this capability as a proof of concept requiring further improvement.
- Histopathology: Histopathology responses align with ground truth for invasive breast carcinoma and a non-tumor lymph-node image, while image context limits certainty about grading and tissue origin.The breast-carcinoma response may omit other grade or subtype information, and the lymph-node interpretation is difficult to confirm without additional context.
- Radiology report generation: Med-Gemini generates reasonable chest X-ray reports covering support devices, normal findings, and acute or chronic abnormalities.The examples include reports describing correctly positioned support devices and no acute cardiopulmonary process.
- Radiology report generation: 3D head CT examples include both correct and incorrect abnormal-case reports, with missed abnormalities alongside mischaracterized or hallucinated findings.The examples illustrate both potential value and clinically important failure modes in volumetric report generation.
- Assistive interaction: Region-specific CXR prompting can recover concepts omitted from reports generated without a hint, including emphysema and pulmonary edema.This experiment directs model attention to a specific region or organ as a proof of concept for assistive use.
6. Related Work
Medical language and multimodal models have progressed from general Transformer-based systems to specialized and generalist medical models. The paper situates Med-Gemini within this expanding ecosystem while emphasizing inconsistent evaluation practices as a barrier to direct comparison.
- Evolution of medical language models: Transformer-based language models and scaling methods enabled increasingly large systems such as PaLM, PaLM 2, and PaLM-E.These developments drove progress in natural language processing and multimodal modeling.
- Evolution of medical language models: Medical language models have expanded to applications including clinical text, biomedical knowledge, trial recruitment, and omic information.Examples include PubMedGPT, BioGPT, Med-PaLM, Clinical Camel, MedAlpaca, BioMistral, and models for genomic or cellular data.
- Multimodal models in medicine: General multimodal models such as Flamingo, PaLI, GPT-4, GPT-4v, LLaVA, and Gemini process text and images, with Gemini extending reasoning across video and audio.These capabilities motivated medical applications of generic multimodal foundation models.
- Multimodal models in medicine: Medical vision-language research includes both multi-modality systems and models specialized for domains such as radiology or histopathology.Representative efforts include Med-Flamingo, BiomedCLIP, Med-PaLM M, BiomedGPT, Flamingo-CXR, LLaVA-Med, and PMC-VQA.
- Multimodal models in medicine: Generalist medical AI models are gaining prominence because they aim to handle broad task and modality ranges and support interaction with medical information.The literature also explores orchestrating medical tools with language models.
- Evaluation benchmarks and metrics: Medical VLM evaluation lacks consistency and standardization across tasks, datasets, and metrics, hindering direct comparison even on the same dataset.Recent benchmark efforts respond to this fragmented evaluation landscape.
7. Discussion
The study shows early potential for Med-Gemini across diverse medical tasks and modalities, while emphasizing that clinical safety, evaluation realism, bias mitigation, and contamination control remain unresolved.
- Med-Gemini models were fine-tuned on predominantly medical data paired with free-text reports, reducing the need for expensive expert labeling.
- The study spans multiple modalities and tasks, including 2D and 3D radiology, histopathology, ophthalmology, dermatology, and genetic risk prediction.
- 3D CT report generation remains a proof of concept because large data volume, architectural limitations, and greater clinical complexity constrain current performance.
- Benchmark gains may not translate to bedside usefulness, so future evaluations should assess realistic clinical scenarios, AI-human collaboration, and patient outcomes.
- Safety evaluations must address inherited data biases and errors before real-world deployment, particularly in safety-critical healthcare settings.
- Potential zero-shot generalization may be overstated if large-model training data contain hidden exposure to evaluation examples.
8. Conclusion
Med-Gemini extends Gemini to specialized medical data through fine-tuning across imaging and genomics. The models show promising performance across classification, VQA, report generation, and disease-risk prediction, while requiring further research before clinical use.
- Med-Gemini models are fine-tuned from Gemini on diverse radiology, histopathology, ophthalmology, dermatology, and genomic data.
- Med-Gemini-2D sets a new standard for expert-evaluated chest X-ray report generation, while Med-Gemini-3D demonstrates LMM-based report generation for 3D CT.
- Med-Gemini-2D performs strongly in visual question answering and classification across medical imaging modalities, while Med-Gemini-Polygenic outperforms conventional polygenic risk-score methods.
- Further rigorous research is needed to ensure safe and effective implementation in real-world clinical settings.
- The results support potential for comprehensive systems that integrate multiple capabilities to assist humans with complex multidisciplinary clinical tasks.
9. Contributions and Acknowledgments
The listed contributors represent research leadership and workstreams spanning Google organizations, Verily, Apollo Radiology International, Northwestern Medicine, EyePACS, and other institutions.
- Authors are listed according to their primary workstreams, although many contributed beyond the workstream under which they are named.
- The affiliations include Google Research, Google DeepMind, Verily Life Sciences, Apollo Radiology International, Northwestern Medicine, EyePACS, and additional clinical and research organizations.
Data Availability
Most datasets used for development and evaluation are publicly accessible with appropriate permissions, while some remain private. The authors plan selective data releases and restrict model code and weights because of medical-safety concerns.
- Except for several named datasets, the datasets used in the report are publicly accessible with appropriate permissions.
- The authors intend to release updated classification labels, custom MIMIC-CXR VQA pairs, dataset splits, and replacement question-answer pairs.
- Model code and weights will not be open-sourced because of safety concerns associated with unmonitored use in medical settings.
- The study was funded by Alphabet, and affiliated authors may own Alphabet stock through standard compensation.
- The manuscript was written manually, with a small number of copy edits performed using Gemini.
A.1.1. Revised MIMIC-CXR classification labels
The MIMIC-CXR classification labels were revised because the dataset lacked ground-truth labels. Med-PaLM 2 and board-certified radiologists were used to refine labels for flagged test cases.
- MIMIC-CXR lacks ground-truth labels for classification evaluation.
- Med-PaLM 2 and US-based board-certified radiologists refined labels for flagged test cases.
- Three radiologists reviewed 1,378 flagged labels, with a fourth thoracic radiologist adjudicating disagreements.
- 66% of Med-PaLM 2 labels matched the adjudicated ground truth, compared with 19% for the original labels.
A.1.2. Prompts for VQA and CXR classification evaluations
The evaluation used task-specific prompt formats for CXR classification and VQA. Binary prompts performed better for Med-Gemini classification and were therefore used at evaluation.
- CXR classification prompts: Binary questions were evaluated for Med-Gemini across the top five CXR conditions and the abnormal/normal class.
- CXR classification prompts: Med-Gemini used binary prompts because they yielded better overall macro F1 scores than multi-select prompts on validation data.
- Prompt templates were crafted separately for each model to optimize evaluation performance.
- VQA prompts: Zero-shot VQA prompts replaced a question placeholder with each individual question for every model and evaluation dataset.
A.1.3. New balanced splits for VQA-Rad dataset
The original VQA-Rad split contained substantial image and patient overlap between training and test data. A replacement split was created with disjoint images and patients while balancing question distributions across subsets.
- Contamination in the official split: 202 of 451 VQA-Rad test image IDs also appeared in the training set, creating train/test contamination.
- New disjoint splits: The replacement split uses image IDs to ensure disjoint images and patients across training, validation, and test sets.
- New disjoint splits: The new splits were made roughly equal-sized and assigned remaining former test images and VQAs to validation.
- Balanced question distributions: Open- versus closed-ended question ratios were balanced across splits for each anatomical region.
- Balanced question distributions: Question-type distributions were also made comparable across the new splits.