Source-linked AI summary
Capabilities of Gemini Models in Medicine
Khaled Saab, Tao Tu, Wei-Hung Weng, Ryutaro Tanno, David Stutz, Ellery Wulczyn, Fan Zhang, Tim Strother, Chunjong Park, Elahe Vedadi, Juanma Zambrano Chaves, Szu-Yeu Hu, Mike Schaekermann, Aishwarya Kamath, Yong Cheng, David G. T. Barrett, Cathy Cheung, Basil Mustafa, Anil Palepu, Daniel McDuff, Le Hou, Tomer Golany, Luyang Liu, Jean-baptiste Alayrac, Neil Houlsby, Nenad Tomasev, Jan Freyberg, Charles Lau, Jonas Kemp, Jeremy Lai, Shekoofeh Azizi, Kimberly Kanada, SiWai Man, Kavita Kulkarni, Ruoxi Sun, Siamak Shakeri, Luheng He, Ben Caine, Albert Webson, Natasha Latysheva, Melvin Johnson, Philip Mansfield, Jian Lu, Ehud Rivlin, Jesper Anderson, Bradley Green, Renee Wong, Jonathan Krause, Jonathon Shlens, Ewa Dominowska, S. M. Ali Eslami, Katherine Chou, Claire Cui, Oriol Vinyals, Koray Kavukcuoglu, James Manyika, Jeff Dean, Demis Hassabis, Yossi Matias, Dale Webster, Joelle Barral, Greg Corrado, Christopher Semturs, S. Sara Mahdavi, Juraj Gottweis, Alan Karthikesalingam, Vivek Natarajan
TL;DR
Medical AI must reason over complex, multimodal information while handling long records and current medical knowledge. Med-Gemini specializes Gemini models for medicine and demonstrates strong performance across diverse medical tasks, while requiring further rigorous evaluation before deployment.
Problem
Medical AI needs advanced reasoning over complex multimodal information, long records, and current medical knowledge across clinical work.
Method
Med-Gemini specializes Gemini with medical fine-tuning, web search integration, long-context processing, and custom encoders for novel modalities.
Results
Med-Gemini shows strong performance across 25 tasks spanning 14 medical benchmarks, including multimodal and long-context medical evaluations.
Takeaways & Limitations
The results suggest Med-Gemini could support medical applications including clinical information processing, summarization, dialogue, research, and education.
Takeaways & Limitations
Further rigorous evaluation is needed because medical AI systems face reliability, safety, regulatory, clinical, and equity-related risks before real-world implementation.
Abstract
from arXiv · showhide
Excellence in a wide variety of medical applications poses considerable challenges for AI, requiring advanced reasoning, access to up-to-date medical knowledge and understanding of complex multimodal data. Gemini models, with strong general capabilities in multimodal and long-context reasoning, offer exciting possibilities in medicine. Building on these core strengths of Gemini, we introduce Med-Gemini, a family of highly capable multimodal models that are specialized in medicine with the ability to seamlessly use web search, and that can be efficiently tailored to novel modalities using custom encoders. We evaluate Med-Gemini on 14 medical benchmarks, establishing new state-of-the-art (SoTA) performance on 10 of them, and surpass the GPT-4 model family on every benchmark where a direct comparison is viable, often by a wide margin. On the popular MedQA (USMLE) benchmark, our best-performing Med-Gemini model achieves SoTA performance of 91.1% accuracy, using a novel uncertainty-guided search strategy. On 7 multimodal benchmarks including NEJM Image Challenges and MMMU (health & medicine), Med-Gemini improves over GPT-4V by an average relative margin of 44.5%. We demonstrate the effectiveness of Med-Gemini's long-context capabilities through SoTA performance on a needle-in-a-haystack retrieval task from long de-identified health records and medical video question answering, surpassing prior bespoke methods using only in-context learning. Finally, Med-Gemini's performance suggests real-world utility by surpassing human experts on tasks such as medical text summarization, alongside demonstrations of promising potential for multimodal medical dialogue, medical research and education. Taken together, our results offer compelling evidence for Med-Gemini's potential, although further rigorous evaluation will be crucial before real-world deployment in this safety-critical domain.
1. Introduction
Medical AI must combine clinical communication, uncertainty-aware reasoning, up-to-date knowledge, multimodal interpretation, and long-context analysis, yet current models still face confabulation, bias, tool-use, and collaboration challenges. Med-Gemini addresses these gaps through medical specialization, web-search capabilities, multimodal fine-tuning, and long-context processing, while requiring rigorous validation before deployment.
- Motivation: Clinicians must integrate patient histories, medical images, diagnostics, and current medical information while communicating diagnoses and treatment plans clearly and empathetically.These demands make medicine a multifaceted setting for AI assistance.
- Challenges: Despite strong medical question-answering results, LLMs still show suboptimal reasoning under uncertainty, confabulation, bias, limited tool use, and difficulty accessing up-to-date information.Effective collaboration with clinicians also remains a challenge.
- Med-Gemini: Med-Gemini is a family of medical models fine-tuned from Gemini to advance clinical reasoning, multimodal understanding, and long-context capabilities.The models inherit Gemini’s language, conversational, multimodal, and long-context strengths.
- Methods: Uncertainty-guided web search within an agent framework is introduced to improve factual accuracy, reliability, and nuance in complex clinical reasoning tasks.Web-search use is enhanced through self-training and applied at inference time.
- Applications: Multimodal fine-tuning targets specialized medical modalities, while long-context configurations analyze complicated electronic health records and videos.The introduction also previews real-world evaluations including medical note summarization, referral letter generation, and EHR question answering.
- Limitations: Rigorous validation remains essential because Med-Gemini should support expert clinicians rather than replace them in real-world medical deployment.The paper emphasizes careful consideration of medicine’s nuances before deployment at scale.
2. Methods
Med-Gemini is developed by adapting Gemini models through specialized fine-tuning, self-training with web search, uncertainty-guided inference, and customized multimodal encoders. The methods also extend to clinical reasoning, medical imaging, long-record understanding, and medical video.
- Model development: Med-Gemini-M 1.0 fine-tunes Gemini 1.0 Pro, whereas Med-Gemini-L 1.0 fine-tunes Gemini 1.0 Ultra for advanced reasoning and web-search use.An uncertainty-guided search strategy is introduced at inference time for complex clinical reasoning tasks.
- Fine-tuning data: Fine-tuning uses MedQA-R with synthetic reasoning explanations and MedQA-RS with reasoning and search, supplemented by 260 expert-crafted long-form responses and 65 clinician-written summaries.The additional responses come from HealthSearchQA, LiveQA, and MedicationQA, while summaries come from MIMIC-III medical notes.
- Clinical reasoning: Self-training generates two reasoning paths per question: one without external information and one incorporating retrieved web-search results.The process uses expert demonstrations, regenerates chain-of-thoughts iteratively, and repeats until performance saturates.
- Clinical reasoning: Uncertainty-based search invokes retrieval using an uncertainty measure based on the Shannon entropy of answer-choice distributions across multiple reasoning paths.The method generates multiple reasoning paths before deciding whether to invoke search.
- Multimodal understanding: For medical imaging, Gemini 1.5 Pro is fine-tuned on task-specific instructions using MIMIC-CXR for chest-X-ray report generation and classification of 13 abnormal radiological conditions.MIMIC-CXR pairs chest radiographs with reports and discrete labels derived using the CheXpert labeler.
- Multimodal understanding: Med-Gemini-M 1.5 is designed to identify visual patterns and understand actions and relationships between events across extended time frames.This capability supports multimodal understanding of medical video and temporally extended events.
3. Evaluation
The evaluation spans text-based reasoning, multimodal question answering, long-form clinical text generation, multimodal dialogue, and long-context medical information processing. It uses benchmark-specific metrics, expert comparisons, uncertainty-guided web search, and in-context evaluation across diverse medical tasks.
- Text-based reasoning: Med-Gemini-L 1.0 is evaluated on MedQA, NEJM CPC, and GeneTuring to assess clinical reasoning, diagnosis, and genomic knowledge.MedQA uses prediction accuracy with four iterations of uncertainty-guided search; NEJM CPC uses top-1 and top-10 diagnosis accuracy; GeneTuring uses prediction accuracy with one search iteration and excludes abstentions.
- Long-form text generation: Clinicians compare Med-Gemini-M 1.0 with human experts on medical summarization, referral letter generation, and medical simplification using blinded side-by-side preferences.The tasks use de-identified medical notes and an expert evaluation panel.
- Multimodal evaluation: Seven multimodal VQA benchmarks cover dermatology, radiology, pathology, cardiology, and cross-specialty health-and-medicine performance.The cross-specialty benchmarks are NEJM Image Challenge, USMLE-MM, and MMMU-HM, and are evaluated without multimodal fine-tuning of Med-Gemini-L 1.0.
- Multimodal evaluation: Evaluation metrics include prediction or exact-match accuracy for close-ended VQA and token-level F1 for open-ended Slake-VQA and Path-VQA.The close-ended tasks include NEJM Image Challenge, USMLE-MM, PAD-UFES-20, and MMMU-HM; ECG-QA uses exact-match accuracy.
- Long-context processing: Med-Gemini-M 1.5 is evaluated on long-context EHR reasoning, medical instructional video QA, and Critical View of Safety assessment in surgical videos.The EHR task uses 200 examples and compares one-shot in-context learning with a heuristic annotation-aggregation baseline using precision and recall; CVS concerns laparoscopic cholecystectomy videos.
4. Results
Med-Gemini achieves state-of-the-art results across diverse medical reasoning, multimodal, and long-context evaluations, while showing promising utility in clinical text generation and multimodal dialogue. Its strongest results include 91.1% MedQA accuracy, broad multimodal gains over GPT-4V, and competitive performance on real-world medical tasks.
- Medical reasoning: 91.1% accuracy on MedQA (USMLE) establishes a new SoTA, surpassing Med-PaLM 2 by 4.5% and MedPrompt by 0.9%.The result uses uncertainty-guided general web search.
- Medical reasoning: 13.2% higher top-10 accuracy than AMIE on NEJM CPC demonstrates generalization of uncertainty-guided search to complex diagnosis.Med-Gemini-L 1.0 also outperforms SoTA models on seven GeneTuring modules.
- Medical reasoning: Self-training adds 3.2% accuracy, while successive uncertainty-guided search rounds improve MedQA performance from 87.2% to 91.1%.Relabeling found 3.8% missing-information, 2.9% likely label-error, and 0.7% ambiguous questions, with high inter-rater agreement.
- Real-world utility: Med-Gemini-M 1.0 responses are rated good or better than expert responses more than half the time across three clinical text-generation tasks.These tasks are after-visit summaries, doctor referral letters, and medical simplification.
- Multimodal benchmarks: Across seven multimodal benchmarks, Med-Gemini improves over GPT-4V by an average relative margin of 44.5%.Med-Gemini-L 1.0 exceeds GPT-4V by 8.7%, 13.1%, and 2.6% on NEJM Image Challenge, USMLE-MM, and MMMU-HM, respectively.
- Long-context and dialogue: Med-Gemini-M 1.5 matches a tuned EHR annotation baseline one-shot, achieves SoTA on two MedVidQA tasks, and exceeds GPT-4V by 21% on CVS assessment.A supervised ResNet3D baseline performs better on CVS assessment; multimodal dialogues also demonstrate dermatology diagnosis and radiology reporting capabilities.
5. Discussion
Med-Gemini demonstrates broad advances across medical reasoning, multimodal understanding, and long-context processing, while discussion emphasizes specialization, search integration, and conversational applications. The authors also stress that rigorous evaluation and responsible-AI safeguards are necessary before broader clinical use.
- Core capabilities: Med-Gemini advances clinical reasoning, multimodal understanding, and long-context processing across 25 tasks spanning 14 medical benchmarks.The benchmarks cover medical knowledge, clinical reasoning, genomics, waveforms, medical imaging, health records, and videos.
- Web search integration: Web search integration may improve factual reliability, but its safety and effectiveness require significant further research.The model was trained to search when uncertain and incorporate results into responses; the discussion notes unresolved restrictions and limitations.
- Multimodal conversational capabilities: Med-Gemini-M 1.5 shows promising multimodal, multi-turn clinical dialogue without medical dialogue fine-tuning.Qualitative examples include requesting additional images when needed and explaining responses.
- Long-context processing: Long-context processing opens new medical-AI applications, including retrieval and verification of conditions, symptoms, and procedures within very long electronic patient records.The paper introduces a “needle-in-a-haystack” EHR task intended to reflect a real-world challenge.
- Medical specialization and fine-tuning: Medical specialization and fine-tuning remain important because medical knowledge and multimodal data are unique and complex despite strong out-of-the-box capabilities.The discussion highlights strong performance on multimodal benchmarks while emphasizing the distinctive nature of medical data.
- Evaluation and responsible AI: Further evaluation should extend beyond 14 benchmarks and incorporate fairness, privacy, equity, transparency, accountability, and sociotechnically grounded risk assessment.The authors identify potential interactions among dataset bias, model bias, search integration, and long-context use, while calling for robust evaluation in specific clinical settings.
6. Conclusion
Gemini and Med-Gemini demonstrate broad opportunities to advance biomedical discovery and assist healthcare delivery. Realizing these benefits requires meticulous attention to model reliability and safety.
- Conclusion: Gemini and Med-Gemini suggest a significant leap forward in opportunities to accelerate biomedical discoveries and assist in healthcare delivery and experiences.The conclusion emphasizes both the depth and breadth of these opportunities.
- Conclusion: Advancing model capabilities must be accompanied by meticulous attention to the reliability and safety of these systems.The passage presents reliability and safety as paramount considerations for progress in health and medicine.
9. Code Availability … C. Additional details on advanced reasoning text-based tasks
The paper withholds model code and weights because of safety concerns, while outlining responsible validation and planned Google Cloud API access. Its related-work discussion situates medical LLMs, reasoning and tool use, multimodal generalist systems, and the underexplored need for long-context medical evaluation.
- 9. Code Availability: Model code and weights are not open-sourced because unmonitored use in medical settings poses safety implications.
- 9. Code Availability: The authors plan to work with research partners, regulators, and providers to validate safe onward uses and eventually offer the models through Google Cloud APIs.
- A. Supplementary Table for Figure 1: The supplementary table aggregates Med-Gemini results against prior SoTA and best GPT-4 methods across text-based, multimodal, and long-context tasks.
- B. Related Works: Large language models use architectures such as transformers and pathways, trained through self-supervision on massive, diverse datasets.
- B. Related Works: Medical LLMs such as Med-PaLM and Med-PaLM 2 are fine-tuned on EHRs, exam questions, and research literature, while generalist systems use prompting or multimodal refinement.
- C. Additional details on advanced reasoning text-based tasks: Reasoning improvements combine model advances, human-reasoning imitation, prompt engineering, improved processes, and external tools or knowledge.
- B. Related Works: Medical multimodal models integrate patient history, imaging, genetic testing, and laboratory results; specialist and generalist approaches target different breadths of application.
- B. Related Works: Most medical LLM evaluations still use relatively short texts and single images, motivating investigation of Med-Gemini on long-context video and EHR-related use cases.
C.1. Text-based fine-tuning & evaluation datasets
This section describes the datasets used for text-based instruction fine-tuning and the evaluation benchmarks for text-based reasoning. The fine-tuning mixture and synthetic data were curated to improve Med-Gemini-L 1.0’s reasoning and web-search use.
- Text-based fine-tuning datasets: The text-based instruction fine-tuning datasets comprise a curated mixture that includes synthetic data.The mixture was designed for Med-Gemini-L 1.0.
- Text-based fine-tuning datasets: The dataset mixture and synthetic data were curated to improve Med-Gemini-L 1.0’s reasoning and ability to use web search.
- Text-based evaluation datasets: The text-based reasoning evaluation uses a set of evaluation benchmarks summarized in an overview table.
C.2. MedQA (USMLE) Relabeling
The MedQA relabeling study used a two-step physician-rating protocol to identify missing information, label errors, and ambiguous questions while minimizing ground-truth bias. It recruited 18 US primary care physicians and found that ambiguity was harder to assess than missing information or label errors.
- Study design: The two-step protocol withheld MedQA’s ground-truth answer while raters assessed missing information, then revealed it to assess label errors and ambiguity.This timing was intended to prevent ground-truth bias in missing-information judgments while enabling disagreement with the answer key.
- Study design: 18 US primary care physicians rated each MedQA question, with at least three independent ratings collected per question.Physicians averaged 255 seconds per question, and 98% took less than 10 minutes.
- Annotation agreement: Average overlap agreement was around 75% after raters saw MedQA’s ground truth, versus around 50% when it was not revealed.Agreement for identifying missing information and label errors was high, whereas answer overlap was substantially lower without the ground truth.
- Annotation agreement: Questions with missing information or label errors could be identified with high confidence, but judging whether multiple answer options were correct was more difficult.The study defined ambiguity as allowing multiple answer options to be correct, despite many questions requesting the best, most likely, or most appropriate option.
C.3. Additional results on NEJM clinico-pathological conference dataset
The NEJM clinical pathology cases are evaluated with Top-1 and Top-10 performance broken down by primary specialty, focusing on specialties with at least 10 cases. In most reported specialties, including Internal Medicine, Pediatrics, and Psychiatry, the best results come from either Med-Gemini-L 1.0 without search or another compared model.
- Performance breakdown: Top-1 and Top-10 performance is broken down by NEJM-identified primary specialty for all specialties with at least 10 cases.The results are reported in Table C4.
- Specialty results: In most specialties, the best Top-1 and Top-10 performance is achieved by either Med-Gemini-L 1.0 without search or another compared model.The passage specifically names Internal Medicine, Pediatrics, and Psychiatry among these specialties.
- Performance breakdown: Table C4 reports performance on NEJM case studies by specialty.Only specialties with at least 10 cases are included.
C.4. Real-world use cases for advanced reasoning on text-based tasks · D. Additional details on multimodal understanding tasks
Med-Gemini-M 1.0 is evaluated on three real-world long-form text-generation tasks: after-visit summaries, referral letters, and plain-language summaries of biomedical systematic reviews. These evaluations compare model outputs with clinician or expert references using blinded physician or clinician ratings across task-specific quality criteria.
- C.4. Real-world use cases for advanced reasoning on text-based tasks: Three real-world tasks evaluate Med-Gemini-M 1.0 on long-form text generation: after-visit summaries, referral letters, and simplified systematic-review summaries.The evaluation is summarized in Figure 5, with additional axes reported in Table C5.
- C.4. Real-world use cases for advanced reasoning on text-based tasks: 31 de-identified H&P notes support the after-visit-summary task, with expert summaries written and refined by U.S.-based clinicians.The model receives a de-identified H&P note and generates a structured summary for the patient.
- C.4. Real-world use cases for advanced reasoning on text-based tasks: The after-visit-summary prompt requires 12 patient-centered fields and instructs the model to use sixth-grade language while avoiding abbreviations and jargon.The fields include visit reasons, diagnoses, measurements, medication changes, next steps, urgent-care conditions, and provider comments.
- C.4. Real-world use cases for advanced reasoning on text-based tasks: After-visit summaries are assessed by U.S.-based physicians for accuracy, coverage, coherence, succinctness, and overall quality.The comparison presents the H&P note, clinician-generated summary, and model-generated summary to blinded physician raters.
- C.4. Real-world use cases for advanced reasoning on text-based tasks: Referral-letter evaluation uses clinician-selected outpatient notes containing referral recommendations and compares model-generated letters with clinician-generated letters.Three U.S. board-certified physicians evaluate all 25 examples, with Likert ratings mapped to [-2,2] and aggregated using the median sign.
- C.4. Real-world use cases for advanced reasoning on text-based tasks: 25 Cochrane technical abstracts and expert plain-language summaries support evaluation of medical simplification for lay audiences.The model generates a plain-language summary from each technical abstract.
- C.4. Real-world use cases for advanced reasoning on text-based tasks: Clinicians compare simplified summaries on grounding, coverage, succinctness, reading level, and overall quality, with ratings aggregated as in the referral-letter task.The comparisons are conducted with the technical abstract, original plain-language summary, and model-generated summary while sources are blinded.
D.1. Multimodal fine-tuning datasets
Med-Gemini’s multimodal fine-tuning uses diverse image-text datasets spanning radiology, pathology, dermatology, and ECG signals. The datasets support visual question answering, classification, captioning, report generation, and sensor-text health-signal encoding.
- Dataset composition: Med-Gemini-M 1.5 uses Slake-VQA, Path-VQA, MIMIC-CXR, PAD-UFES-20, and ROCO for multimodal fine-tuning.Med-Gemini-S 1.0 additionally uses a subset of ECG-QA to develop a health signal encoder for sensor input.
- Radiology datasets: MIMIC-CXR contains 377110 chest X-ray images and PHI-removed reports from 65379 patients, with reports annotated for 13 common radiological conditions.The dataset supports normal-versus-abnormal classification, abnormality-condition VQA, synthetic CXR VQA, and text report generation.
- Dermatology datasets: PAD-UFES-20 supports 6-class classification using images and 14 clinical features, with variants using augmentation and upsampled subsets.The described features include demographics, history, lesion characteristics, and symptom indicators.
- Pathology datasets: Path-VQA contains 998 pathology images with 32799 QA pairs, including 50.2% open-ended and 49.8% close-ended questions.Questions address pathology-image attributes such as color, location, appearance, and shape.
- Captioning and health signals: ROCO provides 29907 image-caption pairs across radiology and non-radiology, restricted to specified Creative Commons licenses for captioning fine-tuning.ECG-QA evaluates cardiac health from waveform signals using single-ECG and paired-ECG question types; this work focuses on single ECG questions.
D.2. Multimodal evaluation datasets
The multimodal evaluation uses three out-of-distribution medical datasets: NEJM Image Challenge, USMLE-MM, and MMMU-HM. Together, they cover image-based clinical diagnosis, multimodal USMLE preparation questions, and health-and-medicine topics across several medical domains.
- NEJM Image Challenge: NEJM Image Challenge images span radiography, skin imaging, electrocardiograms, histopathology, endoscopy, and ophthalmoscopy.The challenge tests diagnostic acumen and visual observation using clinical cases presented weekly.
- NEJM Image Challenge: 942 NEJM Image Challenge cases were collected from 2005 to 2023, with 934 ultimately evaluated for fair comparison.Each case presents a clinical image and brief description, requiring selection of a final diagnosis from five candidates.
- USMLE-MM: 46 USMLE-MM questions were identified from sample exams and include images within the questions.The dataset supports multimodal evaluation in USMLE preparation materials.
- MMMU-HM: 150 MMMU-HM questions cover basic medical science, clinical medicine, diagnostics and laboratory medicine, pharmacy, and public health.MMMU-HM is drawn from the publicly available MMMU validation set.
D.3. Additional results for multimodal tasks · E. Additional details on long-context understanding tasks
The section describes Med-Gemini’s ECG-QA setup, which adds an ECG-specific encoder to Gemini 1.0 Nano and compares frozen and unfrozen training strategies with baseline counterparts. The supplied passage does not provide the corresponding long-context task details or quantitative results.
- D.3. Additional results for multimodal tasks: Med-Gemini extends Gemini 1.0 Nano for ECG-QA by adding an ECG-specific encoder and fine-tuning it for raw biomedical signals.The passage introduces ECG-QA as an expansion of Med-Gemini’s ability to process raw biomedical signals.
- D.3. Additional results for multimodal tasks: Two ECG-QA training approaches keep the Gemini model frozen or fine-tune it jointly.These are explicitly described as frozen and unfrozen configurations.
- D.3. Additional results for multimodal tasks: The ECG-QA experiments use Med-Gemini-S 1.0 as the evaluated model.The passage names Med-Gemini-S 1.0 when introducing the comparisons.
- D.3. Additional results for multimodal tasks: The frozen-Gemini configuration is compared with GPT-4 using SE-WRN ECG features in input prompts.The comparison baseline is attributed to Oh et al. (2023).
- D.3. Additional results for multimodal tasks: The unfrozen-Gemini configuration is compared with an ECG foundation model.The supplied passage ends before naming or evaluating that foundation model.
- E. Additional details on long-context understanding tasks: No substantive findings for the long-context understanding tasks are included in the supplied passage.The input identifies the merged long-context subsection but provides no passage text for it.
E.1. Long-context evaluation datasets
The long-context evaluation uses curated datasets spanning subtle condition retrieval from long EHRs and temporal localization or surgical-video understanding. These datasets provide clinically grounded annotations and comparison settings for evaluating Med-Gemini’s retrieval and multimodal reasoning capabilities.
- MIMIC-III Needle-in-a-Haystack: MIMIC-III Needle-in-a-Haystack evaluates subtle medical-condition search and retrieval over long, unstructured EHRs from ICU patients.The dataset is curated from MIMIC-III and randomly samples medical notes from 44 unique patients.
- MIMIC-III Needle-in-a-Haystack: The MIMIC-III task contains 121 positive and 79 negative examples, with majority-vote labels serving as ground truth.Inter-rater agreement was Krippendorff’s alpha 0.77.
- Medical video datasets: Medical video evaluation includes timestamp-span localization of visual answers, using official splits of 2710 training, 145 validation, and 155 testing questions.The videos have a mean duration of 383.29 seconds.
- Medical video datasets: Cholec80-CVS provides surgeon annotations for 572 video segments, scoring three critical-view-of-safety criteria from 0 to 2.Med-Gemini-M 1.5 is compared with GPT-4V and Resnet3D on this task.
E.2. Rater agreement metrics for the long EHR understanding task
The long EHR understanding benchmark used ratings from three independent raters across 200 questions and showed strong inter-rater consistency. Agreement was substantial for unanimous selections, stronger when at least two raters agreed, and good for identifying medical conditions.
- Rater agreement metrics: Three independent raters evaluated each of the 200 example questions to assess long EHR benchmark reliability.The agreement analysis was designed to demonstrate consistency among raters.
- Rater agreement metrics: 0.83 Jaccard similarity indicated substantial agreement for unanimous selections.This metric measures overlap among the sets of selections made by the three raters.
- Rater agreement metrics: 0.915 Jaccard similarity reflected strong consistency when at least two of three raters made the same selections.The at-least-two-raters measure captures agreement even without unanimity.
- Rater agreement metrics: 0.77 Krippendorff’s alpha indicated good agreement on the existence of medical conditions in the EHR data.Krippendorff’s alpha is a reliability coefficient designed for multiple raters.