Source-linked AI summary
MedSAM2: Segment Anything in 3D Medical Images and Videos
Jun Ma, Zongxin Yang, Sumin Kim, Bihui Chen, Mohammed Baharoon, Adibvafa Fallahpour, Reza Asakereh, Hongwei Lyu, Bo Wang
TL;DR
Medical segmentation needs general-purpose models that handle 3D images and videos, but existing medical foundation models are largely 2D-focused and SAM has a domain gap in medical imaging. MedSAM2 fine-tunes SAM2 on a large, diverse medical dataset and uses iterative annotation workflows, achieving stronger segmentation across diverse modalities and reducing annotation time and costs. Its scope is constrained by bounding-box prompts and fixed eight-frame memory, which can limit segmentation or tracking for complex structures and rapid motion.
Problem
Existing medical foundation models are primarily designed for 2D images and do not generally address both 3D spatial relationships and temporal information in medical videos.
Method
MedSAM2 modifies and fine-tunes SAM2 using a diverse dataset of more than 455,000 3D image–mask pairs and 76,000 annotated video frames, with iterative annotation pipelines.
Results
MedSAM2 achieves better performance across various organs and lesions than existing SAM variants and reduces annotation time by up to 92%.
Takeaways & Limitations
More efficient annotation can support scaling large, high-quality labeled datasets for future diagnostic model development and clinical deployment.
Takeaways & Limitations
Bounding-box prompts limit highly complex elongated or branching 3D structures, while fixed eight-frame memory can impair tracking during rapid or large target movements.
Abstract
from arXiv · showhide
Medical image and video segmentation is a critical task for precision medicine, which has witnessed considerable progress in developing task or modality-specific and generalist models for 2D images. However, there have been limited studies on building general-purpose models for 3D images and videos with comprehensive user studies. Here, we present MedSAM2, a promptable segmentation foundation model for 3D image and video segmentation. The model is developed by fine-tuning the Segment Anything Model 2 on a large medical dataset with over 455,000 3D image-mask pairs and 76,000 frames, outperforming previous models across a wide range of organs, lesions, and imaging modalities. Furthermore, we implement a human-in-the-loop pipeline to facilitate the creation of large-scale datasets resulting in, to the best of our knowledge, the most extensive user study to date, involving the annotation of 5,000 CT lesions, 3,984 liver MRI lesions, and 251,550 echocardiogram video frames, demonstrating that MedSAM2 can reduce manual costs by more than 85%. MedSAM2 is also integrated into widely used platforms with user-friendly interfaces for local and cloud deployment, making it a practical tool for supporting efficient, scalable, and high-quality segmentation in both research and healthcare environments.
INTRODUCTION
Medical segmentation supports diagnosis, planning, monitoring, and anatomical analysis, while foundation models have advanced general-purpose 2D segmentation but remain limited for medical imaging. MedSAM2 addresses the lack of general models for both 3D images and videos by adapting SAM2 with large-scale medical data and evaluating annotation workflows.
- Medical image segmentation delineates organs, lesions, and other anatomies for anatomical analysis, disease diagnosis, surgery planning, and treatment monitoring.
- Foundation models shift segmentation from specialist systems toward general-purpose models, but SAM performs suboptimally on medical images because of the domain gap.
- Existing medical foundation models mainly target 2D images, while 3D and video data require modeling spatial relationships and temporal information.
- MedSAM2 adapts SAM2 into a general model for 3D medical image and video segmentation using more than 455,000 3D image–mask pairs and 76,000 annotated video frames.
- Experiments and user studies assess MedSAM2 across volumetric scans and successive video frames while targeting more efficient high-throughput medical dataset annotation.
RESULTS
MedSAM2 combines a diverse 3D image/video dataset with SAM2-based architecture and domain-specific fine-tuning, achieving strong segmentation across volumetric scans and videos. Its human-in-the-loop pipeline progressively reduces annotation time while expanding CT, MRI, and ultrasound datasets.
- Dataset and model: 363,161 CT, 14,818 PET, and 77,154 MRI image-mask pairs, plus 19,232 ultrasound and 56,462 endoscopy frames, comprise the diverse training dataset.The data span normal anatomy, pathologies, and multiple imaging modalities.
- Dataset and model: MedSAM2 uses SAM2’s image encoder, memory attention, prompt encoder, and mask decoder, with full-model fine-tuning of the lightweight SAM2.1-Tiny variant.The image encoder processes slices or frames, while memory attention conditions current features on prior slices or frames.
- 3D image segmentation: 40 holdout segmentation tasks across CT, MRI, and PET scans evaluated MedSAM2 against SAM2.1 variants and EfficientMedSAM-Top1 using middle-slice bounding-box prompts with bidirectional propagation.The evaluation covered diverse organs and lesions.
- 3D image segmentation: MedSAM2 achieved the highest DSC scores across CT organs, CT lesions, MRI organs, MRI lesions, and PET lesions, including 88.84% for CT organs and 88.37% for MRI lesions.The reported scores were 88.84%, 86.68%, 87.06%, 88.37%, and 87.22%, respectively.
- Video segmentation: 96.13% DSC for left ventricle segmentation and 92.22% DSC on hard polyp cases demonstrate stronger and more consistent video performance than SAM2.1.MedSAM2 also achieved 93.10% for left ventricle epicardium and 95.79% for left atrium, with less spread in heart-chamber results.
- Human-in-the-loop annotation: 65.2 seconds per liver MRI lesion and 8.4 seconds per ultrasound frame were reached after three iterative annotation rounds, alongside expansion to 2,490 MRI lesions and 134,591 ultrasound frames.The ultrasound pipeline reported a 92% reduction versus 102.3 seconds of manual annotation and over 12× throughput improvement.
DISCUSSION
MedSAM2 adapts general segmentation foundation models to 3D medical images and videos while supporting large-scale annotation. The authors report improved accuracy and robustness, reduced annotation effort, broad deployment, and several limitations affecting complex structures, fast motion, and low-resource settings.
- Transfer learning substantially improves segmentation accuracy and robustness when adapting general-domain foundation models to diverse medical imaging modalities.
- Annotation time fell by up to 92%, while the iterative pipeline expanded datasets by more than four times.
- Fine-tuning on domain-specific CT and MRI lesion datasets progressively improves annotation efficiency and segmentation quality, including for heterogeneous lesion types.
- Iterative annotation with transfer learning reduces echocardiography annotation time while improving segmentation quality across patient demographics and ultrasound systems.
- MedSAM2 is deployed through multiple platforms and interfaces, supporting practical access for clinical and research users.
- Bounding-box prompts can limit segmentation of thin, branching, elongated, or curved 3D structures because the model does not explicitly model 3D spatial continuity.
- The fixed eight-frame memory bank may reduce tracking performance for rapidly moving or intermittently visible targets.
- GPU-dependent inference limits applicability on edge devices, point-of-care ultrasound machines, and low-power workstations.
METHODS
The methods combine curated public datasets, a modified SAM2 architecture, full-model fine-tuning, and modality-aware training procedures for 3D images and videos.
- The evaluation uses public 3D multi-phase liver tumor CT and MedSAM testing sets spanning CT organs, CT lesions, MRI organs, MRI lesions, and PET lesions.
- DeepLesion contributes diverse CT lesions, with annotation prioritizing lesions having a minimal diameter of 25mm.
- LLD-MMRI provides 3,984 liver-lesion cases across eight MRI scans per lesion, including contrast-enhanced, T2-weighted, diffusion-weighted, and T1 sequences.
- RVENet contains echocardiography videos annotated for cardiac structures, excluding videos with low image quality or incomplete anatomy.
- MedSAM2 modifies SAM2 by reducing input resolution to 512 × 512 and retaining image encoding, prompt encoding, memory attention, and mask decoding components.
- Training initializes from SAM2.1-Tiny and uses full-model fine-tuning with learning rates of 3.0 × 10^-5 for the image encoder and 5.0 × 10^-5 for other components.
- Training combines 3D images and videos with augmentation, modality reweighting, and randomly perturbed bounding-box prompts.
3D Slicer Integration
MedSAM2 is integrated into 3D Slicer as a client-server plugin supporting medical-image loading, prompting, segmentation, refinement, visualization, and local or remote inference.
- The 3D Slicer plugin reuses built-in modules for loading DICOM and NIfTI data, drawing prompts, refining masks, and visualizing 2D and 3D results.
- Users can choose predefined CT or MRI preprocessing options to normalize image intensity before segmentation.
- Users define regions of interest by selecting start and end slices and drawing bounding-box prompts on a key slice.
- Segmentation controls support middle-slice or full-volume inference, customized models, and a Flask server with a temporary most-recently-used cache for refinement.
- Evaluation reports Dice Similarity Coefficient and Normalized Surface Distance with a 2mm boundary tolerance, using 3D scores for images and averaged frame-wise scores for videos.
- The project maintains a public repository of medical image segmentation datasets for long-term community access.
- Code, model weights, annotated datasets, and the 3D Slicer plugin are publicly available.