Source-linked AI summary
Project Imaging-X: A Survey of 1000+ Open-Access Medical Imaging Datasets for Foundation Model Development
Zhongying Deng, Cheng Tang, Ziyan Huang, Jiashi Lin, Ying Chen, Junzhi Ning, Chenglong Ma, Jiyao Liu, Wei Li, Yinghao Zhu, Shujian Gao, Yanyan Huang, Sibo Ju, Yanzhou Su, Pengcheng Chen, Wenhao Tang, Tianbin Li, Haoyu Wang, Yuanfeng Ji, Hui Sun, Shaobo Min, Liang Peng, Feilong Tang, Haochen Xue, Rulin Zhou, Chaoyang Zhang, Wenjie Li, Shaohao Rui, Weijie Ma, Xingyue Zhao, Yibin Wang, Kun Yuan, Zhaohui Lu, Shujun Wang, Jinjie Wei, Lihao Liu, Dingkang Yang, Lin Wang, Yulong Li, Haolin Yang, Yiqing Shen, Lequan Yu, Xiaowei Hu, Yun Gu, Yicheng Wu, Benyou Wang, Minghui Zhang, Angelica I. Aviles-Rivero, Qi Gao, Hongming Shan, Xiaoyu Ren, Fang Yan, Hongyu Zhou, Haodong Duan, Maosong Cao, Shanshan Wang, Bin Fu, Xiaomeng Li, Zhi Hou, Chunfeng Song, Lei Bai, Yuan Cheng, Yuandong Pu, Xiang Li, Wenhai Wang, Hao Chen, Jiaxin Zhuang, Songyang Zhang, Huiguang He, Mengzhang Li, Bohan Zhuang, Zhian Bai, Rongshan Yu, Liansheng Wang, Yukun Zhou, Xiaosong Wang, Xin Guo, Guanbin Li, Xiangru Lin, Dakai Jin, Mianxin Liu, Wenlong Zhang, Qi Qin, Conghui He, Yuqiang Li, Ye Luo, Nanqing Dong, Jie Xu, Wenqi Shao, Bo Zhang, Qiujuan Yan, Yihao Liu, Jun Ma, Zhi Lu, Yuewen Cao, Zongwei Zhou, Jianming Liang, Shixiang Tang, Qi Duan, Dongzhan Zhou, Chen Jiang, Yuyin Zhou, Yanwu Xu, Jiancheng Yang, Shaoting Zhang, Xiaohong Liu, Siqi Luo, Yi Xin, Chaoyu Liu, Haochen Wen, Xin Chen, Alejandro Lozano, Min Woo Sun, Yuhui Zhang, Yue Yao, Xiaoxiao Sun, Serena Yeung-Levy, Xia Li, Jing Ke, Chunhui Zhang, Zongyuan Ge, Ming Hu, Jin Ye, Zhifeng Li, Yirong Chen, Yu Qiao, Junjun He
TL;DR
Medical imaging lacks large, diverse, unified datasets because public collections are small, fragmented, and costly to curate under clinical, annotation, ethical, and privacy constraints. This paper surveys over 1,000 open-access datasets, analyzes their coverage, and introduces metadata-driven fusion with an interactive portal; it finds a long-tailed, imbalanced landscape and identifies dataset engineering boundaries that require careful integration and re-annotation strategies.
Problem
Public medical imaging datasets are predominantly small, task-specific, and fragmented, limiting the available scale and diversity for medical foundation-model development.
Method
The paper surveys over 1,000 datasets with a taxonomy and gap analysis, then proposes MDFP and an interactive portal for dataset search, analysis, and integration.
Results
The surveyed landscape is long-tailed and imbalanced, with skew across image dimensionality, modalities, organs, and tasks; video datasets are 85.9% endoscopy.
Takeaways & Limitations
Scaling medical imaging corpora requires normalization, balanced sampling, and metadata-guided integration rather than naive concatenation of datasets.
Takeaways & Limitations
Existing datasets often target indirect tasks, while adapting them to clinically oriented tasks requires resource-intensive radiologist re-annotation that scales poorly.
Abstract
from arXiv · showhide
Foundation models have demonstrated remarkable success across diverse domains and tasks, primarily due to the thrive of large-scale, diverse, and high-quality datasets. However, in the field of medical imaging, the curation and assembling of such medical datasets are highly challenging due to the reliance on clinical expertise and strict ethical and privacy constraints, resulting in a scarcity of large-scale unified medical datasets and hindering the development of powerful medical foundation models. In this work, we present the largest survey to date of medical image datasets, covering over 1,000 open-access datasets with a systematic catalog of their modalities, tasks, anatomies, annotations, limitations, and potential for integration. Our analysis exposes a landscape that is modest in scale, fragmented across narrowly scoped tasks, and unevenly distributed across organs and modalities, which in turn limits the utility of existing medical image datasets for developing versatile and robust medical foundation models. To turn fragmentation into scale, we propose a metadata-driven fusion paradigm (MDFP) that integrates public datasets with shared modalities or tasks, thereby transforming multiple small data silos into larger, more coherent resources. Building on MDFP, we release an interactive discovery portal that enables end-to-end, automated medical image dataset integration, and compile all surveyed datasets into a unified, structured table that clearly summarizes their key characteristics and provides reference links, offering the community an accessible and comprehensive repository. By charting the current terrain and offering a principled path to dataset consolidation, our survey provides a practical roadmap for scaling medical imaging corpora, supporting faster data discovery, more principled dataset creation, and more capable medical foundation models.
1 Introduction
Medical imaging foundation models require large, diverse datasets, but public medical imaging data remain fragmented and difficult to assemble. This survey catalogs the landscape, identifies coverage gaps, and proposes metadata-driven integration to support larger resources.
- Motivation: Public medical datasets are typically small and narrowly scoped, while expert annotation and ethical constraints make large, diverse collections difficult to construct.This leaves medical imaging data substantially less scalable than large natural-image corpora.
- Motivation: Existing dataset-integration efforts merge collections sharing modalities, anatomies, or tasks, but often focus on specific imaging types or organ systems.Without a comprehensive overview, integration can reinforce existing biases instead of supporting balanced foundation-model development.
- Motivation: Prior surveys often omit subject- and image-level statistics, recently released large-scale datasets, and systematic links between dataset characteristics and foundation-model requirements.These omissions limit their usefulness for identifying coverage gaps and evaluating integration opportunities.
- Approach: The survey covers over 1,000 open-access datasets from 2000 to 2025 using a taxonomy spanning modality, anatomy, task, and label availability.It uses the taxonomy for gap analysis and prioritization of future dataset creation.
- Approach: The metadata-driven fusion paradigm integrates existing datasets, while an interactive portal supports fine-grained search, statistical analysis, and dataset integration.The authors also release surveyed dataset information, a Python toolkit, and a merged large-scale dataset.
- Paper organization: The paper surveys modality-specific 2D, 3D, and video datasets, applies integration strategies, and discusses challenges for foundation-model development.The organization moves from a broad landscape overview to modality-focused analyses and integration opportunities.
2 An Overview of Medical Image Datasets
The survey organizes over 1,000 open-access medical imaging datasets by dimensionality, modality, task, and anatomy, then analyzes their temporal growth and distribution. The resulting landscape is long-tailed, with substantial skew toward 2D images, selected modalities and organs, and classification and segmentation tasks.
- Dataset scope and organization: Over 1,000 datasets released between 2000 and 2025 were compiled from public repositories and challenge sites, deduplicated, manually verified, and metadata-normalized.The resulting manifest supports the analyses of dataset growth and distributions.
- Dataset scope and organization: Datasets are grouped by imaging dimensionality, then modality, task, and anatomical region to support systematic coverage analysis.The taxonomy includes 2D, 3D, and video data and aligns dimensions, modalities, tasks, and organs with foundation-model training needs.
- Total growth: Dataset releases show clear inflections after 2012 and another surge after 2023, although existing resources remain orders of magnitude smaller than large general-domain datasets, particularly for 3D volumes.Examples include AbdomenAtlas with 1.5 million 2D CT images and 5,195 3D CT volumes, and CT-RATE with 25,692 3D chest CT scans from 21,304 patients.
- Imaging dimensionalities: 2D images dominate absolute scale, while 3D volumes and videos grow more slowly despite providing volumetric or temporal context that can be clinically informative.Higher acquisition, storage, curation, and annotation burdens constrain 3D and video availability.
- Imaging modalities: Pathology, X-ray, CT, and MRI are prominent modalities, whereas PET, mammography, and endoscopy remain comparatively underrepresented in open data.Pathology image counts are amplified because whole-slide images are divided into thousands of patches; MRI accounts for about 10.4% of total images.
- Tasks: Classification and segmentation dominate task-wise image counts, generation rises sharply after 2023, and registration, detection, and tracking remain comparatively small.The imbalance reflects label economics and acquisition or annotation burdens, including temporal labels for tracking and difficult ground truth for registration.
- Anatomical regions: Brain, lung, liver, breast, retina, colon, and whole-body regions receive substantial attention, while many other anatomical regions remain less represented.Brain and lung contributed the largest image volumes before 2023, followed by pronounced post-2023 growth in brain, liver, lung, and breast data.
- Summary: The open medical imaging landscape is long-tailed, so scaling general-purpose models requires normalized counting, balanced sampling, and task-aware objectives rather than naïve concatenation.These measures are intended to avoid amplifying existing modality, organ, and task biases.
3 2D Medical Image Datasets
The survey collected 502 two-dimensional medical image datasets, with labeled datasets showing pronounced imbalance across modalities, anatomies, and tasks. Image volume is concentrated in a few modalities, organs, and common tasks, while several clinical areas and specialized tasks remain sparsely represented.
- Overview: 502 2D medical image datasets were collected, partitioned into 475 labeled and 45 unlabeled datasets.Labeled datasets were analyzed by modality, task, and anatomy; unlabeled datasets were summarized by modality and anatomy.
- Overview: 2D labeled datasets exhibit long-tail distributions across modalities, anatomical structures, and tasks.Figure 11 reports both percentages and image counts for each distribution.
- Modality and anatomy: Pathology and X-ray dominate image volume, followed by CT, MRI, and fundus photography.Endoscopy and other modalities are less representative.
- Modality and anatomy: Full-body views, retina, breast, and brain are comparatively prominent, whereas uterus, heart, esophagus, limb joints, and small substructures remain underrepresented.The survey identifies these less-covered anatomical areas as opportunities for targeted curation.
- Tasks: Generation, classification, segmentation, regression, and detection dominate tasks, while registration, tracking, localization, reconstruction, and visual question answering have far fewer images.The task distribution is therefore heavily concentrated in a limited set of common objectives.
- CT task distribution: 513,900 images across 12 datasets comprise the classification category, the largest CT task group by image volume.Detection/localization is dominated by one large dataset, while segmentation and reconstruction are smaller.
3.3 MRI Slices
MRI datasets provide rich multi-contrast information for neurological, musculoskeletal, and oncological imaging, but their distribution is highly uneven. Image volume is dominated by a large multi-structure classification collection, while segmentation and region-specific datasets are much smaller.
- MRI characteristics: MRI’s multi-contrast sequences support tumor segmentation and tissue characterization without ionizing radiation.Common sequences include T1-weighted, T2-weighted, and FLAIR images.
- Anatomical regions: Brain MRI coverage consists of 2 datasets and 220 images, while abdomen/pelvis coverage comprises 4 datasets and approximately 560 images.These region-specific collections primarily support segmentation and cancer-related classification tasks.
- Anatomical regions: 704,900 images across 8 full-body or multistructure datasets dominate MRI anatomical coverage.RadImageNet (Subset: MR) contributes 673,000 images, while ImageCLEF 2016 contributes 31,000.
- Tasks and coverage: Fourteen additional MRI datasets contain approximately 25,700 images across other or unspecified tasks and anatomical regions.Several datasets lack explicit task labels, leaving them flexible for exploratory research.
- Tasks: Classification accounts for 704,000 MRI images across 3 datasets, largely because RadImageNet (Subset: MR) contributes 673,000 images.This concentration makes classification the dominant MRI task by image volume.
- Tasks: MRI segmentation includes 6 datasets totaling approximately 8,900 images, substantially fewer than classification datasets.These datasets cover cardiac, brain, and abdominal structures and require precise soft-tissue delineation.
3.4 PET Slices
PET datasets are relatively concentrated in brain and abdominal imaging, with classification supplying most images and many collections lacking explicit task or anatomical labels. Ultrasound datasets cover broader anatomy and several task types, but their scale is dominated by a few large collections and some remain unlabeled.
- PET overview: 13 PET datasets contain approximately 41,942 images, focusing mainly on brain and abdominal imaging for segmentation and classification.PET collections commonly combine PET with CT or MRI for anatomical localization.
- PET anatomy: Brain-related PET datasets contain approximately 269 images, limiting their suitability for training large-scale deep learning models.These datasets address cerebral microbleed detection and segmentation.
- PET anatomy: PET full-body or multistructure coverage totals approximately 31,200 images, dominated by ImageCLEF 2016.Many other PET datasets have unspecified anatomy and limited image counts.
- PET tasks: 31,000 PET classification images come from one dataset, while pure segmentation contributes only 255 images from one dataset.This illustrates the strong task imbalance within PET data.
- Ultrasound overview: 19 ultrasound datasets contain approximately 457,663 images, with RadImageNet-US contributing 390,000 images.Ultrasound datasets address segmentation, classification, measurement, tracking, estimation, and reconstruction.
- Ultrasound anatomy: Ultrasound coverage spans breast, skull, heart, thyroid, liver, full-body, and multistructure imaging, but RadImageNet-US dominates with 390,000 full-body images.Its commercial license may restrict accessibility.
- Ultrasound limitations: Three ultrasound datasets are unlabeled, including APOLLO-5 with 6,203 images and AREN0532 with 1,021 images.CMB-LCA has no images available; the collections may support future multimodal fusion studies.
3.6 X-Ray Images
X-ray datasets are large and clinically versatile, but their coverage is concentrated in thoracic imaging and classification, with fewer resources for other anatomies and precise localization tasks.
- Anatomical coverage: X-ray collections contain approximately 46% chest/lung-focused datasets, reflecting their widespread use in respiratory disease screening.Thorax/lung collections include NIH Chest X-ray 14 and CheXpert, among the largest listed datasets.
- Anatomical coverage: Breast imaging includes approximately 248,300 images across three datasets, while musculoskeletal imaging includes eight datasets totaling approximately 15,681 images.VICTRE dominates the breast category but lacks disease annotations; musculoskeletal datasets commonly address fracture detection and classification.
- Task distribution: 31% of X-ray collections provide pixel-level annotations or detection labels, indicating that precise localization resources remain a minority.The survey reports 19 of 61 collections with these annotation types.
- Task distribution: 30 classification datasets contain approximately 670,100 images, making classification the largest X-ray task group.Major collections include CheXpert, NIH Chest X-ray 14, and RANZCR CLiP.
- Task distribution: 10 segmentation datasets contain approximately 708,500 images, dominated by CheXmask with 676,800 images.Other segmentation resources target pneumothorax, lung abnormalities, and clavicle identification.
- Task distribution: 9 detection/localization datasets contain approximately 59,700 images and support applications including surgical planning and cephalometric landmark localization.DENTEX, CL-Detection2023, and CEPHA29 exemplify this smaller task group.
3.7 Optical Coherence Tomography (OCT) Images
OCT datasets are overwhelmingly specialized for retinal imaging, with classification and segmentation dominating the task landscape and segmentation providing pixel-level anatomical annotations.
- Anatomical coverage: Almost all OCT datasets focus on retinal applications, with MedMNIST the only listed exception also applicable to breast and lung imaging.This specialization reflects OCT’s primary clinical use for retinal-layer imaging.
- Task distribution: Segmentation accounts for approximately 50% of OCT datasets, emphasizing pixel-level retinal-layer annotations.These datasets support precise anatomical analysis and thickness measurements.
- Task distribution: Five classification datasets contain approximately 210,200 images, predominantly targeting diabetic retinopathy and glaucoma detection.OCT2017 and Retinal OCT-C8 are major collections, while MedMNIST contributes 100,000 images across multiple modalities.
- Task distribution: Eleven segmentation datasets contain approximately 2,600 images and emphasize retinal-layer delineation and lesion segmentation.SinaFarsiu-009 and SinaFarsiu-018 provide the largest listed annotation collections.
- Task distribution: Three APTOS datasets total 8,500 images for diabetic-retinopathy severity prediction using the International Clinical Diabetic Retinopathy scale.These datasets support quantitative disease-progression monitoring.
3.9 Dermoscopy Images
Dermoscopy datasets are concentrated on skin-lesion analysis, with classification dominating image volume and segmentation supporting precise lesion-boundary delineation; histopathology datasets add large-scale, patch-based tissue analysis but face WSI-scale challenges.
- Anatomical coverage: Thirteen dermoscopy datasets contain approximately 133,600 skin images, while ImageCLEF2016 also spans skin, cell, and breast imaging.The largest skin-focused collections include Monkeypox, ISIC20, and ISIC19.
- Dermoscopy tasks: Ten dermoscopy classification datasets contain approximately 157,500 images, compared with five segmentation datasets containing approximately 9,400 images.Classification collections support multi-class lesion categorization, whereas segmentation collections target lesion boundaries.
- Dermoscopy tasks: Vitiligo is the only listed unlabeled dermoscopy dataset, with 368 images potentially useful for unsupervised learning.Most other collected dermoscopy datasets provide labels suitable for supervised learning.
- Histopathology characteristics: Whole Slide Images can exceed 100,000×100,000 pixels, making direct processing infeasible and motivating patch extraction or tiling.Histopathology pipelines aggregate patch-level predictions into slide-level diagnoses.
- Histopathology tasks: Recent histopathology datasets show increased WSI adoption, with WSI-based collections comprising 32% of recent datasets.Multi-task collections increasingly combine segmentation with classification or counting.
- Histopathology tasks: Histopathology task distributions include 38 classification datasets with approximately 709,000 images and 31 segmentation datasets with approximately 368,000 images.Fine-grained classification remains difficult because of intra-class variation and inter-class similarity at the cellular level.
3.12 Infrared Imaging
Infrared medical imaging is highly specialized in retinal applications and is dominated by classification, while endoscopy datasets span gastrointestinal anatomy and increasingly support multitask video-oriented benchmarks.
- Infrared imaging: 100% of infrared datasets focus on retinal applications, comprising six datasets and approximately 424,532 images.All were created since 2018, suggesting growing interest in this modality.
- Infrared imaging: Five infrared classification datasets contain approximately 424,490 images, while one segmentation dataset contains 42 images.The MRL Eye series covers glasses, eye state, reflections, image quality, and sensor-type classification; RAVIR is the sole segmentation dataset.
- Endoscopy coverage: Endoscopy collections contain approximately 322,200 images and videos across 41 datasets, with 39/41 datasets featuring endoscopic imaging.EndoSlam is the largest collection with 76,837 images, and 16/41 datasets contain over 1,000 images.
- Endoscopy coverage: Six multi-structure gastrointestinal datasets contain approximately 86,000 images, including EndoSlam’s coverage of the esophagus, stomach, and colon.These collections are positioned as resources for developing generalizable endoscopic AI systems.
- Endoscopy tasks: Seventeen endoscopy segmentation datasets contain approximately 20,000 images, while six detection datasets contain approximately 86,600 images.Segmentation commonly targets polyps, whereas detection resources include abnormalities and surgical instruments.
- Endoscopy tasks: Five endoscopy multitask datasets contain 156,000 images and combine annotations such as captioning, classification, localization, and segmentation.HyperKvasir, SUN_SEG, and Endo-FM exemplify the shift toward comprehensive benchmarks.
3.14 Other Modalities
Other 2D modalities add clinically important but unevenly scaled datasets to a landscape dominated by a few large categories. Across this domain, fragmentation, imbalance, annotation inconsistency, and limited 2D context constrain foundation-model development, although scale and multimodal learning offer opportunities.
- Datasets by Tasks: Classification comprises approximately 798,000 images across 15 datasets, whereas segmentation comprises approximately 2,300 images across four datasets.CDD-CESM is identified as a multi-task dataset supporting segmentation and classification.
- Landscape and Challenges: 2D datasets offer substantial pretraining scale but remain fragmented, heterogeneous, and limited by the partial spatial information of two-dimensional representations.These conditions pose challenges for developing robust and generalizable medical foundation models.
- Challenges: 2D data are scattered across repositories with inconsistent protocols, resolutions, and metadata, creating domain shifts between datasets of the same modality.Examples include staining variation in histopathology and projection or exposure differences in chest X-rays.
- Challenges: Pathology, X-ray, and fundus photography dominate the collection, while endoscopy and ultrasound are underrepresented; over 80% of images come from thoracic and breast datasets.Skewed pretraining data may limit generalization to less common modalities or pathologies.
- Opportunities: Millions of 2D images and diverse modalities support self-supervised, multimodal, and clinically scalable applications, while strategic consolidation and balanced coverage are proposed opportunities.Examples include masked auto-encoding, contrastive learning, imaging-text integration, and screening applications using common 2D modalities.
4 3D Medical Image Datasets
The survey identifies a substantial but uneven 3D medical imaging landscape, with 591 datasets containing more than 1.24 million volumes. CT and MRI dominate, while anatomy and task distributions are long-tailed and acquisition, annotation, and integration remain costly.
- Overview: 591 3D datasets contain more than 1,242,022 volumes, providing richer spatial information than 2D data for volumetric analysis and clinical decision-making.Labeled datasets dominate the collection, while unlabeled datasets offer additional self-supervised learning opportunities.
- Overview: MRI and CT are the most prevalent 3D modalities, while the brain, abdomen, and lung contain the largest numbers of datasets.The distributions across modalities, anatomies, and tasks show clear long-tail patterns.
- CT Volumes: CT includes 252 datasets and approximately 516,087 volumes, ranging from small collections such as 3D-IRCADb to large compilations such as CT-RATE.Annotation quality varies from weak or semi-automated supervision to expert-verified labels, including TotalSegmentator’s labels across 104 anatomical structures.
- CT Tasks: CT segmentation comprises 150 datasets and 266,862 volumes, while classification comprises 93 datasets and 206,483 volumes.Detection includes 20 datasets and 52,542 volumes, covering nodules, pulmonary embolism, and lesions.
- Anatomical Coverage: Brain/neuro MRI contains 155 datasets and 356,751 volumes, while whole-body CT includes seven datasets and 123,557 volumes.Whole-body collections support multi-organ segmentation and cross-anatomical learning, whereas brain datasets span tumors, Alzheimer’s disease, stroke, and multiple sclerosis.
- Challenges: High acquisition and annotation costs constrain 3D dataset growth and diversity relative to 2D imaging.Specialized scanners and expert volumetric annotation contribute to the modest growth of 3D datasets.
5 Medical Video Datasets
Medical video datasets provide spatiotemporal information for dynamic clinical tasks, but their distribution is highly uneven. Endoscopy dominates the collection, while anatomy and modality gaps, privacy constraints, domain shift, and computational demands limit broad foundation-model development.
- Tasks: Video datasets support classification, segmentation, detection, tracking, estimation, registration, and other spatiotemporal clinical applications.These tasks address surgical workflow analysis, guidance, disease screening, motion quantification, and cross-modal alignment.
- Overview: Endoscopy accounts for 85.9% of medical videos, while stomach, colon, and esophagus each represent about 30% of the anatomical distribution.Retina, heart, pupil, and iris each constitute less than 2% of the collection, and task frequencies remain long-tailed.
- Dataset Distribution: The survey identifies 77 medical video datasets, including 56 endoscopy, 10 microscopy, and seven ultrasound datasets.Endoscopy’s prevalence is associated with routine recording in surgical suites, high-resolution acquisition, and relatively straightforward annotation of visible structures.
- Challenges: Medical video releases are often restricted in scale or geographic scope because anatomical context can itself identify patients.The paper identifies technical anonymization and standardized regulation as requirements for addressing this privacy barrier.
- Challenges: Device, protocol, and surgical-practice differences can cause models trained on one dataset to fail on another.The paper highlights cross-institutional benchmarks, domain adaptation, and self-supervised pretraining as responses to persistent domain shift.
- Challenges: Hours-long, high-resolution recordings make storage, annotation, and real-time analysis resource-intensive, requiring accuracy–efficiency trade-offs for deployment.The burden increases when benchmarks combine detection, segmentation, and tracking.
6 Paradigm for Dataset Fusion
The Metadata-Driven Fusion Paradigm addresses fragmented medical imaging data through metadata-centered discovery, harmonization, alignment, and composition. Its case study produces a structured integrated pool, while an interactive portal supports search, auditing, statistical analysis, and dataset refinement.
- MDFP: MDFP systematizes dataset discovery, auditing, and composition primarily through structured metadata rather than raw pixels.This design targets lower handling overhead and privacy risk while strengthening reproducibility and auditability.
- Dataset Representation: The survey organizes datasets in a multidimensional database by dimension, modality, anatomy, case count, label availability, and task.Dataset-level and annotation-level JSONL files preserve metadata, provenance, media geometry, and task-specific annotations.
- MDFP Workflow: MDFP uses four sequential phases, including metadata-schema harmonization and semantic alignment grounded in UMLS and MeSH.The schema standardizes modality, dimensionality, hierarchical anatomy, provenance, and task vocabulary across heterogeneous datasets.
- Harmonization and Alignment: The harmonized metadata table enables cross-dataset comparison, reproducible filtering, and integration, while the aligned task vocabulary supports goal-oriented filtering and benchmarking.These outputs are intended to improve consistency, completeness, interoperability, and clinical interpretation.
- Fusion Composition: The MDFP case study integrates 57 datasets and 2,135,301 validated 2D images across CT, MR, and Fundus.The composition contains 10 CT datasets with 1,173,965 images, five MR datasets with 681,025 images, and 42 Fundus datasets with 280,311 images.
- Fusion Composition: All CT and MR datasets are fully annotated, while Fundus datasets have a labeled_ratio of 0.952 within the 2D CT/MR/Fundus case study.The integrated pool covers classification, segmentation, detection, and regression under the stated dimensional and modality constraints.
7 Discussion
Current medical imaging datasets are insufficiently aligned with foundation-model needs: they are narrow in task scope, scarce across modalities, and constrained by privacy, licensing, temporal, and clinical-context requirements.
- Task Definition: Most datasets target proxy tasks such as segmentation, classification, or detection rather than direct clinical goals, making them difficult to reuse for diagnosis or treatment recommendation.Re-annotation requires expert knowledge and scales poorly; a lung-nodule mask does not encode malignancy or biopsy need.
- Multimodal Data: Multimodal datasets combining imaging with reports, genomics, and temporal records remain exceedingly rare and lack standardized collection, alignment, and validation frameworks.Modalities operate at different spatial and temporal scales, complicating semantic consistency and cross-modal evaluation.
- Scale and Diversity: Foundation-model training also requires broader coverage of diseases, protocols, specialties, demographics, rare conditions, and atypical presentations than current narrow datasets provide.The gap is especially acute in pediatric imaging, rare diseases, underrepresented populations, and longitudinal treatment monitoring.
- Governance: Privacy regulations and institutional intellectual-property policies fragment medical data access even when synthetic augmentation can improve training resources.These constraints differ from general-domain data sharing and limit the broader reuse of enhanced datasets.
- Clinical Context: Medical AI must incorporate workflows, resource constraints, patient context, prior treatments, and disease progression, but current training paradigms inadequately support this temporal intelligence.The paper calls for governance, federated learning, standardized licensing, and clinically grounded evaluation benchmarks.
8 Conclusion
The survey finds that open-access medical imaging data remain fragmented, imbalanced, small-scale, and concentrated in conventional tasks and modalities. It proposes broader data release, synthetic data, annotation-efficient learning, and public foundation models as routes toward more generalizable medical AI.
- Conclusion: The survey of over 1,000 datasets finds a fragmented and imbalanced landscape dominated by small-scale, task-specific, modality-restricted resources.Disparities are pronounced across anatomical regions and imaging modalities.
- Conclusion: Segmentation and classification dominate released tasks, while visual question answering and multimodal reasoning remain underrepresented.The imbalance highlights the field’s incomplete transition from task-oriented to foundation-oriented data engineering.
- Future Directions: Broader public release, synthetic data generation, annotation-efficient learning, and public release of models trained on private data are identified as complementary priorities.These strategies address access, privacy, scarcity, partial labeling, and institutional data-sharing constraints within the paper’s proposed roadmap.
B Tables of 3D Medical Image Datasets
The tables organize medical imaging datasets across modalities, tasks, anatomies, and dimensionalities, with the supplied entries spanning 2D and 3D resources and diverse clinical applications.
- Modality Coverage: The tables include 2D CT, MRI, PET, ultrasound, X-ray, OCT, fundus, dermoscopy, histopathology, microscopy, infrared, endoscopy, and other-modality datasets.The modality listings also include multimodal resources combining several imaging types.
- Multimodal Resources: Several listed resources combine modalities, including CT, MR, PET, ultrasound, X-ray, pathology, and whole-slide imaging.The multimodality annotations identify the imaging components associated with named resources such as AREN0534, ImageCLEF, and CMB-CRC.
- Task Labels: Dataset records characterize tasks using labels such as segmentation, detection, classification, reconstruction, registration, localization, estimation, prediction, generation, and tracking.The supplied abbreviation entries define the task labels used throughout the tables.
- 3D Dataset Coverage: The 3D entries cover CT and CT/MR or CT/PET resources across brain, head and neck, chest, lung, abdomen, spine, pelvis, breast, and other anatomies.Examples include segmentation, registration, and classification datasets for cancer, stroke, respiratory motion, and other conditions.
C Tables of Medical Video Datasets
Table 26 compiles medical video datasets spanning endoscopy and microscopy, with entries covering segmentation, detection, classification, tracking, retrieval, and estimation across varied anatomies and surgical targets.
- Video datasets cover endoscopy, microscopy, ultrasound, RGB, and combined video–3D or video–2D formats.
- Segmentation, detection, and classification recur across the catalog, often combined with tracking, surgical-phase recognition, or instrument analysis.
- The listed anatomies and targets include gallbladder, colon, retina, abdomen, liver, kidney, prostate, uterus, placenta, and surgical instruments.
- Dataset sizes vary substantially, from 4 kidney images in KBD and 8 retina images in CatRelDet to 4,019 polyp images in EndoCV 2021 and 18,000 surgical-skill examples.
- Several datasets target surgical understanding through labels for phases, actions, instruments, anatomy, skill, abnormalities, or nursing procedures.