Source-linked AI summary

A Review on Deep-Learning Algorithms for Fetal Ultrasound-Image Analysis

Maria Chiara Fiorentino, Francesca Pia Villani, Mariachiara Di Cosmo, Emanuele Frontoni, Sara Moccia

arXiv:2201.12260v1eess.IVcs.CVcs.LG

TL;DR

The review provides a compact overview of deep-learning research across fetal ultrasound standard-plane detection, anatomical-structure analysis, and biometry estimation. It surveys these applications, datasets, metrics, methods, and open issues, finding that anatomical-structure analysis is the largest task category while comprehensive end-to-end fetal assessment remains underdeveloped.

  • Problem

    A compact overview is needed to organize fetal ultrasound deep-learning applications, datasets, metrics, and open issues for researchers in the field.

  • Method

    The review surveys recent fetal ultrasound deep-learning papers across standard-plane detection, anatomical-structure analysis, and biometry estimation, while summarizing datasets, metrics, methodologies, and applications.

  • Results

    Anatomical-structure analysis accounts for 46.6% of surveyed papers, while standard-plane detection and biometry parameter estimation account for 19.9% and 21.9%, respectively.

  • Takeaways & Limitations

    The review identifies current challenges for translating fetal ultrasound deep-learning research into clinical practice, including the incomplete development of unified end-to-end analysis.

  • Takeaways & Limitations

    The review highlights limited exploitation of anatomical shape priors and a lack of annotated datasets containing both healthy and pathological cases.

Abstract

from arXiv · show

Deep-learning (DL) algorithms are becoming the standard for processing ultrasound (US) fetal images. Despite a large number of survey papers already present in this field, most of them are focusing on a broader area of medical-image analysis or not covering all fetal US DL applications. This paper surveys the most recent work in the field, with a total of 145 research papers published after 2017. Each paper is analyzed and commented on from both the methodology and application perspective. We categorized the papers in (i) fetal standard-plane detection, (ii) anatomical-structure analysis, and (iii) biometry parameter estimation. For each category, main limitations and open issues are presented. Summary tables are included to facilitate the comparison among the different approaches. Publicly-available datasets and performance metrics commonly used to assess algorithm performance are summarized, too. This paper ends with a critical summary of the current state of the art on DL algorithms for fetal US image analysis and a discussion on current challenges that have to be tackled by researchers working in the field to translate the research methodology into the actual clinical practice.

I. INTRODUCTION

Fetal ultrasound is clinically valuable but difficult to analyze because artifacts and tissue interactions degrade image quality. This review surveys deep-learning research across fetal standard-plane detection, anatomical-structure analysis, biometry estimation, datasets, and evaluation metrics.

  • Ultrasound is widely used during pregnancy to evaluate fetal growth, development, pregnancy progression, and clinical suspicion.
  • Artifacts including acoustic shadows, speckle noise, motion blurring, and missing boundaries make fetal ultrasound analysis challenging for clinicians.
  • Deep learning, particularly convolutional neural networks, has increasingly been used to provide decision support in fetal ultrasound image analysis.
  • Existing surveys often address broader medical-image analysis, nonfetal ultrasound, or narrower fetal applications rather than the full fetal ultrasound deep-learning field.
  • The review organizes recent work around standard-plane detection, anatomical-structure analysis, and biometry estimation, while also summarizing datasets, metrics, limitations, and open issues.
  • Classification, segmentation, localization, detection, regression, and saliency tasks use different evaluation measures, including accuracy, recall, specificity, AUC, DSC, IoU, AP, MSE, and MAE.

B. Publicly-available datasets

Public fetal-ultrasound datasets support deep-learning development but remain difficult to collect, annotate, and share. The available challenges vary substantially in task, scale, and annotation scope.

  • High-quality annotated fetal-ultrasound datasets are difficult to share because labeling is time-consuming and privacy concerns limit data distribution.
  • The 2012 Challenge US dataset contains 270 images for fetal abdomen, head, and femur segmentation, a size considered insufficient for generalizable deep-learning algorithms.
  • The HC18 challenge provides 999 training and 335 testing two-dimensional ultrasound images for fetal-head-circumference estimation.
  • The A-AFMA challenge targets amniotic-fluid and maternal-bladder detection plus landmark identification for maximum vertical pocket measurement.
  • A publicly available screening dataset contains 7129 training and 5271 testing two-dimensional images from 896 patients across six standard-plane-related classes.

III. FETAL STANDARD-PLANE DETECTION

Fetal standard-plane detection methods use deep learning to identify clinically relevant ultrasound planes, spanning classification, anatomical-structure detection, and 2D or 3D plane localization. The reviewed approaches cover multiple fetal planes and include both image- and video-based analysis.

  • Clinical role: Standardized fetal planes support reproducible biometry and fetal evaluation, especially abdomen, brain, and femur planes.Brain standard-plane analysis includes trans-ventricular and trans-thalamic views.
  • Detection strategies: Early approaches used cascaded CNNs to localize fetal anatomy and then classify plane-specific structures such as the stomach bubble and umbilical vein.One fetal-abdomen system processed 2606 images from 219 ultrasound videos.
  • Detection strategies: Classification CNNs detected individual planes, including facial and brain standard planes, with reported mean AUC values of 0.99 in representative studies.One facial-plane study also reported Acc=0.96, Prec=0.96, Rec=0.97 and F1=0.97 on 2418 images.
  • Detection strategies: Video-based methods combined CNNs with recurrent models to use temporal information, achieving mean Acc, Prec, Rec and F1 values of 0.87, 0.71, 0.64 and 0.64 in one study.Another video study reported 0.85 for each of these four metrics.
  • Detection strategies: The literature also includes anatomical-structure-based detection, motivated by the limitation that classification alone does not explicitly identify anatomical landmarks.This approach is intended to better reflect how clinicians recognize standard planes.
  • 3D localization: Alternative methods predicted plane geometry from 3D ultrasound, reporting plane-centre differences of 3.44 mm and rotation angles of 11.05° in fetal-brain volumes.Other 3D approaches reported DF values of 3.03 mm, 2.31 mm and 1.82 mm, with corresponding angles of 9.36°, 10.36° and 7.20°.

A. Limitations and open issues for fetal standard-plane detection

The review identifies limited plane coverage, inconsistent patient-level experimental splits, and heterogeneous datasets as major barriers to evaluating fetal standard-plane detectors fairly. It also calls for broader reporting to assess model bias and variance.

  • Scope: Many studies consider only a few standard planes, and some investigate a single plane, limiting the scope of standard-plane detection evidence.The cited examples include studies focused on one plane or a small subset of planes.
  • Experimental design: Few studies report patient-level splits, although images from the same woman should not occur in both training or validation and test sets because this can introduce bias.The review emphasizes patient-based reporting as crucial for experimental validity.
  • Evaluation: Different datasets, including very small ones, prevent fair quantitative comparison across proposed approaches.Dataset heterogeneity makes reported performance values difficult to compare directly.
  • Evaluation: The first publicly released dataset improved evaluation, but it covers only FASP, FBSP, 4CH, maternal cervix and FFESP.The review additionally recommends training and validation curves and visual explanation techniques such as Grad-CAM to assess bias and variance.

IV. ANATOMICAL-STRUCTURE ANALYSIS

The anatomical-structure analysis literature is organized around fetal heart, brain, placenta, amniotic fluid, other individual structures, and multi-structure analysis. Figure 5 summarizes the principal tasks addressed across these areas.

  • Organization of the literature: The review surveys fetal heart, brain, placenta and amniotic-fluid analysis, while grouping other structures and multi-structure methods separately.The heart section covers cardiac evaluation, the brain section covers fetal-brain analysis, and the remaining sections address additional or combined structures.

A. Heart

Fetal-heart deep-learning studies address structural detection, segmentation, cardiac measurements, and congenital-heart-disease classification. The reviewed methods range from CNN-based detectors to encoder-decoder, recurrent, ensemble, and generative models.

  • Clinical motivation: Fetal cardiac evaluation supports detection of congenital heart diseases and intrauterine growth restriction through cardiac-function and anatomical assessment.Assessment includes heart dimensions and shape.
  • Detection: Detection models identify cardiac structures and landmarks in four-chamber views, with one SSD-based approach achieving mAP=0.93 on 1991 images.The reported structures included the left atrial pulmonary vein angle, apex cordis, moderator band and multiple ribs.
  • Segmentation: Encoder-decoder CNNs are used for segmentation because heart and cardiac-structure shape can indicate possible pathology.A U-Net study segmented cardiac standard planes with IoU=0.94 and Acc=0.96 on 106 images containing atrial and ventricular septal defects.
  • Segmentation: Video segmentation methods dynamically fine-tune CNNs across sequential echocardiographic frames to track the fetal left ventricle.The method uses multiscale information and shallow tuning to adapt to the latest frame.
  • Segmentation: Instance-segmentation networks with category, mask and category-attention branches segment the four cardiac chambers while correcting instance misclassification.Five-fold cross-validation on 638 images from 319 fetuses produced chamber-specific mean DSC values from 0.75 to 0.82.
  • Disease diagnosis: Congenital-heart-disease studies combine standard-plane detection with classifiers, cardiac-parameter estimation, feature visualization, augmentation, and one-class classification.One multi-dataset approach classified normal and abnormal hearts with mean AUC=0.94.

B. Brain

Deep-learning fetal-brain analysis spans segmentation, localization, measurement, gestational-age estimation, and pathology support. The reviewed approaches use 2D and 3D architectures, including encoder-decoder, attention, Bayesian, and multi-task models.

  • Brain analysis: Fetal-brain analysis addresses anatomical localization, segmentation, classification, measurement, gestational-age estimation, and brain-development assessment.These tasks support evaluation of fetal growth and detection of brain pathologies despite high intra- and inter-structure variability.
  • Segmentation: Encoder-decoder architectures are used to segment brain structures and support downstream measurements or clinical actions.Examples include middle cerebral artery segmentation with Doppler gate positioning and CSP segmentation with width measurement.
  • 3D analysis: 3D U-Net approaches process volumetric information to segment and localize multiple fetal-brain structures across standard planes.One hybrid-attention model achieved DSC = 0.96, IoU = 0.92 and HD = 0.46 mm on 50 test volumes.
  • Multi-task analysis: A multi-task FCN jointly performs 3D brain localization, structural segmentation, and alignment using skull, eye-socket, and head-pose references.Testing used 2D axial slices sampled from 140 volumes and reported Acc = 0.99.
  • Clinical support: DL pipelines also support gestational-age estimation and brain-lesion diagnosis through regression, segmentation, classification, and Grad-CAM localization.The lesion-diagnosis pipeline reported classification Acc, Rec, Spec and AUC of 0.96, 0.97, 0.96 and 0.99, while heatmaps localized lesions correctly in 61.60% of images.

C. Placenta & amniotic fluid

Placenta, amniotic fluid, and lungs are analyzed because their condition or measurements relate directly to fetal health, development, and clinical assessment. Reviewed DL methods use 2D and 3D segmentation, classification, and gestational-age estimation.

  • Placenta: Placental function supports fetal oxygenation, nutrition, thermoregulation, and waste removal, making placenta assessment relevant to fetal health.Abnormal placental function may affect fetal development and, in severe cases, endanger fetal life.
  • Placenta: 2D placenta methods include U-Net-inspired segmentation with acoustic-shadow detection and separate normal-versus-abnormal classification.One model tested on 205 images obtained mean DSC = 0.92 and Acc = 0.93.
  • Placenta: 3D placenta approaches use 3D U-Nets and multi-view acquisition, voxel-wise fusion, segmentation, and volume estimation.Reported results include DSC = 0.84 and HD = 14.6 mm on 1196 volumes, and mean DSC = 0.80 on 12 volumes.
  • Amniotic fluid: Amniotic-fluid analysis combines segmentation with amniotic-fluid-index assessment, reflecting the fluid’s roles in protection, infection prevention, and organ development.A VGG16-based model tested on 400 images obtained Acc = 0.78 and IoU = 0.54 for amniotic-fluid segmentation.
  • Lungs: Fetal-lung DL methods include gestational-age classification from four-chamber-plane images, motivated by the link between lung immaturity and neonatal respiratory morbidity.A DenseNet evaluated on images from 1023 pregnancies achieved overall Acc = 0.83 across three gestational-age classes.

2) Kidney:

The reviewed anatomical-structure applications extend beyond brain, placenta, and fluid to kidneys, spine, abdomen, femur, head, and multi-organ screening. Methods combine 3D segmentation, landmark localization, multimodal inputs, recurrent models, and multi-task learning.

  • Kidney: Kidney assessment is clinically relevant because poor fetal kidney development is associated with increased risk of kidney disease in adulthood.Abnormal fetal growth is also associated with reduced kidney functionality after birth.
  • Spine: Spine analysis targets centerline identification in 3D ultrasound to assess growth-related malformations.The cited approach uses a multi-scale CNN for spine segmentation.
  • Spine: Rec, Prec, Acc, and IoU were 0.93, 0.96, 0.94, and 0.91, respectively, with mean standard error of 4.12 mm and average running time of 12.15 seconds.The study included 24 spina-bifida cases among 3300 total cases and did not divide results among spina-bifida types.
  • Other structures: Fetal abdomen, femur, and head methods address circumference, landmark, volume, and skull-segmentation tasks using CNNs, deformable models, and cascaded 3D U-Nets.The head model adds ultrasound-wave incidence-angle and shadow-casting maps to a second 3D U-Net.
  • Multi-organ analysis: Multi-organ analysis seeks to reproduce comprehensive clinical screening across anatomical structures through classification, localization, segmentation, and landmark detection.The surveyed work includes 14-structure classification, text-image integration, cross-device adaptation, and multitask abdominal-organ analysis.
  • Multi-organ analysis: A 3D CNN–RNN pipeline combines spatial intensity concurrency with spatial-sequential encoding to refine boundaries, achieving average DSC = 0.80 and HD = 14.13 mm on 44 volumes.The recurrent component encodes spatial sequences for boundary refinement.

F. Limitations and open issues for anatomical-structure analysis

Anatomical-structure analysis remains concentrated in segmentation, heart and brain studies, while advanced architectures, 3D/4D imaging, multi-organ modeling, and public datasets remain underused or insufficient for fair comparison.

  • F. Limitations and open issues for anatomical-structure analysis: Segmentation is the most addressed task, but most methods use 2D U-Net architectures while adversarial and spatio-temporal approaches remain less explored.Shape priors have also not been fully exploited, although recent adversarial strategies impose shape constraints on segmentation outputs.
  • F. Limitations and open issues for anatomical-structure analysis: 3D, 4D, and spatiotemporal image correlation techniques have not yet been exploited for more detailed fetal-heart anatomical and functional assessment.This leaves newer volume-sonography techniques underused in the surveyed heart-analysis literature.
  • F. Limitations and open issues for anatomical-structure analysis: 30.3% of surveyed anatomical-structure papers concern the heart and 24.2% concern the brain, whereas few address multiple organs.The literature on fetal brain analysis is heterogeneous and fragmented across localization, segmentation, classification, and gestational-age estimation.
  • F. Limitations and open issues for anatomical-structure analysis: The lack of public datasets makes comparisons among heterogeneous fetal-brain approaches challenging.Limited annotation also restricts characterization of inter-organ relations and differentiation between pathological and physiological images.
  • F. Limitations and open issues for anatomical-structure analysis: The surveyed anatomical analyses rely on varied datasets and metrics, including DSC, IoU, MAE, HD, and average perpendicular distance.Reported evaluations span 2D and 3D data, with examples including IoU = 0.92 for abdomen and head segmentation and DSC = 0.98 with HD = 1.14 mm.

A. Limitations and open issues for biometry parameter estimation

Biometry estimation is constrained by dataset limitations, incomplete end-to-end automation, and the absence of a unified multi-region framework. Additional work also addresses adipose-tissue analysis, but these applications remain comparatively limited.

  • A. Limitations and open issues for biometry parameter estimation: The HC18 dataset advanced head-circumference research but its limited image count and single anatomical focus restrict advanced methodological study.It contains only fetal-brain images, limiting evaluation across broader anatomical structures.
  • A. Limitations and open issues for biometry parameter estimation: Only one surveyed work proposes an approach that first identifies the scan plane and then estimates the associated biometrics, so further research is needed.Such end-to-end automation is described as a potential support tool for clinicians.
  • A. Limitations and open issues for biometry parameter estimation: A unified multi-region biometry framework has not yet been proposed because fetal structures vary greatly in shape, morphology, contrast, and size across gestational trimesters.The passage specifically contrasts abdomen, femur, and cerebellum in their appearance and suitability for ellipse fitting.
  • A. Limitations and open issues for biometry parameter estimation: Fetal adipose-tissue analysis includes U-Net segmentation and Faster R-CNN localization followed by threshold-based thickness measurement.Reported test performance includes DSC = 0.80 for adipose-tissue segmentation and DSC = 0.93 for fetal-thigh cross-sectional analysis.

2) Face analysis:

The review covers fetal ultrasound applications spanning face analysis, pose estimation, cervical assessment, simulation, workflow analysis, and volume reconstruction. These approaches address anatomical visualization, operator support, image-quality improvement, and broader clinical analysis tasks.

  • Face analysis: A 3D CNN segments fetal faces from ultrasound volumes, achieving a mean ED of 1.72 mm in five-fold cross-validation.
  • Pose estimation: Fetal pose estimation localizes 16 landmarks using self-supervised refinement, achieving ED of 4.92 mm and AUC of 62.90% over 52 ultrasound volumes.
  • Cervical analysis: U-Net-based cervix segmentation supports cervical-length and anterior-cervical-angle estimation through centerline and iterative geometric algorithms.
  • US simulation: A patch-based GAN improves simulated ultrasound quality while preserving anatomical structures and acoustic shadows at constant computational time, obtaining mean KLD of 13.80.
  • Clinical support and reconstruction: DL systems analyze clinical workflows, guide probe movement, and reconstruct 3D volumes from 2D frames to extend fetal ultrasound coverage.
  • Application distribution: Across the reviewed field, anatomical analysis is concentrated on fetal heart and brain applications, while lungs, kidneys, and spine receive comparatively marginal attention.

1) Multi-expert image annotation:

The review identifies annotation quality, dataset limitations, and computational burden as major constraints on robust and clinically useful fetal-ultrasound DL. It also highlights the need for comprehensive models spanning fetal development and analysis tasks.

  • Multi-expert image annotation: Multi-clinician annotation is crucial for robust algorithms, fair comparisons, and measuring inter-clinician variability, but is costly in time and money.Manual annotation as the primary reference therefore requires careful interpretation.
  • Comprehensive analysis: Fetal modeling is difficult because growth causes substantial structural and physiological changes across trimesters and variability within and between organs.An end-to-end system for plane classification followed by biometry or defect evaluation remains insufficiently exploited; one unified approach was limited to the heart.
  • Semi, weak and self-supervised learning: Only 13 of 145 surveyed papers investigate semi-, weakly, or self-supervised learning, leaving these approaches relatively unexplored for small annotated datasets.Applications include plane detection, fetal pose estimation, probe movement estimation, reconstruction, multi-organ analysis, and biometry estimation.
  • Model efficiency: Most surveyed papers do not report training or deployment cost, limiting assessment of on-device efficiency and associated CO2 consumption.The review proposes model efficiency as an additional evaluation metric.

7) Use of federated learning:

Federated learning is presented as a privacy-preserving paradigm for multi-institutional fetal-US collaborations, addressing barriers created by centrally sharing clinical images. The review also notes that ethical principles were not addressed in the surveyed papers.

  • Use of federated learning: Most surveyed studies use single-center or internationally shared datasets, while sufficiently large and diverse fetal-US datasets remain difficult to access.Centrally shared images raise privacy and ownership concerns.
  • Use of federated learning: Federated learning enables data-private collaboration among clinical institutions without relying on centrally shared ultrasound images.The passage introduces it as a novel paradigm for multi-institutional collaborations.
  • Ethics: The review reports that surveyed papers made no attempts to apply ethical principles and guidelines across DL design and deployment.The discussion refers to the European Commission’s 2018 Ethics Guidelines for Trustworthy AI.
  • Review contribution: The review summarizes methods, performance metrics, training and test sizes, and annotator counts while highlighting each method’s advantages and disadvantages.It aims to provide an overview useful to both young researchers and established researchers.

NOMENCLATURE

The nomenclature defines abbreviations for fetal ultrasound planes, biometric measurements, model architectures, evaluation metrics, clinical structures, and related organizations or methods.

  • Biometry: AC, FL, MVP, OFD, TCD, and LVR denote fetal biometric measurements or related clinical measurements.The terms include abdominal circumference, femur diaphysis length, maximum vertical pocket, occipito-frontal diameter, trans-cerebellar diameter, and lateral ventricle ratio.
  • Metrics: AUC, AUC-J, DSC, IoU, KLD, NSS, and RMSE denote performance metrics used in fetal ultrasound analysis.They cover ROC-based, segmentation, saliency, divergence, and regression measures.
  • Models and learning methods: CNN, FCN, GAN, GRU-RCN, LSTM, RNN, NAS, SVM, SE, SSL, and DRL denote deep-learning architectures, learning paradigms, or model-design methods.The nomenclature includes convolutional, recurrent, adversarial, self-supervised, reinforcement-learning, and architecture-search terminology.
  • Clinical terms: CSP, CTR, FV, LVOT, RVOT, and 3VT denote anatomical structures, fetal conditions, or cardiac ultrasound views.The list includes cavum septum pellucidum, cardio-thoracic ratio, fetal ventriculomegaly, outflow tracts, and three-vessel trachea.
  • Standard planes: FASP, FBSP, FCSP, FFASP, FFSP, FFESP, FLVSP, FTSP, and FVSP denote fetal ultrasound standard planes.These abbreviations cover abdomen, brain, face, femur, lumbosacral spine, trans-thalamic, and trans-ventricular planes.
  • Organizations and measures: ISBI, ISUOG, and MICCAI denote biomedical-imaging or obstetric-ultrasound organizations and conferences.The nomenclature also includes DIFF, the mean plane centres difference.
Loading 2201.12260v1…