Source-linked AI summary
TotalSegmentator: robust segmentation of 104 anatomical structures in CT images
Jakob Wasserthal, Hanns-Christian Breit, Manfred T. Meyer, Maurice Pradella, Daniel Hinck, Alexander W. Sauter, Tobias Heye, Daniel Boll, Joshy Cyriac, Shan Yang, Michael Bach, Martin Segeroth
TL;DR
Segmentation of anatomical structures supports radiologic biomarkers, pathology detection, tumor quantification, and clinical planning, but robust whole-body tools are needed. The authors developed and evaluated a model for 104 structures across diverse CT data, achieving high accuracy and applying it to age-related volume and attenuation analyses.
Problem
Anatomical segmentation is important for extracting radiologic biomarkers, detecting pathologies, quantifying tumor load, and supporting surgical or radiotherapy planning.
Method
The authors created an nnU-Net-based tool for segmenting 104 anatomical structures using manually refined annotations and trained models at 1.5 mm and 3 mm isotropic resolution.
Results
0.943 Dice was achieved on diverse clinical CT data, while age-related volume and attenuation changes were evaluated across more than 4000 examinations.
Takeaways & Limitations
The tool provides robust segmentation for most structures and supports large radiologic population studies examining organ volume and attenuation.
Takeaways & Limitations
Male patients were overrepresented in the study datasets, potentially reflecting the hospital population.
Abstract
from arXiv · showhide
We present a deep learning segmentation model that can automatically and robustly segment all major anatomical structures in body CT images. In this retrospective study, 1204 CT examinations (from the years 2012, 2016, and 2020) were used to segment 104 anatomical structures (27 organs, 59 bones, 10 muscles, 8 vessels) relevant for use cases such as organ volumetry, disease characterization, and surgical or radiotherapy planning. The CT images were randomly sampled from routine clinical studies and thus represent a real-world dataset (different ages, pathologies, scanners, body parts, sequences, and sites). The authors trained an nnU-Net segmentation algorithm on this dataset and calculated Dice similarity coefficients (Dice) to evaluate the model's performance. The trained algorithm was applied to a second dataset of 4004 whole-body CT examinations to investigate age dependent volume and attenuation changes. The proposed model showed a high Dice score (0.943) on the test set, which included a wide range of clinical data with major pathologies. The model significantly outperformed another publicly available segmentation model on a separate dataset (Dice score, 0.932 versus 0.871, respectively). The aging study demonstrated significant correlations between age and volume and mean attenuation for a variety of organ groups (e.g., age and aortic volume; age and mean attenuation of the autochthonous dorsal musculature). The developed model enables robust and accurate segmentation of 104 anatomical structures. The annotated dataset (https://doi.org/10.5281/zenodo.6802613) and toolkit (https://www.github.com/wasserth/TotalSegmentator) are publicly available.
1. Introduction
CT segmentation supports biomarker extraction, pathology detection, tumor quantification, and clinical planning, but developing such models requires tedious annotation and technical expertise. The study therefore aimed to provide a publicly available, easy-to-use model covering anatomically relevant structures across clinical settings.
- CT segmentation supports biomarker extraction, pathology detection, tumor quantification, and surgical or radiotherapy planning.
- Building and training segmentation algorithms requires tedious manual annotation and technical expertise.
- The proposed model was designed to be publicly available, easy to use, broadly anatomical, and robust across clinical settings.
- The tool was also applied to 4004 whole-body CT examinations to analyze age-dependent changes in structure volume and attenuation.
2. Materials and Methods
The study assembled and annotated a diverse CT dataset, trained nnU-Net models at two resolutions, and evaluated them against human-approved ground truth and another publicly available model. The workflow included iterative annotation, independent final-model training, and age-association analyses.
- Materials and Methods: Two datasets were aggregated: a training dataset for model development and an aging-study dataset for application analysis.
- Materials and Methods: 1204 CT series were retained after exclusions and divided into 1082 training, 57 validation, and 65 test patients.
- Data Annotation: 104 anatomical structures were manually segmented or refined under physician supervision using an iterative learning workflow.
- Data Annotation: The final model was trained independently of intermediate annotation models to reduce test-set bias.
- Model: nnU-Net models were trained at 1.5 mm and 3 mm isotropic resolution to balance accuracy with lower memory requirements.
- Evaluation: Dice and NSD were calculated against human-approved ground truth, and performance was compared with an nnU-Net trained on the BTCV dataset.
- Evaluation: Age associations with structure volume and mean attenuation were calculated after excluding structures with failed segmentations.
3. Results
The diverse training dataset supported robust segmentation across clinical and pathologic cases, while performance varied with image resolution. The aging analysis identified multiple associations between patient age and anatomical volume or attenuation.
- Study sample: The training data varied in slice thickness, resolution, contrast phase, reconstruction kernel, site, and scanner.The dataset included images from 8 sites and 16 scanners.
- Segmentation performance: 0.943 Dice was achieved by the 1.5 mm model, with NSD of 0.966 on the test set.The 3 mm model achieved Dice 0.840 but retained NSD of 0.966, indicating less precise borders at lower resolution.
- Segmentation performance: 0.932 Dice versus 0.871 was achieved by the proposed model versus the BTCV model on the test set.The difference was statistically significant for both Dice and NSD, with p<0.001.
- Pathologic cases: Robust results were obtained for distorted, displaced, missing, and duplicated structures in pathologic cases.Examples included broken bones, hernia-displaced bowels, splenectomy, and transplant kidney.
- Age-related differences: Age was negatively correlated with attenuation in several bones and muscles, including hips and autochthonous dorsal musculature.Aortic volume showed a positive correlation with age, with rs = 0.64; p<0.0001.
4. Discussion
The model addresses limited prior coverage of anatomical structures and nonrepresentative training data by providing broad segmentation and supporting large-scale radiologic analyses. The authors position it as a basis for population studies while noting demographic imbalance and planned methodological extensions.
- Contribution: The tool segments 104 anatomical structures from 1204 CT datasets spanning scanners, acquisition settings, and contrast phases.It achieved Dice score 0.943 and outperformed other freely available segmentation tools.
- Prior work: Previous models covered smaller subsets of structures and were trained on datasets that were not representative of routine clinical data.The authors connect this limitation to the need for broader segmentation resources.
- Implications: The model supports large radiologic population studies, including analysis of organ volumes and disease-related segmentation applications.The authors specifically describe creating new reference values for organ volumes as an example.
- Limitations: Male patients were overrepresented in the study datasets, possibly because males comprise more of the overall hospital population.The authors plan future aging analyses with more patients and confounder correction.
Supplemental Materials
The supplemental materials enumerate the anatomical structures included in the segmentation model, spanning organs, vessels, bones, muscles, and other body structures.
- Organs and vessels: The listed organs and vessels include the spleen, kidneys, liver, stomach, aorta, vena cava, portal and splenic veins, pancreas, and lungs.The list also includes bowel, duodenum, colon, gallbladder, adrenal glands, esophagus, trachea, and heart structures.
- Muscles: The listed muscles include bilateral gluteus maximus, gluteus medius, gluteus minimus, autochthonous, and iliopsoas muscles.Urinary bladder and additional body structures are also included in the supplied class list.
S2: List of all pretrained models which were used during data annotation
The supplemental materials list pretrained models used during annotation, covering abdominal organs, ribs, vertebrae, lungs, and thoracic structures.
- Pretrained models: The annotation resources include Multi-Atlas Labeling Beyond the Cranial Vault - Abdomen (nnU-Net Task 17).The listed classes include aorta, esophagus, heart, and trachea.
- Pretrained models: The abdominal pretrained model includes spleen, kidneys, gallbladder, esophagus, liver, stomach, aorta, vena cava, portal and splenic veins, pancreas, and adrenal glands.The class list covers bilateral kidneys and adrenal glands.
- Pretrained models: Additional resources include RibFrac 2020, the Large Scale Vertebrae Segmentation Challenge, Johof Lung Segmentation, and SegTHOR.RibFrac provides 24 rib classes that were split into individual rib instances through postprocessing.
S3: More details on data annotation
Heart subpart segmentations were extended across CT sequences by transferring labels from contrast-enhanced images to aligned native images and training an nnU-Net on both types.
- The resulting model predicted left and right atria, left and right ventricles, myocardium, and pulmonary artery across CT sequences.
- An in-house dataset provided ground-truth segmentations of heart subparts on arterial contrast-enhanced CT images.
- Aligned native CT images enabled transferring contrast-image segmentations to noncontrast images.
- An nnU-Net was trained using heart-subpart images with and without contrast agent.
S4: Model details
The model used targeted nnU-Net adaptations to handle left-right distinctions, large training data, and the memory demands of segmenting 104 classes.
- Mirroring was removed from data augmentation because it impaired distinction between left and right anatomical structures.
- 4000 training epochs replaced the default 1000 because of the dataset’s large size.
- 104 classes were split into five parts, with four models covering 21 classes and one covering 20 classes, to reduce memory consumption.
- The 3 mm model combined all 104 classes into one model and used 8000 training epochs because training took longer.
Challenge” which we use as baseline
The referenced baseline includes portal and splenic veins among its listed structures.
- The baseline’s listed structures include the portal vein and splenic vein.
S6: Reasons for low performance on some structures
The model performed accurately overall, but errors remained for anatomically difficult or poorly visible structures, including colon, heart subparts, ribs, vertebrae, and iliac vessels.
- The proposed model showed highly accurate and robust results for most structures, although minor errors occurred frequently.
- Neighboring ribs or vertebrae can be labelled incorrectly when only a subset is visible, making anatomical numbering difficult.
- Iliac vessel segmentation can miss parts on images without contrast agent because the vessels are hardly visible to the human eye.
- Colon segmentation can miss parts because variable shape, size, texture, and position make the colon difficult to distinguish from small bowel.
- Heart-subpart segmentation is inaccurate on native images because the chambers are difficult to see without contrast.
S7: Details of evaluation on BTCV dataset
The BTCV evaluation compared the proposed model with a separately trained nnU-Net baseline on 30 resampled subjects, while the aging analysis examined volume and attenuation across 104 structures and age quartiles.
- BTCV comparison: 30 subjects from the BTCV dataset were resampled to the training dataset’s 1.5 mm isotropic resolution before comparison.The resampled dataset was used to train a nnU-Net with 5-fold cross-validation and generate predictions for all subjects.
- BTCV comparison: 5-fold cross-validation predictions were compared with the proposed model, using an ensemble of five models for the study test set.The evaluation compared the proposed model with the BTCV model on both the study test set and the BTCV dataset.
- Aging analysis: Age-quartile analyses displayed Spearman correlations and boxplots for volume and attenuation across all 104 segmented structures.Quartiles were Q1 < 41 years, Q2 41–59 years, Q3 59–78 years, and Q4 > 78 years.
- Aging analysis: Kruskal-Wallis tests assessed age-quartile differences, with significance reported below p<0.0001 after Bonferroni correction.Post-hoc Mann-Whitney-U differences were indicated in the graphs, and segmentations below structure-specific lower volume cutoffs were excluded.