Source-linked AI summary

Vision Foundation Models for Computed Tomography

Suraj Pai, Ibrahim Hadzic, Dennis Bontempi, Keno Bressem, Benjamin H. Kann, Andriy Fedorov, Raymond H. Mak, Hugo J. W. L. Aerts

arXiv:2501.09001v2eess.IVcs.CV

TL;DR

Radiology requires diverse analytical tasks, yet dedicated vision-centric encoders have been lacking. This study introduces CT-FM, a self-supervised 3D foundation model that outperforms several state-of-the-art baselines across relevant tasks and exhibits anatomical awareness in its embeddings.

  • Problem

    Radiology encompasses diverse analytical tasks, while dedicated vision-centric encoders have been lacking.

  • Method

    The study presents CT-FM, a large-scale 3D foundation model pretrained with a self-supervised strategy tailored to medical imaging.

  • Results

    CT-FM outperforms several state-of-the-art baselines across relevant tasks and exhibits anatomical awareness in its embeddings.

  • Takeaways & Limitations

    The findings support CT-FM as a robust foundation model for diverse radiological tasks and embedding-based anatomical analysis.

  • Takeaways & Limitations

    Diagnostic and prognostic quantification abilities remain to be explored.

Abstract

from arXiv · show

Foundation models (FMs) have shown transformative potential in radiology by performing diverse, complex tasks across imaging modalities. Here, we developed CT-FM, a large-scale 3D image-based pre-trained model designed explicitly for various radiological tasks. CT-FM was pre-trained using 148,000 computed tomography (CT) scans from the Imaging Data Commons through label-agnostic contrastive learning. We evaluated CT-FM across four categories of tasks, namely, whole-body and tumor segmentation, head CT triage, medical image retrieval, and semantic understanding, showing superior performance against state-of-the-art models. Beyond quantitative success, CT-FM demonstrated the ability to cluster regions anatomically and identify similar anatomical and structural concepts across scans. Furthermore, it remained robust across test-retest settings and indicated reasonable salient regions attached to its embeddings. This study demonstrates the value of large-scale medical imaging foundation models and by open-sourcing the model weights, code, and data, aims to support more adaptable, reliable, and interpretable AI solutions in radiology.

INTRODUCTION · RESULTS

CT-FM is a vision-centric foundation model pre-trained on 148,000 CT scans with task-agnostic self-supervised learning for diverse radiological tasks. It outperformed compared baselines and state-of-the-art approaches while producing localized, interpretable embeddings and supporting retrieval, with its pipeline, dataset, and code open-sourced.

  • INTRODUCTION: Radiology requires multiple specialized models for localization, diagnosis, monitoring, segmentation, comparison, and prognosis, creating substantial operational overhead.Foundation models are introduced as a unified solution because they can learn and adapt across multiple tasks.
  • INTRODUCTION: Vision-centric encoders remain lacking because many state-of-the-art systems jointly train vision with other modalities rather than learning solely from visual data.A vision-only encoder could better represent visual feature distributions and complement existing foundation models.
  • INTRODUCTION: 2D slice-wise pre-training dominates radiological image-based pre-training, while existing native 3D methods are not focused on spatial semantics in 3D imaging data.This limitation motivated a self-supervised design focused on large-scale unannotated 3D imaging data.
  • RESULTS: 148,000 CT scans from the Imaging Data Commons were used to pre-train CT-FM for whole-body and tumor segmentation, head CT triage, image retrieval, and semantic understanding.The model was evaluated through fine-tuning on clinically relevant tasks and zero-shot evaluations.
  • RESULTS: CT-FM quantitatively outperformed compared baselines and state-of-the-art approaches across the evaluated radiological tasks.The evaluations included whole-body segmentation of 117 labels, heterogeneous tumor segmentation, head CT triage, content-based retrieval, and semantic understanding.
  • RESULTS: CT-FM learned localized representations that supported anatomical clustering, concept identification, and retrieval of full CT scans or scans containing specific structures.These embedding properties were associated with greater interpretability and semantic understanding.
  • INTRODUCTION: The CT-FM model pipeline, complete dataset, and implementation code were open-sourced in a reproducible and transparent framework.The stated aim was to accelerate development and evaluation of CT interpretation models across clinical use-cases.
  • RESULTS: The foundation model used a task-agnostic self-supervised learning strategy tailored to medical imaging data and was tested in both fine-tuned and zero-shot settings.The source cohort represented different cancer types, phenotypes, and associated comorbidities.

Segmentation of 117 anatomical structures in whole-body CT scans · Segmentation of cancer across anatomical sites on CT scans

CT-FM improved whole-body segmentation across 117 anatomical structures, outperforming architectural and supervised baselines on overall and most label-level comparisons, including few-shot settings. When transferred through Auto3DSeg, CT-FM weights improved tumor segmentation for hepatic, vessel-adjacent hepatic, pancreatic, and lung tumors, with gains varying by metric and site.

  • Segmentation of 117 anatomical structures in whole-body CT scans: CT-FM achieved a mean Dice coefficient of 0.8981 (95% CI: 0.8959-0.9004), exceeding the architectural baseline at 0.8959 and SuPREM at 0.8695.It also surpassed Auto3DSeg at 0.882 and VISTA3D at 0.893.
  • Segmentation of 117 anatomical structures in whole-body CT scans: Foundational pre-training increased mean Dice from the baseline’s 0.9017 (95% CI: 0.8885-0.9149) to 0.9058 (95% CI: 0.8929-0.9186).The baseline also exceeded Merlin’s 0.862.
  • Segmentation of 117 anatomical structures in whole-body CT scans: CT-FM outperformed the baseline and SuPREM across ribs, muscles, cardiac structures, and vertebrae, while the baseline performed better in the organ group.Across individual labels, CT-FM performed better in 73.5% of cases than the baseline and 91.5% than SuPREM.
  • Segmentation of cancer across anatomical sites on CT scans: For hepatic tumors adjacent to vessels, CT-FM weights increased Dice to 0.709 (95% CI: 0.656-0.759) from 0.698 and reduced ASD to 7.4 (95% CI: 4.9-10.2) mm from 9.4 (95% CI: 4.7-16.4) mm.The separate task targeted heterogeneous tumors adjacent to hepatic vessels.
  • Segmentation of cancer across anatomical sites on CT scans: CT-FM weights improved hepatic tumor segmentation, raising Dice to 0.696 (95% CI: 0.612- 0.772) from 0.681 and reducing ASD to 2.8 (95% CI: 1.7-3.7) mm from 3.6 (95% CI: 2.0- 6.0) mm.These comparisons used Auto3DSeg initialized with CT-FM weights versus standard initialization.
  • Segmentation of cancer across anatomical sites on CT scans: In pancreatic tumors, CT-FM weights showed no significant Dice improvement, with 0.475 (95% CI: 0.395-0.549) versus 0.482 (95% CI: 0.402-0.561), but reduced ASD to 7.8 (95% CI: 4.8-11.5) mm from 13.8 (95% CI: 6.2-23.7) mm.Visual inspection indicated fewer false positives and more true positives with CT-FM weights.
  • Segmentation of cancer across anatomical sites on CT scans: For lung tumors, CT-FM weights improved Dice to 0.609 (95% CI: 0.445-0.754) from 0.532 and ASD to 48.3 (95% CI: 10.8-103.4) mm from 62.1 (95% CI: 20.3-114.7) mm.Visualizations showed reduced over-segmentation and improved coverage of tumor extent, although challenging sites were sometimes missed.

Triage classification in head CT · Medical image retrieval · Anatomical Clustering

CT-FM supported head-CT triage, content-based retrieval, and anatomical region awareness, outperforming or matching relevant baselines across transfer and embedding-based evaluations. Its embeddings retrieved scans effectively and formed clusters aligned with anatomical regions.

  • Triage classification in head CT: On SinoCT, CT-FM achieved F1 0.776 and AUC-ROC 0.836, below SuPREM’s F1 0.798 and AUC-ROC 0.868 but above non-pretrained baselines.The non-pretrained model achieved F1 0.754 and AUC-ROC 0.802.
  • Triage classification in head CT: In zero-shot CQ500 transfer, CT-FM achieved F1 0.754 and AUC 0.794, surpassing the non-pretrained model and closely matching SuPREM.SuPREM achieved F1 0.745 and AUC-ROC 0.793, while the non-pretrained model achieved F1 0.728 and AUC-ROC 0.767.
  • Medical image retrieval: For OrganMNIST3D retrieval at k=3, CT-FM achieved AP=0.932, HR=0.968, and F1=0.944 versus SuPREM’s AP=0.923, HR=0.959, and F1=0.935.CT-FM outperformed SuPREM across top-k percentages for scans with similar organ field-of-view.
  • Medical image retrieval: On 3D-MIR at k=3, CT-FM retrieved lesions at the same anatomical site with precision 0.916 and similar lesion groups with precision 0.595.The corresponding reported baselines were 0.902 for lesion-site retrieval and 0.534 for lesion-group identification.
  • Medical image retrieval: Across up to 10 3D-MIR matches, CT-FM improved average precision to 0.923 for lesion identification and 0.660 for lesion grouping, versus 0.914 and 0.646.The comparison used the dataset authors’ baseline.
  • Anatomical Clustering: Using OrganMNIST3D embeddings and unsupervised tSNE, CT-FM and SuPREM attributed clusters to anatomical regions and distinguished different anatomical regions.OrganMNIST3D contains resampled cubic volumes from eight anatomical regions with differing fields of view.

Semantic Concept Search … DISCUSSION

CT-FM learned anatomically meaningful and stable 3D embeddings, outperforming SuPREM in fine-grained semantic search and retaining robustness across test-retest scans. Ablations supported the pre-training design, while the discussion highlighted broad task performance, limitations, and open-sourced resources.

  • Semantic Concept Search: CT-FM showed superior anatomical concept consistency to SuPREM, linking heart, kidney, bowel, and cervical concepts across scans at the micro-level.Patch-based queries assessed whether embeddings of matching structures were more similar than embeddings of different structures.
  • Semantic Concept Search: 5.61 ± 5.44 cm versus 23.44 ± 7.06 cm, CT-FM achieved lower organ centroid distance than SuPREM; 97.7% versus 81.9% of matched patches remained within the same heart structure.The metric was organ centroid distance, and the same-structure comparison concerned matched embedding patches in target scans.
  • PCA semantic analysis: PCA visualization produced consistent anatomical colors across scans, with heart red, lungs blue, and bones green, linking embedding features to anatomical concepts.The first three principal components were visualized by assigning each component a color.
  • Saliency and stability of CT-FM features: Saliency mapping localized CT-FM features to lungs, heart, diaphragm, vertebral column, and pelvic girdle across respiratory, cardiovascular, gastrointestinal, and musculoskeletal systems.The method mapped feature-space deviations from occluded regions back to specific CT locations.
  • Saliency and stability of CT-FM features: CT-FM embeddings remained highly similar across RIDER test-retest scans despite acquisition-parameter variation and identified outliers associated with larger positioning differences.The result indicates robustness across repeated scans while retaining sensitivity to registration or positioning differences.
  • Pre-training ablations: 0.257, 0.3246, and 0.154, intra-sample objective modifications improved micro Dice scores in SimCLR, SimSiam, and VicReg, respectively.Increasing SimCLR contrastive crops from N=5 to N=15 increased micro Dice scores from 0.654 to 0.732.
  • Pre-training ablations: The best macro- and micro-averaged Dice scores occurred at epoch 449, showing that longer training did not necessarily produce optimal transfer representations.The ablations examined representation quality across training checkpoints.

METHODS

The study assembled a large Imaging Data Commons CT pre-training cohort and multiple task-specific datasets spanning segmentation, head CT classification, retrieval, and semantic analysis. These datasets included annotated scans, external test cohorts, and standardized preprocessing and splitting protocols.

  • Pre-training dataset: 148,000 CT scans from 81,148 studies and 32,643 patients formed the Imaging Data Commons pre-training dataset.The cohort was selected using quality-based criteria and included 69 cohorts.
  • Whole-body segmentation: TotalSegmentator comprised 1,228 CT scans annotated for 117 structures, split into 928 training, 52 validation, and 248 testing scans.The scans originated from 1,368 CT examinations and were manually segmented by physicians with review and correction.
  • Tumor segmentation: Medical Segmentation Decathlon data covered four tumor cohorts: hepatic, hepatic tumor adjacent to vessels, pancreas, and lung.The cohorts included annotated liver, vessel-adjacent liver, pancreatic, and lung tumors, with custom task-specific splits and Auto3DSeg preprocessing.

Pre-training of the CT Foundation Model.

CT-FM used intra-sample contrastive learning, modifying SimCLR to contrast patches within the same CT scan and learn invariance at an anatomy-informed concept level. The model was trained with augmented 3D patches, a SegResNet encoder, and scan-balanced objectives.

  • Intra-sample contrastive learning: CT-FM replaced batch-derived negatives with patches from within the same sample to address limited variation between 3D medical-imaging scans and views.This modification was termed intra-sample contrastive learning.
  • Patch sampling and augmentation: 24x128x128 (z,y,x) patches were sampled from CT scans after resampling to 3mm slice thickness and 1mm in-plane resolution.The concept level was a 12.8cm x 12.8cm x 6.4cm cube, approximately the dimensions of the human heart.
  • Patch sampling and augmentation: Each selected patch generated two augmented views using standard and medical-imaging-specific transformations before encoding.Augmentations included random resize and crop, histogram shifting, random affine transformations, and intensity scale variations.
  • Encoder and objective: A SegResNet encoder produced 512-dimensional latent representations, which a projection head mapped into the similarity-learning space.The normalized temperature-scaled cross-entropy objective used cosine similarity, with temperature value 0.1.

Adaptation of the Vision Foundation model

CT-FM was adapted through end-to-end fine-tuning for segmentation and classification, with task-specific architectural additions. Retrieval used precomputed training embeddings and found minimum patch-wise aggregation to perform best.

  • General adaptation: CT-FM was fine-tuned end-to-end for segmentation and classification tasks, adapting all layers.Segmentation added the SegResNet decoder because pre-training trained only the encoder.
  • Whole body segmentation adaptation: Whole-body segmentation used a custom fine-tuning algorithm drawing on TotalSegmentator and MONAI Auto3DSeg design choices.Training used 300 epochs, AdamW with learning rate 0.0002, patch size [96, 160, 160], and a weighted Dice–cross-entropy objective.
  • Multi-region tumor segmentation: Multi-region tumor segmentation plugged CT-FM directly into Auto3DSeg and fine-tuned it with a reduced learning rate to preserve pre-trained weight spaces.The CT-FM model used learning rate 0.0005, whereas baselines used Auto3DSeg’s default 0.0002.
  • Head abnormality binary classification: Head CT triage used a feature-extraction backbone and classification head, with adaptive average pooling, fully connected layers, and a binary-logit output.CT-FM used SegResNet as its backbone; preprocessing concatenated blood, subdural, stroke, and bone intensity windows.
  • Content-based Retrieval of organs and lesion characteristics: Minimum-value aggregation across patch embeddings provided the best retrieval across tested strategies, although differences between retrieval methods were small.Training embeddings were precomputed, test embeddings were computed on the fly, and matches were ranked by cosine similarity.

Analysis Metrics

The study evaluated algorithms with task-specific metrics spanning segmentation, classification, and retrieval. Whole-body segmentation used macro-averaged and label-specific Dice, while tumor, head abnormality, and retrieval tasks used complementary measures.

  • Whole-body multi-class segmentation was evaluated with macro-averaged Dice across all labels and specific label groups, alongside individual-label Dice comparisons.
  • Tumor segmentation used Dice for predicted tumor masks and average surface distance to assess segmentation-boundary efficacy.Average surface distance was described as a stringent boundary-focused metric.
  • Binary head-abnormality classification was evaluated using AUC and F1 scores.
  • Medical image retrieval was evaluated using average precision, hit rate, and F1 score.

Semantic Concept Clustering

Semantic concept clustering projects CT-FM’s 512-dimensional embeddings into two dimensions using t-stochastic Neighbourhood Embeddings, enabling analysis of anatomical and contextual organization across 11 organs.

  • Projection method: t-stochastic Neighbourhood Embeddings project CT-FM’s 512-dimensional embeddings into 2 dimensions using different perplexity settings.The resulting low-dimensional projections support clustered semantic concept analysis.
  • Anatomical contexts: The clustering dataset contains 11 anatomical organs with contexts describing abdominal location, left-right position, and nearby structures.Embeddings are colored according to the anatomical label they represent.
  • Cross-method analysis: Clusters formed from embeddings across several methods can be meta-analyzed to determine shared semantic organization.This analysis relies on clustered low-dimensional projections of high-dimensional embeddings.

Semantic Concept Search

CT-FM supports semantic concept search by matching a selected 3D patch with semantically similar regions across target CT scans. The framework uses patch embeddings and cosine-based comparison to produce heat maps for exploring high-similarity regions.

  • Semantic concept search: Semantic concept search performs content-based retrieval at the micro-scale of concepts such as organs within a CT scan.
  • Semantic concept search: Users select a 3D patch from any source-scan field of view and search for semantically similar regions in one or more target scans.
  • Semantic concept search: The framework embeds the selected patch, compares it with sliding-window patches across target scans, and identifies matches using cosine distance or similarity.
  • Semantic concept search: Best-match patches generate heat maps of cosine-similarity rankings, allowing users to scroll through target scans and inspect highly similar regions across the field of view.
  • Semantic concept search: Evaluation selected scans sharing the same body-part DICOM tag from the TotalSegmentator dataset.

Organ Centroid Distance … DATA AVAILABILITY STATEMENT

The paper introduces interpretability analyses for CT-FM, including organ-localization evaluation, PCA-based visualization of semantic concepts, and occlusion-based saliency estimation. It also documents LLM use and provides publicly accessible datasets and retrieval details to support reproducibility.

  • Organ Centroid Distance: Organ Centroid Distance (OCD) measures the distance between a target organ’s true center and the center of the best-matching semantic-search region.Smaller OCD indicates more accurate localization from a single reference point.
  • Organ Centroid Distance: OCD was computed across TotalSegmentator by matching a heart point from one thoracic scan against comparable scans containing heart segmentations.The source scan was randomly selected from scans imaging the thoracic region and containing a 3D heart segmentation.
  • PCA analysis of learned semantic concepts: Patches of size 24x64x64 from two chest scans with similar FOVs were embedded and reduced from 512 to 3 dimensions using PCA.The first principal component was thresholded to remove background variance before visualization.
  • PCA analysis of learned semantic concepts: Foreground PCA values were mapped to CIELAB, with components representing lightness, red-green, and blue-yellow variation, then overlaid on the original images.CIELAB was selected for perceptual uniformity, so perceived color changes correspond more closely to numerical changes.
  • Salient regions captured by the embeddings: The authors propose Occlusion-based Feature Deviation (OFD) to identify regions influencing CT-FM’s feature embeddings.The method addresses saliency-map sensitivity to selected hyperparameters observed with methods such as RELAX.
  • Salient regions captured by the embeddings: OFD slides an 8x8x8 occlusion window across an image, comparing original and occluded embeddings with cosine distance to quantify regional saliency.Greater cosine distance indicates greater contribution of that region to the overall embedding.
  • LLM Usage Disclosure: The authors state that LLMs were used only to edit and refine author-written content, which was thoroughly verified and reviewed.The disclosure distinguishes editing assistance from original authorship.
  • DATA AVAILABILITY STATEMENT: Pre-training and evaluation datasets are publicly available, with IDC query parameters and versioning documented for reproducible retrieval.Evaluation datasets include TotalSegmentator, Medical Segmentation Decathlon, SinoCT, CQ500, 3D-MIR, and OrganMNIST3D.

CODE AVAILABILITY STATEMENT

The CT-FM code is publicly available through GitHub and the accompanying project website. The repository provides pipelines for preprocessing, pre-training, transfer learning, and evaluation, using the open-source Lighter framework.

  • CT-FM code is publicly available on GitHub through the accompanying website https://aim.hms.harvard.edu/ct-fm.
  • The repository includes pipelines for data preprocessing, CT-FM pre-training, transfer learning across demonstrated use cases, and evaluation scripts for comparison metrics.
  • The training framework uses the in-house open-source framework Lighter, available at https://github.com/project-lighter/lighter.

EXTENDED DATA FIGURES

The extended data figures provide additional evaluations of CT-FM across segmentation, retrieval, anatomical clustering, pre-training, embedding stability, and semantic search workflows. They include detailed comparisons, ablations, failure cases, and test-retest analyses.

  • Segmentation: Additional whole-body CT segmentation comparisons evaluate CT-FM against baselines, including few-shot performance across structure groups.These analyses are presented for anatomical structures and the TotalSegmentator dataset.
  • Segmentation: Detailed MSD analyses provide patient-wise comparisons and visualize failure cases.The extended comparison focuses on CT-FM and other approaches on the MSD dataset.
  • Representation analysis: Extended retrieval and clustering figures examine content-based retrieval for anatomical regions and compare anatomical clustering between CT-FM and a baseline.These figures expand analysis of structural and anatomical representation quality.
  • Pre-training and stability: Pre-training analyses assess ablations, representation-quality evolution, intra-sample adaptation for SimCLR, and patch-wise embedding stability on RIDER test-retest scans.The figures address both training behavior and embedding consistency across repeated scans.
  • Semantic search: The semantic search application workflow selects source and target scans, identifies a region of interest, crops a 3D box, and returns search results.The workflow proceeds through five stated steps from source-scan selection to semantic-search results.
Loading 2501.09001v2…