Source-linked AI summary
CARDINAL Predicts Cardiovascular Risk From Non-contrast Cardiac CT
Roy Gabriel, Nattakorn Kittisut, Jamshid Hassanpour, Michael Galarnyk, Abanoub Abdelmalak, Marly van Assen, Carlo N. De Cecco, Arshed Quyyumi, Ali Adibi
TL;DR
Cardiovascular risk stratification is limited by incomplete clinical data and handcrafted imaging features. CARDINAL learns clinically grounded compact representations from routine non-contrast cardiac CT, improving long-horizon MACE prediction and risk stratification beyond established clinical and imaging baselines.
Problem
Current risk stratification is imperfect, while engineered CT features may not fully leverage information in the full CT volume.
Method
CARDINAL uses clinically supervised nested latent representations to encode routine electrocardiogram-gated non-contrast cardiac CT while preserving anatomical and calcium-related information.
Results
CARDINAL improved MACE discrimination across horizons relative to PCE, PREVENT, CAC, and CT biomarker baselines, with its strongest overall performance at 10 years.
Takeaways & Limitations
Non-contrast cardiac CT contains prognostic information beyond conventional risk equations, CAC scoring, and engineered imaging biomarkers.
Takeaways & Limitations
Gains were strongest at longer horizons and in follow-up-restricted analyses, while motion, truncation, reconstruction-kernel shifts, and external scanner variation require dedicated validation.
Abstract
from arXiv · showhide
Cardiovascular risk prediction remains limited by incomplete clinical data and imaging biomarkers that reduce computed tomography (CT) to a small number of handcrafted features. We developed CARDINAL (Cardiovascular Assessment via Representation learning from Deep Imaging with Nested Anatomical Latent embeddings), a clinically grounded framework that learns compact representations from routine non-contrast cardiac CT for major adverse cardiovascular event (MACE) prediction. In 17,659 patients, CARDINAL was evaluated for 1-, 3-, 5-, and 10-year MACE prediction against American Heart Association (AHA) pooled cohort equations (PCE), AHA predicting risk of cardiovascular disease events (PREVENT), coronary artery calcium (CAC), segmentation-derived CT biomarkers, and 70-feature structural radiomics. Gains were largest at longer horizons. At 10 years, CARDINAL (joint) achieved an area under the receiver operating characteristic curve (AUROC) of 0.866 $\pm$ 0.020 and an area under the precision-recall curve (AUPRC) of 0.890 $\pm$ 0.015, compared with an AUROC of 0.826 $\pm$ 0.023 and an AUPRC of 0.826 $\pm$ 0.022 for structural radiomics, the strongest baseline. CARDINAL also achieved the highest survival concordance index (C-index), 0.753 $\pm$ 0.015, and high-versus-low risk-tertile hazard ratio, 10.78 $\pm$ 3.16, with favorable reclassification and exploratory calibration. These findings suggest that non-contrast cardiac CT contains prognostic information beyond conventional risk equations, CAC scoring, and engineered imaging biomarkers.
1 Introduction
CARDINAL addresses limitations of clinical risk equations and handcrafted CT biomarkers by learning compact, clinically grounded representations directly from non-contrast cardiac CT for MACE prediction.
- Clinical risk equations rely on variables that may be missing, outdated, or misaligned with the imaging encounter.
- CAC captures calcified plaque burden but reduces the broader cardiothoracic information in CT to a scalar summary.
- Most CT prediction pipelines use segmentation and predefined features, which may not fully leverage information in the complete CT volume.
- CARDINAL learns a low-dimensional representation from electrocardiogram-gated non-contrast cardiac CT using nested embeddings and anatomy- and calcium-related supervision.
- Frozen CARDINAL embeddings were evaluated for MACE classification across 1-, 3-, 5-, and 10-year horizons and for time-to-event modeling.
- The framework was designed to preserve clinically meaningful information while improving MACE discrimination, survival discrimination, and patient-level risk stratification relative to conventional and engineered baselines.
2 Results
Across held-out evaluations, CARDINAL preserved clinically meaningful phenotypes and showed strongest performance at longer MACE horizons, with advantages over several clinical and imaging baselines. It also supported survival risk stratification, compact latent representations, reclassification, calibration, and subgroup analyses, while performance varied by horizon and evaluation convention.
- Cohort and evaluation: 17,659 patients remained after eligibility filtering, with 12,370 training, 1,768 validation, and 3,521 held-out test patients.
- Clinically grounded representations: At latent dimension d = 8, CARDINAL (joint) achieved ICCs of 0.885 ± 0.007 for heart volume, 0.907 ± 0.012 for lung volume, and 0.873 ± 0.007 for myocardium volume.
- Clinically grounded representations: 0.882 ± 0.011 and 0.891 ± 0.025 were the ICCs for CAC volume and CAC score, respectively, in the fused CARDINAL (MoE) representation.
- MACE classification: At 10 years under withFU, CARDINAL (joint) achieved AUROC = 0.866 ± 0.020 and AUPRC = 0.890 ± 0.015, exceeding structural radiomics at AUROC = 0.826 ± 0.023 and AUPRC = 0.826 ± 0.022.
- MACE classification: At shorter horizons, CARDINAL variants remained competitive, while PREVENT BOTH achieved the highest AUROC in the 1-year analyses.
- Survival and risk stratification: CARDINAL (joint) achieved C-index = 0.753 ± 0.015 and the largest high-versus-low risk-tertile HR of 10.78 ± 3.16.
- Compact representations: Approximately 96% of peak 10-year withFU AUROC was retained by d = 64 for CARDINAL (joint) and d = 128 for CARDINAL (MoE).
- Reclassification and calibration: At 10 years under withFU, CARDINAL (joint) showed positive NRI against every non-CARDINAL comparator.
3 Discussion
CARDINAL uses clinically grounded latent representations from routine non-contrast cardiac CT to improve long-horizon MACE prediction, survival discrimination, and patient-level risk stratification beyond conventional scores, CAC, and engineered CT biomarkers. Its strongest results occurred in 10-year follow-up-restricted analyses, while performance depended on follow-up quality and preserved anatomy.
- Results: CARDINAL’s strongest gains occurred at longer horizons, particularly with follow-up-restricted labeling, whereas gains were more modest at shorter horizons and under ignoreFU.WithFU restricts evaluation to patients observed through the target horizon or an earlier event; ignoreFU includes a larger cohort but introduces greater label uncertainty.
- Representation learning: The clinically grounded latent space preserved heart, lung, myocardial, aortic, and calcium-related phenotypes while supporting downstream MACE prediction.Anatomy- and calcium-grounded supervision shaped the nested representation, strengthening the interpretation that it captures clinically meaningful cardiothoracic information.
- Representation learning: The nested representation retained at least 95% of maximum 10-year withFU AUROC at 64 joint dimensions and 128 MoE dimensions, enabling compact storage and reuse.The 1024-dimensional representation reduces an approximately 1.6-million-voxel input by more than 1,500-fold.
- Model comparison: The joint model was strongest overall, especially at 10-year withFU, while the MoE model achieved the highest AUROC at 3- and 5-year ignoreFU horizons.These results indicate complementary contributions from globally integrated and organ-specific image information.
- Translation: CARDINAL’s multisite, multivendor cohort included 11 sites and scanners from five manufacturer labels, supporting relevance to varied U.S. imaging practice.The study also emphasizes opportunistic cardiovascular risk assessment from scans already acquired without additional radiation or acquisition.
- Limitations and future work: External validation, prospective impact studies, and evaluation under motion, truncation, reconstruction-kernel shifts, and external scanner variation remain necessary.Calibration improvements were exploratory because no independent calibration cohort was used, and robustness depended on preserved anatomy.
4 Methods
This retrospective cohort study developed CT-based models to predict MACE from non-contrast cardiac CT and tested whether a clinically grounded CT representation improves prediction against established approaches. The analysis used linked clinical and imaging data from 17,659 patients, with patient-level separation of training, validation, and held-out test sets.
- Study design and population: The pipeline learned CT representations, predicted MACE downstream, and evaluated discrimination, calibration, reclassification, and survival modeling.All modeling and evaluation steps used strictly separated patient-level splits to prevent information leakage.
- Study design and population: The study followed TRIPOD+AI reporting guidance and was interpreted using PROBAST+AI principles.The retrospective study used de-identified clinical and imaging data, and the institutional review board waived informed consent.
- Study design and population: 17,659 patients formed the final analytic cohort after eligibility assessment and index-examination selection.Adults underwent electrocardiogram-gated non-contrast cardiac CT across 11 affiliated sites, with imaging linked to longitudinal EHR data.
- Study design and population: Patients with inadequate imaging, preprocessing incompatibility, or insufficient data linkage were excluded, and no formal prospective sample-size calculation was performed.Study size was determined by all eligible patients in the retrospective source cohort.
- Study design and population: The cohort was partitioned into training (70%), validation (10%), and held-out test (20%) sets with stratification by MACE outcome and baseline calcium burden.No patient appeared in more than one partition during model development or evaluation.
4.4 Outcome definition
MACE was defined using cardiovascular events and all-cause mortality identified from linked clinical data, with manual review validating outcome extraction. Horizon-specific classification and survival analyses used distinct follow-up and censoring rules.
- Event definition: MACE comprised stroke, myocardial infarction, late percutaneous coronary intervention, late coronary artery bypass grafting, or all-cause mortality.Revascularization events counted when occurring more than 90 days after the index CT.
- Event definition: Diagnosis and procedure codes supplemented by mortality data identified events, while clinicians manually reviewed over 10% of the cohort.Review included every event category and non-event cases and was independent of downstream model predictions.
- Horizon-based classification: Horizon labels at 1, 3, 5, and 10 years were positive when an event occurred within the specified time window.These labels supported horizon-based MACE classification.
- Horizon-based classification: The withFU strategy included patients observed through the horizon or an earlier event, whereas ignoreFU treated absent recorded events as negative regardless of follow-up completeness.Thus, withFU used observed event-free follow-up for negative labels, while ignoreFU introduced greater label uncertainty.
- Survival analysis: Survival time ran from index CT to first MACE or last documented follow-up, with right-censoring for patients without observed events and administrative censoring at 10 years.Survival eligibility was independent of horizon-specific classification labels.
4.5 Clinical and imaging baselines
CARDINAL was compared with clinical risk equations, CAC, segmentation-derived CT biomarkers, and structural radiomics using common held-out patients. The baselines represented clinical variables, calcium burden, anatomical measurements, and 70 engineered morphology features.
- Clinical baselines: PREVENT ASCVD was the primary clinical comparator, calculated from EHR variables recorded at or within 6 months before index CT.Inputs included demographics, body mass index, blood pressure, lipids, kidney function, treatments, diabetes, and smoking.
- Clinical baselines: PCE and other PREVENT risk-equation results were retained for supplementary comparison, with cases missing required inputs excluded from the corresponding analysis.Variable-specific available-case counts were reported separately.
- Imaging baselines: CAC burden was quantified using the Agatston score from a validated automated pipeline.CAC provided a calcium-based imaging comparator.
- Imaging baselines: The CT biomarker baseline used six segmentation-derived features, while structural radiomics used 70 three-dimensional shape features across five cardiothoracic structures.The six features included organ volumes, CAC volume, and CAC score; radiomics covered the aorta, atria, and ventricles.
- Comparative evaluation: All displayed models were evaluated on the same held-out patients with available predictions for each horizon and labeling convention.Models included CARDINAL joint and MoE variants, structural radiomics, CT biomarkers, PREVENT variants, and CAC score.
4.6 CT preprocessing
The preprocessing pipeline standardized non-contrast cardiac CT volumes and used anatomical and calcium-derived masks to guide processing and auxiliary representation-learning targets. CARDINAL encoded the resulting volumes into compact latent representations without using MACE outcomes.
- Image standardization: CT volumes were converted to NIfTI, reoriented to a canonical anatomical axis, intensity-normalized, and resampled to fixed spatial resolution.These steps standardized imaging inputs before model processing.
- Image standardization: Volumes were resized to a consistent grid, and deterministic slice selection preserved cardiothoracic structures and global thoracic context.The same preprocessing and quality-control criteria were applied across demographic groups.
- Anatomical guidance: TotalSegmentator masks and CAC annotations guided preprocessing and supplied auxiliary anatomical targets for representation learning.These targets incorporated anatomical and calcium-related information into training.
- Representation learning: CARDINAL used a Swin Transformer and Matryoshka Representation Learning to encode non-contrast CT into compact nested latent representations.Progressively larger latent prefixes retained additional clinically meaningful anatomical and calcium-related information.
- Representation learning: Representation learning used anatomical and calcium-related CT targets rather than MACE outcomes, preventing label leakage from downstream prediction tasks.The representation was therefore clinically grounded without direct access to MACE labels during pretraining.
4.8 Joint and mixture-of-experts (MoE) representations
CARDINAL representations were constructed either as a single shared latent representation or through specialized expert encoders combined by learned fusion. Encoders were trained without MACE outcomes and frozen before downstream modeling.
- Joint representation: A joint formulation trained one encoder to predict six anatomical and calcium-related targets, producing one unified nested latent representation.The targets were aorta, heart, lung, and myocardium volumes, plus CAC volume and CAC score.
- Single-expert representation: Six separate encoders independently represented aorta volume, heart volume, lung volume, myocardium volume, CAC volume, and CAC score.
- Mixture-of-experts representation: The fused/MoE formulation concatenated six 1024-dimensional expert embeddings and learned a multilayer module to produce one 1024-dimensional fused representation.This fused representation was comparable with the joint representation and did not average expert-level phenotype metrics.
- Downstream use: All encoders and the fusion model were trained without MACE outcomes and frozen before downstream event modeling.The frozen encoders were applied to all CT studies to extract latent embeddings.
- Downstream use: Downstream selection used predefined model classes, including FT-Transformer, multilayer perceptron, XGBoost, CatBoost, and random forest.Survival modeling used penalized Cox proportional hazards models with elastic-net regularization.
4.10 Survival modeling
Survival analyses modeled time to MACE using frozen latent representations and baseline inputs as Cox covariates. Performance was assessed through discrimination, risk separation, calibration, and reclassification procedures across fixed horizons.
- Modeling: Penalized Cox proportional hazards models with elastic-net regularization modeled time to MACE across five fixed training-validation partitions.Each model was evaluated on the same held-out test partition.
- Modeling: Patients without observed events were right-censored at last follow-up, with follow-up beyond 10 years administratively censored.The same fitted Cox model generated predicted risks at 1, 3, 5, and 10 years.
- Evaluation: Harrell’s C-index quantified overall survival discrimination, while Brier scores at 1, 3, 5, and 10 years were secondary metrics.Latent representations and baseline inputs served as Cox covariates.
- Evaluation: High-versus-low predicted-risk tertile hazard ratios, two-sided log-rank tests, and cumulative-incidence separation quantified risk stratification.Patient risks were averaged across five fixed models before division into model-specific tertiles.
- Evaluation: Classification evaluation included AUROC, AUPRC, sensitivity, specificity, F1 score, Brier score, calibration measures, and decision-curve analysis.
- Reclassification: Continuous, category-free NRI quantified whether CARDINAL assigned higher risk to patients with events and lower risk to patients without events than comparators.Positive NRI values favored CARDINAL, with confidence intervals estimated by event-stratified bootstrap.
- Calibration: Raw calibration, decision-curve analysis, and reclassification used held-out test predictions, while Platt and isotonic calibrators were fitted across prediction folds.Platt scaling preserves rank-based discrimination; isotonic score ties can slightly change rank metrics.
4.12 Subgroup analyses
Prespecified subgroup analyses evaluated held-out test predictions across sex, age, race, and baseline CAC burden. Robustness testing stressed the joint representation using imaging-relevant perturbations.
- Subgroup analyses: Subgroup analyses stratified held-out test predictions by sex, age, race, and baseline CAC burden.The groups included women and men; age below or at least 60 years; White and Black or African American participants; and CAC = 0, 1–99, or ≥100.
- Subgroup analyses: Within each labeling convention and horizon, all models were evaluated on the same patients in each subgroup.AUROC was computed for CARDINAL joint, CARDINAL MoE, and each available baseline.
- Robustness analysis: Robustness testing assessed the 1024-dimensional joint representation across five held-out folds and six phenotype targets.
- Robustness analysis: Intensity scaling, slice dropout, and spatial masking represented acquisition or coverage stresses with direct imaging analogues.Gaussian-noise perturbations were excluded from the reported analysis.
4.14 Statistical analyses
Reported results aggregate performance across independently trained fixed-fold models and use paired held-out predictions for reclassification and statistical comparisons. Analyses used prespecified tests and significance thresholds, with supplementary materials providing additional results.
- Reporting: Classification and survival metrics were computed for five independently trained fixed-fold models and reported as mean ± standard deviation.
- Reclassification: Continuous NRI used event-stratified nonparametric bootstrap resampling of paired patient-level predictions to estimate 95% confidence intervals.The analyses were category-free and did not depend on prespecified clinical risk thresholds.
- Statistical testing: Pairwise AUROC comparisons used two-sided DeLong tests, while paired model-level metrics used exact two-sided Wilcoxon signed-rank tests.Survival risk groups used two-sided log-rank tests and Cox model contrasts, with statistical significance defined as α = 0.05.
- Implementation: All analyses were conducted using Python with standard scientific computing libraries.
- Supplementary analyses: Supplementary information provides additional classification, reclassification, survival, calibration, subgroup, phenotype-recovery, and robustness results.
5 Declarations
The study reports institutional review board oversight with waived informed consent, a related provisional patent application, and research funding received by Dr. De Cecco.
- The retrospective study received institutional review board oversight, with informed consent waived.
- A provisional patent application related to the work has been filed.
- Dr. De Cecco received research funding from Siemens, Cleerly, Elucid, and Pfizer.