Source-linked AI summary
The MYOSAIQ Challenge: Myocardial Segmentation with Automated Infarct Quantification
Olivier Bernard, William A. Romero R., Cyprien Bouton, Celia Goujat, Hang Jung Ling, Pierre-Marc Jodoin, Fumin Guo, Calder Sheagren, Graham Wright, Abdul Qayyum, Moona Mazher, Steven A. Niederer, Hairui Wang, Xiaomei Wu, Franz Thaler, Gernot Plank, Martin Urschler, Ricardo M. Rosales, Esther Pueyo, Nicolas Duchateau, Frederic Cervenansky, Patrick Clarysse, Loic Belle, Thomas Bochaton, Nathan Mewton, Magalie Viallon, Pierre Croisille
TL;DR
Automated myocardial infarction quantification remains difficult because LGE images contain ambiguous, artifact-prone lesion boundaries and current studies often lack heterogeneous data. The MYOSAIQ challenge evaluates six teams and foundation models on a diverse multicenter dataset, finding strong LV performance but persistent difficulty for MI and MVO segmentation.
Problem
Myocardial infarction quantification is clinically valuable but rarely routine because LGE images make myocardial and lesion boundaries difficult to segment accurately.
Method
The study releases a heterogeneous, fully annotated 439-image LGE MR dataset, reports six challenge teams, analyzes multiple segmentation factors, and compares challengers with SAM-based foundation models.
Results
Segmentation performance decreases from LV to MYO to MI, with average Dice scores of 0.927, 0.832, and 0.668, respectively; fine-tuned foundation models may help address limitations of UNet-inspired approaches.
Takeaways & Limitations
The benchmark indicates that LV segmentation is highly performant, whereas MI and MVO remain the most challenging structures for automatic quantification.
Takeaways & Limitations
Conclusions are limited by only six participating teams, fewer than 500 dataset samples, and unresolved inter-observer disagreement between two experts.
Abstract
from arXiv · showhide
Late gadolinium enhancement (LGE) cardiac magnetic resonance (MR) imaging is the modality of choice to assess myocardial infarction (MI) lesions. Nowadays MI volume quantification is not performed routinely in clinical practice. Numerous deep learning (DL) methods have been developed to automate the segmentation of the myocardium and infarct regions. However, most studies rely on relatively small datasets which typically undergo pre-processing steps to standardize images and focus on a specific phase of myocardial infarction following reperfusion therapy. These limitations have impeded the development of models that are generalizable across diverse conditions and thus suitable for routine clinical use. To advance research and establish benchmarks in generalizable learning for myocardial infarct quantification, this paper presents findings from the Myocardial Segmentation with Automated Infarct Quantification (MYOSAIQ) challenge. The dataset set up for the challenge combines 439 CMR volumes from two multicenter clinical trials, with representative data acquired in acute and chronic phases after acute MI. Data were acquired in 16 centers using MRI scanners from three different vendors. Six teams participated until the end of the challenge, employing various baseline models, data augmentation techniques, and confidence strategies. To enhance the significance of this study, we compare the challengers' results with those of fine-tuned foundation models. Our results indicate that well-designed UNet-based techniques outperform fully automatic foundation models for LGE MR segmentation. While the best methods achieve high-quality and stable delineations of the left ventricle and myocardium under various conditions, they remain improvable in accurately segmenting infarct regions.
1. Introduction
MYOSAIQ addresses the limited generalizability of prior LGE-MR infarct-segmentation studies by evaluating automatic methods on heterogeneous, multi-phase clinical data. The challenge releases a broad dataset, compares six teams and foundation models, and analyzes segmentation quality across several dimensions.
- Motivation: LGE-MR infarct quantification is clinically valuable but rarely routine because infarct and MVO boundaries are difficult to segment accurately.Blurred, time-dependent boundaries, artifacts, small lesion size, and substantial manual input complicate delineation.
- Generalization gap: Prior studies commonly use small, pre-processed datasets focused on a single infarction phase, limiting validation across realistic multicenter and multi-vendor conditions.Earlier challenges used datasets of no more than 45 patients, while broader validation across acute and chronic cases remains needed.
- Challenge scope: The MYOSAIQ challenge evaluates automatic deep-learning methods for myocardial LGE-lesion quantification across acute and delayed or chronic disease phases.Acute data represent 4–8 days after MI and reperfusion, whereas delayed or chronic data represent 1 and 12 months.
- Contributions: The released dataset contains 439 fully annotated LGE-MR images acquired across 16 centers and 3 vendors, without pre-processing, at acute and chronic time points.Its heterogeneity includes major LGE imaging variants and is intended to support generalization studies.
- Evaluation: Results from six teams are compared across baseline models, augmentation techniques, confidence strategies, and foundation models based on the SAM architecture.The analysis covers segmentation quality, reperfusion timing, CNN versus foundation models, and confidence strategies.
2. Challenge framework
The challenge framework combines 439 examinations from two multicenter cohorts, covering acute, one-month, and twelve-month post-MI phases. Its acquisition protocols and cohort structure preserve clinically relevant variation for evaluating generalization.
- Dataset composition: The complete dataset contains 439 examinations from two multicenter cohorts spanning D8, M1, and M12 post-MI and reperfusion assessments.D8 represents the acute phase; M1 and M12 represent later follow-up phases.
- D8: The D8 subgroup includes 123 patients, with 98 for training and 25 for testing, from the MIMI cohort.These images depict acute myocardial infarction within 8 days after MI.
- M1: The M1 subgroup includes 187 patients, with 155 for training and 32 for testing, from the HIBISCUS-STEMI cohort.Images were acquired one month after PCI and reperfusion.
- M12: The M12 subgroup includes 129 patients, with 105 for training and 24 for testing, from the HIBISCUS-STEMI cohort.Images were captured twelve months after PCI and reperfusion.
- Clinical cohorts: The source cohorts include a multicenter randomized MIMI trial and HIBISCUS-STEMI follow-up imaging at one and twelve months after acute MI.HIBISCUS-STEMI patients underwent coronary angiography, primary PCI, and longitudinal LGE cardiac MR imaging.
- Imaging protocol: LGE imaging uses inversion-recovery prepared T1-weighted sequences acquired about 10 minutes after gadolinium administration.Optimized settings can produce a five-fold intensity difference between viable and nonviable myocardium.
2.2 Data annotation procedure
Experts manually annotated four cardiac regions using 3D Slicer and a supervised FWHM-guided procedure. The resulting NIfTI labels distinguish the LV cavity, healthy myocardium, MI, and MVO.
- Expert annotation: Every LGE cardiac MR image was manually annotated in 3D Slicer by an experienced cardiac radiologist, with independent test-set annotations from a second expert.The second annotation enabled evaluation of inter-observer variability.
- Annotation guidance: The semi-automatic FWHM approach guided annotations, but experts systematically supervised and corrected its results.FWHM is described as currently recommended in the literature.
- Regions of interest: The annotation covers the LV cavity, entire LV myocardium, MI, and MVO as four principal regions.MYO includes healthy tissue and regions affected by MI or MVO when present.
- Contour rules: The LV cavity must be completely covered, including papillary muscles, while MI must lie inside the myocardium.These contouring rules constrain the spatial relationship between cardiac structures.
- MVO rule: MVO is defined as hypo-enhanced black regions surrounded by a bright rim, excluding black signals caused by noise or artifacts.This rule distinguishes MVO from non-pathological dark regions.
- Output format: Reference segmentations use NIfTI labels: background = 0, LV cavity = 1, healthy myocardium = 2, MI = 3, and MVO = 4.The numeric labels encode the five classes used in the stored masks.
2.3 Image-based clinically relevant descriptors
The clinically relevant descriptors characterize infarct burden by coronary territory and time after reperfusion. Resolution and slice-thickness distributions are also compared between training and test data across D8, M1, and M12.
- Lesion descriptors: Lesion descriptors are categorized by LAD, LCX, and RCA territories and by D8, M1, and M12 time points.This organizes infarct characteristics across vascular territories and longitudinal disease stages.
- Clinical measures: Infarct extent is expressed relative to total myocardium or total endocardial surface area.The table reports lesion size and endocardial surface length as proportions of corresponding cardiac structures.
- Image properties: In-plane resolution is approximately 1.3–1.9 mm and slice thickness is centered around 5 mm across training and test datasets.The reported similarity indicates matched spatial sampling between the splits.
- Infarct location: Figure 2 summarizes average infarct locations for LAD, LCX, and RCA infarcts after 8 days, 1 month, and 1 year post reperfusion.The averages are computed across subjects in each subgroup.
- Image properties: Figure 3 compares in-plane resolution and slice-thickness distributions between training and test datasets for D8, M1, and M12.The figure evaluates whether spatial sampling differs across experimental configurations.
2.4 Evaluation platform
The challenge evaluated segmentation geometrically and clinically using 3D masks, with separate rankings for the two metric groups. Clinical quality assessment also incorporated prediction confidence and accounted for missing MVO outputs.
- Dice, HD, and ASSD quantified 3D geometric segmentation quality for the LV cavity, full myocardium, MI, and MVO.
- CC, MAE, and LOA assessed extracted infarction volumes, whereas CRPS evaluated confidence quality in those predictions.
- The CRPS calculation compared predicted cumulative distributions with reference volumes across patients, using volumes in milliliters and an assumed upper bound of 599 mL.
- Missing MVO predictions received no corresponding metric values and were excluded from averages, while missed-case counts were reported for sensitivity assessment.
- Separate geometric and clinical leaderboards summed per-method ranks across metrics and structures, excluding MVO from ranking computation.
2.5 Backbones architectures
All participating teams used UNet-inspired architectures, spanning diverse model families and complexities. Several teams added cascaded stages to progressively segment cardiac structures and distinguish damaged tissue.
- All teams employed UNet-inspired backbones, including nnUNet, 2D and 3D UNet, UNet++, xLSTM-UNet, and MedNeXt.
- Base-model complexity ranged from 1.4 million parameters for P4 to 50.3 million for P5.
- P5 used eight UNet++ backbones and fused their class predictions with STAPLE before applying a multi-class continuous min-cut algorithm.
- P2, P5, and P6 cascaded stages, first detecting the LV and myocardium before segmenting healthy versus damaged tissue; D8 added a third MI–MVO separation stage.
- P2 and P5 trained cascaded models end-to-end, whereas P6 trained each base UNet separately.
2.6 Data augmentation
All participants used data augmentation to improve model development, combining spatial transformations with intensity-based changes. The preprocessing and synthetic-data pipeline included resampling, padded cropping, normalization, and label deformation.
- All participants applied data augmentation using spatial transformations or intensity-based techniques.
- Spatial augmentation increased sample variation through rotations, flips, scaling, or deformations, while intensity augmentation altered appearance without changing anatomy.
- Each sample underwent resampling, padded cropping, and Z-score normalization before model use.
- Synthetic data were generated by deforming training labels with rigid and non-rigid transformations.
2.7 Confidence strategy
Five teams introduced strategies to improve confidence estimates for clinical volume metrics. These strategies included calibration, probability-based confidence modeling, and ensemble methods.
- Five teams developed strategies to enhance confidence in clinical metric estimates.
- P1 applied temperature scaling to nnUNet outputs, then used calibrated probabilities to derive volume-wise probability and cumulative density functions for CRPS.
- P2, P3, P5, and P6 used different ensemble strategies to estimate prediction confidence.
2.8 Foundation models
The study evaluates foundation models for fully automatic LGE-MRI cardiac segmentation, adapting mcMedSAM by removing bounding-box prompts and fine-tuning its decoder. The approach tests whether foundation-model representations can support multi-structure segmentation without precise initialization.
- Foundation-model design: mcMedSAM extends MedSAM with structure-specific tokens and a hierarchical architecture for multi-class cardiac segmentation.Its original design segments multiple cardiac structures from a single left-ventricle-centered bounding box.
- Foundation-model design: Competitive performance relative to nnU-Net requires mcMedSAM’s bounding box to be positioned within 3 mm in the x and y axes.This spatial-accuracy requirement motivates evaluating the model without a bounding-box prompt.
- Study adaptation: The study removes the bounding-box prompt to evaluate mcMedSAM in a fully automated setting.The encoder inherited from MedSAM is frozen, while the decoder is fine-tuned on the MYOSAIQ training dataset.
3. Results
The challenge methods produced broadly comparable and stable results, but segmentation quality declined as the target became more complex, especially for infarct and MVO regions. Fully automatic foundation models lagged behind CNN-based methods on these difficult structures, whereas confidence quality depended strongly on the estimation strategy.
- MYOSAIQ challenge results: The six participant methods produced very similar scores across backbone architectures and data-augmentation strategies.Statistical analysis of geometrical metrics confirmed comparable performance levels.
- MYOSAIQ challenge results: Dice scores averaged 0.927 for LV, 0.832 for MYO, 0.668 for MI, and 0.606 for MVO.Corresponding average HD scores were 7.3, 9.6, 19.5, and 16.2 mm, respectively.
- MYOSAIQ challenge results: Clinical-score correlations averaged 0.977 for LV, 0.937 for MYO, and 0.890 for MI.The decreasing values follow the same increase in segmentation difficulty across structures.
- Foundation models’ results: Fully automatic foundation models were competitive for LV, slightly lower for MYO, and clearly inferior for MI and MVO compared with CNN-based models.The semi-automatic foundation model achieved the lowest tested HD values: 4.8 mm for LV, 7.4 mm for MYO, 8.4 mm for MI, and 5.2 mm for MVO.
- Effect of time elapsed since reperfusion therapy: LV and MYO Dice and HD scores remained stable over time, whereas MI performance dropped at day 8 and became more variable.The authors associate the greater day-8 MI complexity with the presence of MVO structures.
- Confidence in the clinical scores: Temperature scaling yielded CRPS scores 3 to 7 times lower than those of other participants.Ensemble strategies produced similar confidence scores to a team without a specific confidence-estimation strategy, while potentially improving segmentation accuracy.
4. Discussion and Conclusion
The challenge shows that LV segmentation is strong, while MYO, MI, and MVO remain progressively more difficult, especially for infarct-related regions. CNN-based methods remain competitive, foundation models offer potential but automatic approaches do not consistently improve performance, and generalization across sequence types is encouraging but bounded by study limitations.
- Model comparisons: CNN-based approaches produce comparable geometric scores across backbone architectures and augmentation strategies, suggesting UNet-inspired models may have reached a performance plateau.The study also reports that performance is largely similar regardless of when reperfusion therapy occurred, with some differences for MI.
- Segmentation performance: MI and MVO remain the most challenging structures because of lower contrast and higher variability, requiring further segmentation improvement.Best-performing methods are competitive with experts for LV and MYO volume estimation, but MI and MVO still need improvement.
- Model comparisons: Fully automatic mcMedSAM accurately segments the LV but does not improve other structures over CNN methods, whereas fine-tuned MedSAM reaches performance comparable to inter-observer variability.The semi-automatic model uses a lightweight decoder but depends strongly on bounding-box initialization.
- Confidence estimation: Temperature scaling outperforms ensemble-based uncertainty estimation, yielding substantially lower CRPS scores through improved probability calibration.The authors attribute this to limited ensemble diversity and the dominance of aleatoric uncertainty from ambiguous MI–MVO boundaries.
- Generalization: Methods generalize across 2D PSIR and 3D-IR-GRE sequences, with ASSD differences below 0.4 mm for all anatomical structures despite sequence imbalance.The test set contains 93% 3D-IR-GRE and 7% 2D-PSIR sequences.
Ethical Standards
The study follows appropriate ethical standards for research involving human or animal subjects. It also complies with applicable laws and regulations governing such research.
- Ethical compliance: The research follows appropriate ethical standards for conducting the study and writing the manuscript.
- Ethical compliance: The study follows all applicable laws and regulations concerning the treatment of human or animal subjects.
- Ethical compliance: Ethical compliance covers both research conduct and manuscript preparation.