Source-linked AI summary
Patch-Based Diffusion Reconstruction for Accelerated Cardiac Cine
Xuan Lei, Philip Schniter, Juliet Varghese, Rizwan Ahmad
TL;DR
Highly accelerated 2D RT cine CMR reconstruction requires a method that can recover image detail from undersampled measurements across varied acquisition settings. The paper introduces CineDiff, a patch-based diffusion reconstruction framework using diffusion posterior sampling for data consistency, and reports improved reconstruction quality across retrospective and prospective evaluations, including mid-field and porcine data.
Problem
Highly accelerated 2D RT cine CMR reconstruction must recover high-quality images from undersampled measurements across variable acquisition settings.
Method
CineDiff combines an unconditional patch-based spatiotemporal diffusion model with diffusion posterior sampling to enforce consistency with measured k-space data.
Results
CineDiff achieved superior or competitive image quality compared with CS and CineVN across retrospective and prospective undersampled studies.
Takeaways & Limitations
CineDiff enabled high-quality reconstruction of highly accelerated 2D RT cine CMR and showed robustness to mid-field and porcine acquisitions.
Takeaways & Limitations
The image-quality assessments do not directly assess clinical utility.
Abstract
from arXiv · showhide
Purpose: To develop and evaluate a diffusion-based reconstruction framework for highly accelerated 2D real-time (RT) cine cardiovascular magnetic resonance imaging (CMR). Methods: We trained an unconditional patch-based diffusion model and incorporated it into a reconstruction framework, termed CineDiff, using diffusion posterior sampling for data consistency. CineDiff was evaluated in four settings: (i) 30 retrospectively undersampled breath-held cine at 1.5T and 3T from healthy participants across multiple acceleration rates, (ii) 15 prospectively undersampled free-breathing RT cine at 1.5T and 3T from patients indicated for clinical CMR, and (iii) 10 prospectively undersampled mid-field (0.55T) free-breathing scans, including five from healthy subjects and five from porcine models. For retrospective undersampling, reconstruction quality was assessed using peak signal-to-noise ratio (PSNR), structural similarity index measure (SSIM), learned perceptual image patch similarity (LPIPS), and deep image structure and texture similarity (DISTS). For prospective undersampling, image quality was evaluated by blinded expert scoring on a 5-point Likert scale. Results: In retrospectively undersampled breath-held cine data, CineDiff achieved higher PSNR and SSIM and lower LPIPS and DISTS than the comparison methods across all evaluated acceleration rates. In prospectively undersampled free-breathing RT cine data, CineDiff received higher expert image-quality scores. Qualitatively, CineDiff reduced block-like artifacts and preserved finer anatomical detail compared with traditional compressed sensing and a variational network method, termed CineVN. Conclusion: CineDiff enabled high-quality reconstruction of highly accelerated 2D RT cine CMR. The method also demonstrated robustness to out-of-distribution data, including mid-field and porcine acquisitions.
Abbreviations
The paper uses abbreviations for cardiac CMR acquisition, reconstruction methods, evaluation metrics, and MRI concepts.
- CMR denotes cardiovascular magnetic resonance, while RT denotes real-time imaging.
- CineVN denotes a variational network for cardiac cine, and CS denotes compressed sensing.
- DPS denotes diffusion posterior sampling, while DDPM and DDIM denote diffusion-model formulations.
- PSNR, SSIM, LPIPS, and DISTS are the reported image-quality metrics.
1 Introduction
High acceleration is necessary for real-time cine CMR but challenges conventional reconstruction quality and flexibility. The paper introduces CineDiff, a patch-based diffusion framework intended to improve quality across varied acquisition settings.
- Motivation: Real-time cine CMR requires substantial k-space undersampling to maintain spatial and temporal resolution.
2 Methods
CineDiff addresses the ill-posedness of highly accelerated RT cine reconstruction with a patch-based diffusion prior and measurement-guided data consistency. Its design models local spatiotemporal structure while enforcing agreement with the measured k-space data.
- Problem: High acceleration makes the RT cine inverse problem ill-posed, limiting least-squares reconstruction quality.Compressed sensing addresses this with regularized least squares and a sparsity-promoting prior.
- Measurement-guided reconstruction: Diffusion posterior sampling alternates a generative diffusion update with data-consistency correction through the MRI forward model.The data-consistency step is applied after converting patches back into full image series.
- Patch-based diffusion prior: CineDiff trains a diffusion denoiser on overlapping spatiotemporal patches rather than full cine images.Patch modeling supports simpler distributions and allows more unique training patches to be extracted from each cine series.
- Patch aggregation: CineDiff uses ramp-weighted averaging when aggregating overlapping patches to suppress boundary artifacts.The patch-to-image operator is applied before k-space data consistency is enforced.
- Inference acceleration: A CS-based warm start can reduce diffusion inference steps by supplying a preliminary reconstruction for subsequent refinement.This strategy is used to reduce the number of time steps during inference.
2.4 Implementation details of CineDiff
CineDiff was trained on fully sampled cine data after preprocessing complex-valued temporal inconsistencies and used normalized overlapping patches for reconstruction. Evaluation combined retrospective image metrics with blinded expert scoring of prospective scans across field strengths and acquisition settings.
- Training data and preprocessing: Training used fully sampled cine data from the OCMR and CMRxRecon repositories.The training data underwent preprocessing to correct temporal phase and magnitude inconsistencies in complex-valued cine images.
- Training data and preprocessing: The correction procedure addressed frame-dependent complex scaling ambiguities that produced phase drift or inconsistent magnitude.These artifacts included intermittent sign flips and frame-to-frame fluctuations in phase and magnitude.
- Model implementation: The denoiser used overlapping 64 × 64 × 8 spatiotemporal patches with two channels representing real and imaginary image components.Patch selection was biased toward central anatomical structures using a truncated 2D Gaussian, and each series was normalized by its maximum magnitude.
- Inference implementation: Eight posterior samples were averaged pixel-wise to improve image SNR, requiring approximately 5 to 8 minutes per series in Study I.The samples used different noise realizations and were averaged as complex-valued image series.
- Evaluation: Retrospective evaluation used PSNR, SSIM, LPIPS, and DISTS, while prospective free-breathing data were scored blindly on 5-point image-quality and sharpness scales.Retrospective metrics focused on the central cardiac region, and three reconstruction methods were presented in randomized side-by-side order for reader assessment.
- Evaluation: The evaluation included retrospective accelerations R ∈ {8, 12, 16, 20}, prospective 1.5T and 3T scans at R = 9, and 0.55T scans at R = 10.The mid-field dataset comprised five healthy-volunteer slices and five porcine-study slices.
3 Results
CineDiff generally outperformed CS and CineVN across retrospective and prospective real-time cine evaluations. Its strongest results combined CS initialization with posterior averaging, while preserving sharper anatomical detail and reducing artifacts.
- CineDiff generally outperformed CS and CineVN across evaluated configurations and acquisition settings.
- Retrospectively undersampled data: T = 50 with CS initialization and NAVG = 8 yielded the highest PSNR and SSIM among CineDiff configurations.T = 50 uses truncated sampling and offers faster generation than T = 1,000.
- Retrospectively undersampled data: CS initialization with NAVG = 8 preserved finer structures and produced sharper reconstructions than CS and CineVN at R = 16.
- Prospectively undersampled data: CineDiff achieved the highest overall quality and sharpness scores in prospectively undersampled 1.5T/3T free-breathing cine.Scores were averaged across 15 cine series and two experienced readers.
- Prospectively undersampled data: At 0.55T, CineDiff achieved the highest human-study scores and matched CineVN in porcine overall quality while leading in sharpness.
- Prospectively undersampled data: Representative 0.55T human and porcine examples showed sharper boundaries and improved visualization of anatomical details.
4 Discussion
CineDiff's discussion emphasizes strong reconstruction quality across acceleration, prospective acquisition, field strength, and species. Posterior averaging, warm starts, patch-based training, and data consistency shape its quality, efficiency, and generalization.
- Framework design: Patch-based training addresses limited training data, while ramp-weighted de-patchification suppresses patch-boundary artifacts.Simple averaging can produce structured discontinuities at patch edges.
- Retrospective reconstruction: At R = 16, CineDiff achieved PSNR/SSIM of 30.54/0.9149, versus 29.36/0.9007 for CineVN and 28.67/0.8676 for CS.
- Retrospective reconstruction: Posterior averaging improved PSNR and SSIM but slightly smoothed images, whereas NAVG = 1 produced lower LPIPS/DISTS and higher perceptual realism.At R = 16, LPIPS/DISTS were 0.0737/0.0968 for NAVG = 1 and 0.0755/0.1058 for NAVG = 8.
- Sampling efficiency: CS warm-starting reduced reverse diffusion steps by a factor of 20 while consistently improving PSNR and SSIM.The comparison was T = 50 versus Gaussian initialization at T = 1,000.
- Prospective reconstruction: CineDiff achieved the highest expert scores for image quality and sharpness in prospectively undersampled free-breathing studies without a reference standard.In Study II, mean scores were 4.97/4.93 for CineDiff, compared with 4.27/4.30 for CineVN and 4.30/4.27 for CS.
- Generalization: CineDiff generalized to 0.55T human and porcine acquisitions despite training exclusively on 1.5T/3T breath-held cine data.In Study III, it outperformed both CS and CineVN in human subjects and porcine models.
5 Limitations
The study's prospective evaluations used relatively few slices, reconstruction remained computationally demanding, and image-quality assessments did not directly establish clinical utility. Patch and sampling design choices were empirical, leaving room for further optimization.
- Dataset scope: Prospective Studies II and III included relatively few slices, limiting statistical power and potentially restricting generalizability.Larger, more diverse, multicenter datasets are needed to confirm robustness across patient populations and acquisition settings.
- Computational cost: CineDiff reconstruction remained computationally demanding, requiring several minutes per cine series on a high-end GPU despite CS warm-start.The warm-start reduces diffusion steps, but posterior averaging still contributes to reconstruction time.
- Clinical validation: The evaluation used image-quality metrics and blinded reader scores that do not directly assess clinical utility.Additional validation on ventricular volumes, ejection fraction, and regional wall motion is needed.
- Design choices: Patch size, overlap, and sampling parameters were selected empirically rather than through systematic optimization.These choices may affect the trade-off between local detail preservation and global consistency, and further tuning may improve performance.
6 Conclusions
The paper presents CineDiff, a patch-based diffusion reconstruction framework for highly accelerated 2D real-time cine CMR. Across retrospective and prospective undersampling, it improved image quality and sharpness, including in mid-field and porcine acquisitions.
- Framework: CineDiff is a patch-based diffusion reconstruction framework for highly accelerated 2D real-time cine CMR.The framework trains on complex-valued spatiotemporal patches and uses ramp-weighted de-patchification to reduce boundary artifacts.
- Results: Across retrospective and prospective undersampled studies, CineDiff consistently improved image quality and sharpness at high acceleration rates.The reported improvements included mid-field and porcine acquisitions.
- Scope: CineDiff demonstrated robust reconstruction performance in challenging acquisition settings, including mid-field and porcine studies.These settings were included among the prospective undersampled evaluations.
Ethics Declarations
The paper reports no competing interests, documents institutional review board approval and informed consent for human data, and states that code and a sample dataset will be publicly available after peer review.
- Competing interests: The authors declared no competing interests.
- Ethics approval: Human-subject data received Institutional Review Board approval from The Ohio State University.The cited approval identifiers are 2020H0402 and 2019H0076.
- Consent: Informed consent to participate and publish results was obtained from all individual participants.
- Data and code: Code and a sample dataset will be made publicly available after the peer-review process is complete.
- Author contributions: The paper identifies contributions spanning CineDiff implementation, optimization, study design, image analysis, and supervision.
A Patchification and De-patchification Operations
CineDiff converts spatiotemporal image series into overlapping patches for inference and reconstructs them with ramp-weighted de-patchification. The weighting scheme combines overlapping pixels while reducing boundary artifacts.
- Patchification: Training samples 64×64×8 patches using spatially central selection from a truncated Gaussian and uniform temporal sampling.The sampling strategy preferentially selects patches from the image's central region.
- Patchification: During inference, the full spatiotemporal image series is converted into overlapping 64 × 64 × 8 patches.The number of patches is kept as small as possible along each spatiotemporal dimension.
- Operations: Inference uses patchification P(·) and de-patchification P−1(·) to move between the image series and patch representations.
- De-patchification: De-patchification combines overlapping pixels with a ramp-weighted scheme rather than simple neighboring-patch averaging.The approach is illustrated for one dimension using patch index, pixel location, and relative patch weights.
- De-patchification: Ramp-weighted averaging effectively eliminates boundary artifacts from the patches.
B Supplementary figures
The supplementary figures compare CS, CineVN, and CineDiff reconstructions across retrospective and prospective real-time cine settings, including human and porcine acquisitions at 0.55T. They emphasize error or time-profile comparisons and details better preserved in CineDiff.
- Retrospective reconstruction: Reference images, R = 16 reconstructions, error maps, and time profiles are compared in one supplementary figure.The time profiles are measured along a dashed red line.
- Prospective human reconstruction: CS, CineVN, and CineDiff are compared on prospectively undersampled free-breathing real-time cine data.The figure includes time profiles along the dashed red line.
- Mid-field reconstruction: A supplementary figure shows CS, CineVN, and CineDiff reconstructions from a free-breathing real-time cine acquisition in a human subject at 0.55T.Time profiles are included along the dashed red line.
- Porcine reconstruction: The porcine 0.55T comparison includes reconstructions and time profiles, with red arrows marking details better preserved in CineDiff.The acquisition used free-breathing real-time cine imaging.