Source-linked AI summary
A cross-modal generative model for incomplete and degraded prostate MRI with multicentre clinical validation
Siyuan Ma, Liang He, Mengying Zhu, Yi Chai, Mengyao Lyu, Haowei Wang, Qizhen Lan, HaoBo Sun, Qixin Zhang, Jingli Chen, Xiaobing Wei, Jiaming Liu, Guiqin Liu, Qianwen Zhang, Yang Liu, Dacheng Tao, Guangyu Wu
TL;DR
Missing or degraded sequences can limit prostate multiparametric MRI. This study evaluates MSCNet, a sequence-conditioned cross-modal generative framework, which achieved higher cross-task similarity than comparators while supporting image-quality non-inferiority for DWI, ADC and T2W but not T1W.
Problem
Missing or degraded sequences limit prostate multiparametric MRI, motivating evaluation of reconstruction for unavailable contrasts and unreliable acquisitions.
Method
MSCNet reconstructs unavailable contrasts and restores artefact- or degradation-affected acquisitions using complementary prostate MRI sequences.
Results
0.818 versus 0.798 mean SSIM across ten completion tasks; image-quality non-inferiority was met for DWI, ADC and T2W, but not T1W.
Takeaways & Limitations
These retrospective findings support quality-controlled cross-modal reconstruction as an adjunct to acquired prostate MRI while retaining acquired examinations as the clinical reference.
Takeaways & Limitations
Generated ADC maps require separate calibration for quantitative measurement, and missing anatomical coverage and dynamic contrast kinetics cannot be recreated.
Abstract
from arXiv · showhide
Missing or degraded sequences can limit prostate multiparametric MRI. We developed MSCNet, a sequence-conditioned cross-modal generative framework for reconstructing unavailable contrasts and restoring degraded acquisitions. Across ten completion tasks, task-specific MSCNet achieved mean structural similarity of 0.818 versus 0.798 for the strongest task-matched comparators; matched-capacity analyses showed larger differences in lesion fidelity and boundary preservation. In a blinded 1,000-case reader study, overall image quality met the prespecified non-inferiority criterion for DWI, ADC and T2W completion, but not T1W. In a separate 200-case diagnostic assessment, AUCs for clinically significant cancer were 0.860 with acquired images, 0.841 with MSCNet and 0.797 with baseline-generated images. A locked 186-case three-hospital cohort supported multicentre transportability. These retrospective results support quality-controlled cross-modal reconstruction as an adjunct to acquired prostate MRI.
Results
MSCNet achieved the highest reconstruction fidelity across ten completion tasks, improved degraded acquisitions while reducing safety events, and narrowed the quality gap to acquired images. Diagnostic discrimination remained close to acquired imaging and exceeded baseline-generated images, with multicentre analyses supporting transportability.
- Sequence completion: MSCNet achieved the highest SSIM across all ten completion tasks and mean SSIM of 0.818 for optimized task-specific models.The model outperformed the strongest task-matched comparators, which achieved mean SSIM of 0.798.
- Anatomical fidelity: Lesion fidelity increased from 0.87 to 0.92 and boundary SSIM from 0.853 to 0.908 versus the parameter-matched Transformer, while LPIPS decreased from 0.108 to 0.086.Regional analyses also found lower absolute residuals within the gland and boundary ring, with adjacent-slice consistency of 0.942 versus 0.908 for DynUNet and 0.874 for the GAN baseline.
- Safety: Any safety event occurred in 9.4% of MSCNet outputs versus 28.7% of baseline outputs, with lower rates of false lesions, reduced conspicuity, blurred margins, partial erasure and non-diagnostic outputs.The corresponding rates were 1.4% versus 5.8%, 3.6% versus 11.2%, 5.1% versus 15.6%, 2.3% versus 8.4%, and 3.2% versus 9.3%, respectively.
- Clinical validation: AUCs for clinically significant prostate cancer were 0.860 with acquired images, 0.841 with MSCNet images and 0.797 with baseline-generated images.The MSCNet-minus-reference difference was −0.019, compared with −0.063 for the strongest baseline; the pooled three-hospital external cohort had SSIM 0.791, LPIPS 0.124 and MAE 0.061.
Discussion
The discussion frames incomplete prostate mpMRI as information routing, clinical fidelity and selective use, while showing persistent but task- and sequence-specific benefits and defining boundaries for shared models, validation and deployment. It supports reliability-gated cross-modal reconstruction, with prospective workflow, economic and calibration studies still needed.
- Task-specific MSCNet retained the highest SSIM after comparison with unified synthesis, matched-capacity Transformers, diffusion, convolutional, retrieval and physics-aware baselines.The persistent advantage under stronger controls indicates that capacity and optimization explain only part of the result.
- 0.018 mean SSIM advantage occurred in the physically coupled group, versus 0.024 in cross-contrast T2W synthesis and 0.020 for T1W completion.DWI and ADC share diffusion-derived information, whereas T2W reconstruction reflects anatomy; SSIM changes are not universally clinically meaningful.
- 158.3-million-parameter shared MSCNet reproduced all ten tasks with a 0.015 mean SSIM reduction versus optimized task-specific models.Unified pretraining followed by task-specific fine-tuning recovered the optimized task-specific mean, but zero-shot unseen-combination performance remained lower.
- 0.714 to 0.883 SSIM improvement in 204 same-patient repeat acquisitions supported recovery from genuine degraded scans against a clean repeat reference.Repeat examinations can differ in positioning, physiology and interval change, so the retrospective reference is not perfect ground truth.
- DWI, ADC and T2W met overall image-quality non-inferiority, whereas T1W did not; diagnostic confidence showed the same sequence-specific pattern.In the 200-case analysis, the MSCNet-versus-acquired-image AUC difference was small and its confidence interval remained above the supportive −0.05 boundary, but formal non-inferiority depends on prespecified boundary and power.
- 0.79 to 0.92 diagnostic usability increase occurred when coverage decreased from full coverage to the validation-selected 50% operating point.Externally transferring native threshold 𝜏= 0.50 retained 48.9% of cases, with failure detection 0.75, false rejection 0.16, retained-case AUC 0.86 and ECE 0.036; multicentre calibration and prospective workflow studies remain necessary.
Online Methods … MSCNet
MSCNet was developed and evaluated using public and institutional prostate MRI cohorts, with patient-level splits, sequence-specific completion tasks, standardized preprocessing, correspondence checks and quality control. The architecture fused modality-specific multiscale features with location-dependent gating, attention-based decoding, deep supervision and an edge branch, while MSCNet-Shared used masked modality-conditioned training.
- Cohorts, governance and task definition: Public PI-CAI and PROSTATEx datasets and multiple Ren Ji Hospital cohorts supported development, reader evaluation, diagnostic assessment and artefact analysis.The clinical artefact cohort contained 182 unique patients and 262,306 DICOM slices.
- Cohorts, governance and task definition: Development splits were patient-level, while external, reader, diagnostic and repeat-scan cohorts were excluded from principal model-weight selection.Hospital data were retrospectively de-identified and processed under institutional ethics approval.
- Cohorts, governance and task definition: MSCNet independently trained principal models for three-sequence and four-sequence target configurations, with MSCNet-Shared using one masked, target-conditioned model across tasks.The shared experiment used an available-modality mask, target-modality token and training-time modality dropout.
- Preprocessing, matching and quality control: DICOM series underwent metadata identification, manual verification, orientation standardization, common-grid resampling, within-examination alignment and sequence-specific intensity normalization.The prostate region and adjacent context were retained, while cases with specified file, metadata, registration, annotation or diagnostic issues were excluded.
- Preprocessing, matching and quality control: Sequence correspondence was checked at patient, examination, slice and anatomical-region levels using rigid or affine alignment followed by deformable correction when required.Quality metrics included landmark displacement, normalized mutual information, prostate-region Dice and centre displacement.
- MSCNet: Modality-specific residual encoders with squeeze-and-excitation recalibration extracted features at five scales with 32, 64, 128, 256 and 512 channels.A gating network allowed each sequence’s contribution to vary by feature scale and spatial location.
- MSCNet: The deepest fused representation passed through a multi-head self-attention bottleneck, while decoder attention gates combined upsampled and fused encoder features.Intermediate-scale auxiliary predictions enabled deep supervision, and an edge branch compared prediction and acquired-target gradients.
- MSCNet: MSCNet-Shared represented absent sequences with zero-filled tensors, binary availability masks and a learned target token, and used valid-combination sampling with modality dropout.The shared model used 158.3 million parameters; unified pretraining used 3,750 development cases.
Optimization and comparison models
MSCNet was optimized with a weighted composite reconstruction loss and trained using a specified AdamW regimen. Comparisons included legacy GAN, convolutional, and diffusion models alongside expanded unified, Transformer, and diffusion-based alternatives.
- Optimization: The objective combined character, SSIM, MS-SSIM, perceptual, edge, wavelet, frequency and deep-decoder terms with specified weights.The loss was L = Lchar + 0.15LSSIM + 0.10LMS-SSIM + 0.02Lperc + 0.05Ledge + 0.10Lwave + 0.05Lfreq + 𝜆dsLdeep.
- Optimization: 154.1 million parameters and 200 epochs defined task-specific MSCNet training with AdamW, a 2 × 10−4 initial learning rate, batch size 2, cosine annealing and 10 warm-up epochs.𝜆ds denoted the normalized contribution of auxiliary decoder outputs.
- Comparison models: The legacy comparison set comprised DynUNet, Pix2Pix, a residual generator, LSGAN with PatchGAN and spectral normalization, and conditional diffusion with 1,000 diffusion steps and 50-step DDIM sampling.DynUNet was implemented in MONAI.
- Comparison models: The expanded comparison set added unified missing-modality, multi-contrast Transformer, modality-masked diffusion, structure-aware latent-diffusion, parameter-matched Transformer, wide DynUNet and DynUNet t
Image, spatial and lesion-level evaluation
The evaluation assessed image fidelity, spatial error, volumetric consistency, boundary preservation, lesion fidelity, and lesion localization using complementary quantitative and reader-based measures.
- Image and spatial evaluation: Image fidelity was measured with PSNR, SSIM, MS-SSIM, LPIPS, FID and MAE.These metrics quantified reconstructed-image fidelity.
- Image and spatial evaluation: Spatial evaluation partitioned error into gland, boundary-ring and peri-gland regions and stratified results by apex, mid-gland and base.This captured regional and anatomical-location differences in error.
- Image and spatial evaluation: Volumetric consistency and boundary preservation were evaluated using adjacent-slice similarity, slice-to-slice variation, volumetric smoothness, z-trend correlation, edge preservation, boundary SSIM, gradient similarity and capsule sharpness.The measures targeted interslice coherence and preservation of anatomical boundaries.
- Lesion-level evaluation: Lesion fidelity was assessed with contrast-to-noise ratio, signal-to-noise ratio, contrast ratio, line-profile agreement and radiologist conspicuity, while localization used Dice overlap, centre distance, volume overlap and boundary distance.Evaluation definitions and aggregation rules were provided in Supplementary Note 5.
Artefact transfer and same-patient repeat-scan evaluation
The study evaluated artefact restoration using clinically observed degradation patterns transferred onto anatomically corresponding clean images, and separately assessed degraded scans against same-patient repeat acquisitions. The repeat-scan analysis included 204 degraded–repeat pairs from 182 unique patients, with clinical and acquisition correspondence recorded.
- Artefact transfer: Five clinically observed artefact patterns were categorized and transferred in image or frequency space after screening for topology and texture.The categories were susceptibility, motion ghosting, banding, Gibbs ringing and chemical shift.
- Artefact transfer: Pairing required anatomical correspondence, lesion-boundary preservation and SSIM of approximately 0.8 between clean targets and controlled degraded images.
- Same-patient repeat-scan evaluation: 204 degraded–repeat pairs from 182 unique patients formed the distinct repeat-scan analysis.The clean reference was a same-patient repeat acquisition obtained within four weeks; one patient could contribute more than one adjudicated stratum.
- Same-patient repeat-scan evaluation: Seven repeat-scan degradation categories were recorded alongside treatment status, interval, scanner and protocol correspondence.Categories included motion blur, susceptibility distortion, low-SNR high-b-value DWI, geometric distortion, zipper artefact, ghosting and radiofrequency inhomogeneity.
Reader study, diagnostic assessment and external evaluation
The evaluation combined a blinded, block-randomized reader study, a biopsy- or follow-up-based diagnostic assessment, and external transportability analyses without hospital-specific model adjustment. Image quality, diagnostic discrimination and calibration were assessed using prespecified or exploratory benchmarks as applicable.
- Reader study: 1,000 held-out cases were assessed across four reader-study tasks, with 250 cases per task and de-identified, block-randomized acquired, MSCNet and baseline images.Three radiologists scored overall quality, anatomical fidelity, lesion conspicuity, diagnostic confidence and artefact absence on five-point scales.
- Reader study: The reader study compared acquired-reference images using a prespecified −0.5-point margin for primary task-level differences and 95% confidence intervals.Readers completed sessions with washout before scoring the five-point endpoints.
- Diagnostic assessment: 200 cases with biopsy or clinical follow-up supported diagnostic assessment using PI-RADS version 2.1 categories and binary clinically significant prostate cancer judgements.AUCs were compared with paired methods, while calibration used expected calibration error, calibration slope and Brier score; the −0.05 AUC boundary was exploratory and non-confirmatory.
- External evaluation: No hospital-specific training, fine-tuning or threshold recalibration was performed in the external evaluation, with site-level transportability results reported separately.Completion and diagnostic endpoints were reported separately even when the same case contributed to both.
Safety adjudication, uncertainty and rejection · Statistical analysis and reporting
Safety adjudication defined reconstruction-related events and attributed them to acquisition, lesion, registration/model, mixed or unassigned contributors. Quality control used a prespecified reconstruction-risk score, validation-selected thresholds and held-out internal and external evaluation with case-level statistical reporting.
- Safety adjudication, uncertainty and rejection: Safety events included suspected false lesions, reduced conspicuity, blurred margins, partial erasure, non-diagnostic outputs and input-output inconsistency.Radiologists adjudicated each event and assigned acquisition-dominant, lesion-dominant, registration/model-dominant, mixed or unassigned contributors.
- Safety adjudication, uncertainty and rejection: U(x) was a normalized voxel-wise reconstruction-risk map, fitted and operating-point selected using development and validation data only.The score was fixed before held-out internal and external evaluation and interpreted operationally as reconstruction risk, not as aleatoric or epistemic uncertainty decomposition.
- Safety adjudication, uncertainty and rejection: The primary case-level score was the maximum 95th-percentile value across the global image, prostate boundary, peripheral zone and transition zone.Lesion-region summaries were retrospective and required lesion annotations; failure detection used AUROC for motion artefact, sequence mismatch, boundary blur and lesion reconstruction error.
- Safety adjudication, uncertainty and rejection: Thresholds from 0.2 to 0.7 were assessed using failure-detection sensitivity.The supplied passage identifies the threshold range and sensitivity measure but does not provide further operating-point results.
- Safety adjudication, uncertainty and rejection: 0.5 was the locked internal threshold, selected with conservative, balanced and permissive policies on validation data without test-set retuning.It was transferred directly to the 186-case external cohort without hospital-specific recalibration, with coverage, retained count, failure detection, false rejection, retained-case AUC and calibration reported jointly.
- Statistical analysis and reporting: Paired t-tests, Wilcoxon signed-rank tests, Spearman’s ρ, paired DeLong or bootstrap procedures, weighted κ and ICC were used for prespecified outcome types.Continuous summaries used means with standard errors or medians with interquartile ranges, while patient clustering was retained for repeat acquisitions.
- Statistical analysis and reporting: Held-out internal and external risk–coverage results used validation-selected thresholds without test-set retuning, with external transport assessed using the native internal threshold.Summaries were calculated at case level, task groups were prespecified as physically coupled DWI–ADC, cross-contrast T2W and auxiliary T1W, and no formal health-economic, cost-effectiveness or decision-curve analysis was conducted.
Ethics approval and consent
The retrospective institutional study received ethics approval, followed the Declaration of Helsinki and applicable data-protection requirements, and had written-consent requirements waived for de-identified data.
- Ethics approval and consent: Ethics approval was granted by Ren Ji Hospital’s Ethics Committee (KY2018-212-s2), with retrospective use of de-identified institutional data exempted from written informed consent.The study complied with the Declaration of Helsinki and applicable data-protection requirements.
Data availability
PI-CAI and PROSTATEx are available through public repositories under applicable access terms. Institutional imaging and clinical data are restricted because they contain protected health information, with access requests reviewed by the relevant data-access committee.
- PI-CAI and PROSTATEx are available from their public repositories subject to applicable access terms.
- Institutional imaging and clinical data are not publicly released because they contain protected health information.
- Requests for de-identified institutional data should be directed to the co-corresponding authors and reviewed by the relevant institutional data-access committee.
Funding
The work was supported by Chinese national and municipal science funding agencies.
- Support came from the National Natural Science Foundation of China (No. 82371912) and the Science and Technology Commission of Shanghai Municipality (No. YDZX20243100003001).
Extended Data
The Extended Data define patient-level cohort splits and analytical roles, while summarizing quantitative, reader-study, and diagnostic-validation analyses. PROSTATEx is treated as a vendor-specific public cohort rather than independent external validation.
- Study cohorts: Patient-level splits define study cohorts and analytical roles, with PROSTATEx treated as a vendor-specific public cohort rather than independent external validation.
- Quantitative performance: The main quantitative MSCNet performance table reports PSNR, SSIM, LPIPS and MAE, with parenthetical values denoting absolute changes from the strongest non-MSCNet baseline.PSNR and SSIM are higher-is-better, whereas LPIPS and MAE are lower-is-better.
- Clinical validation: Clinical validation summarizes reader scores from 1,000 held-out cases and diagnostic discrimination from a separate 200-case set.