Source-linked AI summary

Advancing machine learning for MR image reconstruction with an open competition: Overview of the 2019 fastMRI challenge

Florian Knoll, Tullie Murrell, Anuroop Sriram, Nafissa Yakubova, Jure Zbontar, Michael Rabbat, Aaron Defazio, Matthew J. Muckley, Daniel K. Sodickson, C. Lawrence Zitnick, Michael P. Recht

arXiv:2001.02518v1eess.IVcs.CV

TL;DR

The paper addresses the need to evaluate machine-learning MR reconstruction methods in large-scale, realistic settings rather than only on small individual studies. It describes an open challenge using clinical knee k-space data, multiple coil tracks, and quantitative plus radiologist evaluation. The challenge received 33 submissions, and winning deep-CNN entries outperformed human performance, while unstable behavior for severe abnormalities remained a hurdle for clinical adoption.

  • Problem

    Prior machine-learning MR reconstruction studies were trained and validated on small individual datasets, limiting evaluation in large-scale realistic settings.

  • Method

    The challenge provided raw k-space data from 1,594 consecutive clinical knee exams and evaluated multi-coil and single-coil reconstruction tracks using quantitative metrics followed by radiologist assessment.

  • Results

    33 challenge submissions were received, and current winning deep-CNN entries outperformed human performance.

  • Takeaways & Limitations

    The challenge advanced machine-learning MR reconstruction, provided insight into the field’s state of the art, and highlighted hurdles for clinical adoption.

  • Takeaways & Limitations

    MR reconstruction methods can react unpredictably and unstably to cases with severe abnormalities.

Abstract

from arXiv · show

Purpose: To advance research in the field of machine learning for MR image reconstruction with an open challenge. Methods: We provided participants with a dataset of raw k-space data from 1,594 consecutive clinical exams of the knee. The goal of the challenge was to reconstruct images from these data. In order to strike a balance between realistic data and a shallow learning curve for those not already familiar with MR image reconstruction, we ran multiple tracks for multi-coil and single-coil data. We performed a two-stage evaluation based on quantitative image metrics followed by evaluation by a panel of radiologists. The challenge ran from June to December of 2019. Results: We received a total of 33 challenge submissions. All participants chose to submit results from supervised machine learning approaches. Conclusion: The challenge led to new developments in machine learning for image reconstruction, provided insight into the current state of the art in the field, and highlighted remaining hurdles for clinical adoption.

1 | INTRODUCTION

The introduction frames fastMRI as an open competition intended to address limited accessibility and reproducibility in medical image reconstruction research. It aimed to stimulate machine learning for MR reconstruction through realistic data, broad participation, and radiologist evaluation.

  • Background: Medical imaging research has adopted machine learning and data science, while broader machine learning has achieved major advances in vision and other domains.Deep convolutional neural networks have driven progress from image classification to championship-level gaming.
  • Research gap: Prior reconstruction methods were often trained and validated on small, privately collected datasets, limiting reproducibility and comparison across approaches.Restricted data access also limited participation to researchers connected with institutions possessing imaging data.
  • Motivation: The fastMRI challenge sought to stimulate machine learning research in MR image reconstruction aimed at reducing MR examination times.The project followed the release of a large-scale database of raw MRI scanner data from clinical patients.
  • Contribution: The challenge provided a large-scale, realistic setting where researchers could evaluate reconstruction methods using clinical radiologist assessment.It was designed to offer both realistic data and access to evaluation by clinical radiologists.
  • Contribution: The article describes the challenge’s design, results, and lessons learned from organizing it.The stated scope covers both the challenge outcomes and organizational lessons.

2 | METHODS

The challenge used realistic raw MR k-space data and multiple tracks to test image-reconstruction methods while balancing clinical relevance, validation, and accessibility. Evaluation combined quantitative image metrics with radiologist assessment.

  • Challenge design: 1,594 consecutive clinical knee MRI acquisitions supplied realistic raw k-space data for image reconstruction.The dataset included multiple contrasts, scanners, field strengths, and clinical multi-channel receive coils.
  • Challenge design: Retrospective undersampling of fully sampled acquisitions provided matched ground truth for reconstruction comparisons.Allowed undersampling patterns were predefined, using one-dimensional pseudo-random phase-encoding sampling with a fully sampled central k-space region.
  • Evaluation: SSIM, NRMSE, and PSNR supported online quantitative ranking, while musculoskeletal radiologists determined the final winners.The radiologist panel was added because quantitative metrics provide limited insight into diagnostic quality.
  • Operating modes: The challenge tested acceleration factors of R=4 and R=8, with R=4 selected for the single-coil track.The lower acceleration scenario targeted challenging but potentially clinically acceptable reconstructions, whereas R=8 was intended to probe performance beyond reasonable limits and analyze failure modes.
  • Challenge tracks: Multiple submission tracks balanced realistic multi-coil reconstruction with a lower-barrier single-coil setting.The multi-coil track used true multi-channel scanner data, while the single-coil track simplified data handling for groups from machine learning, computer vision, and image processing.

3 | RESULTS

The challenge attracted 33 submissions using supervised machine learning, and results varied across multi-coil and single-coil tracks. Quantitative metrics and radiologist rankings showed both agreement and important discrepancies across tracks.

  • Participation: All submissions used exclusively fastMRI data for training their approaches.
  • Qualitative findings: A subchondral osteophyte and a meniscal tear were not well seen or visible in the accelerated reconstructions.The osteophyte was not visible in any accelerated reconstruction, while the meniscal tear was not well seen in any.
  • Quantitative evaluation: 0.924, 0.895 and 0.707 were the average SSIM values for multi-coil R=4, multi-coil R=8 and single-coil R=4, respectively.The reported averages showed a substantial difference between the multi-coil and single-coil tracks.
  • Quantitative evaluation: The multi-coil R=8 submission with SSIM=0.874 significantly outperformed the highest-ranking single-coil R=4 submission.
  • Radiologist evaluation: In multi-coil R=8, the highest-ranked submission also had the highest SSIM, RMSE and PSNR, whereas single-coil R=4 aligned with radiologist rankings only for SSIM.For multi-coil R=4, the top four submissions were close across metrics, with SSIM differences below 1%.
  • Radiologist evaluation: Radiologists strongly preferred one submission in multi-coil R=8 and single-coil R=4, ranking it first by 5 readers and second by the remaining 2.In multi-coil R=4, rankings were less consistent: the two highest-rated submissions were each ranked worst by one reader, while the lowest-rated submission was ranked best by 3 readers and worst by 4.
  • Radiologist evaluation: Radiologist rankings generally followed summed category scores, with all readers doing so for multi-coil R=8, 5 of 7 for multi-coil R=4, and 6 of 7 for single-coil R=4.The categories were artifacts, sharpness, perceived contrast-to-noise ratio and diagnostic confidence, rated on a 4-point scale where 1 was best.

4 | DISCUSSION

The challenge showed that supervised deep-learning methods can perform strongly across reconstruction tracks, while exposing metric, data-distribution, pathology, and clinical-evaluation limits that constrain conclusions about generalization and adoption.

  • Dataset and generalization: The challenge did not systematically separate anatomical or pathological variation across training, validation, and test sets.The authors suggest future challenges could partition data by pathology, age, height, weight, body mass index, or gender.
  • Dataset and generalization: Randomly partitioned coronal knee data from a limited set of scanners from one vendor substantially limited insight into robustness and generalization.The challenge also lacked substantial variation in receive-coil geometries, using standard knee arrays from one vendor.
  • Clinical translation: Radiologist evaluation should occur at diagnostic-interpretation level, because image-quality ratings alone did not assess diagnostic interchangeability with fully sampled reconstructions.The authors identify diagnostic interpretation as necessary for domain knowledge to provide substantial additional information and regard clinical translation as unresolved.
  • Evaluation: The first- versus fourth-ranked multi-coil submission differed by less than 1%, and top-team differences were almost negligible across all three tracks.In multi-coil R=4, radiologist preferences substantially disagreed, with the lowest-ranked submission preferred by 3 of 7 radiologists and ranked worst by 4.
  • Evaluation: SSIM aligned with radiologist preferences in tracks with clear winners, but no quantitative metric reproduced radiologists’ rank order in every track.For single-coil R=4, SSIM followed the radiologist trend while RMSE and PSNR showed almost opposite trends; radiologists still rated image quality subjectively rather than diagnostic interchangeability.
  • Methods and submissions: Multi-coil and single-coil tracks posed substantially different inverse problems, yet most groups used essentially the same core method after fine-tuning or retraining.The single-coil problem is undetermined, whereas the multi-coil problem is overdetermined but ill-posed because coil elements are not independent.
  • Clinical translation: No submission deteriorated for the severe metallic-implant artifact case in multi-coil R=8, although dedicated studies are required to investigate robustness.The authors describe this result as encouraging while noting that subtle pathology was not correctly identified even in multi-coil R=4 results.
  • Methods and submissions: 33 submissions all used supervised deep neural-network approaches, and winners combined learned reconstruction with additional data or model-based elements rather than purely end-to-end learning.Six of eight multi-coil participants also submitted to the single-coil track, and the top three single-coil submissions came from groups submitting to multi-coil.
Loading 2001.02518v1…