Source-linked AI summary

Identifying the Best Machine Learning Algorithms for Brain Tumor Segmentation, Progression Assessment, and Overall Survival Prediction in the BRATS Challenge

Spyridon Bakas, Mauricio Reyes, Andras Jakab, Stefan Bauer, Markus Rempfler, Alessandro Crimi, Russell Takeshi Shinohara, Christoph Berger, Sung Min Ha, Martin Rozycki, Marcel Prastawa, Esther Alberts, Jana Lipkova, John Freymann, Justin Kirby, Michel Bilello, Hassan Fathallah-Shaykh, Roland Wiest, Jan Kirschke, Benedikt Wiestler, Rivka Colen, Aikaterini Kotrotsou, Pamela Lamontagne, Daniel Marcus, Mikhail Milchenko, Arash Nazeri, Marc-Andre Weber, Abhishek Mahajan, Ujjwal Baid, Elizabeth Gerstner, Dongjin Kwon, Gagan Acharya, Manu Agarwal, Mahbubul Alam, Alberto Albiol, Antonio Albiol, Francisco J. Albiol, Varghese Alex, Nigel Allinson, Pedro H. A. Amorim, Abhijit Amrutkar, Ganesh Anand, Simon Andermatt, Tal Arbel, Pablo Arbelaez, Aaron Avery, Muneeza Azmat, Pranjal B., W Bai, Subhashis Banerjee, Bill Barth, Thomas Batchelder, Kayhan Batmanghelich, Enzo Battistella, Andrew Beers, Mikhail Belyaev, Martin Bendszus, Eze Benson, Jose Bernal, Halandur Nagaraja Bharath, George Biros, Sotirios Bisdas, James Brown, Mariano Cabezas, Shilei Cao, Jorge M. Cardoso, Eric N Carver, Adrià Casamitjana, Laura Silvana Castillo, Marcel Catà, Philippe Cattin, Albert Cerigues, Vinicius S. Chagas, Siddhartha Chandra, Yi-Ju Chang, Shiyu Chang, Ken Chang, Joseph Chazalon, Shengcong Chen, Wei Chen, Jefferson W Chen, Zhaolin Chen, Kun Cheng, Ahana Roy Choudhury, Roger Chylla, Albert Clérigues, Steven Colleman, Ramiro German Rodriguez Colmeiro, Marc Combalia, Anthony Costa, Xiaomeng Cui, Zhenzhen Dai, Lutao Dai, Laura Alexandra Daza, Eric Deutsch, Changxing Ding, Chao Dong, Shidu Dong, Wojciech Dudzik, Zach Eaton-Rosen, Gary Egan, Guilherme Escudero, Théo Estienne, Richard Everson, Jonathan Fabrizio, Yong Fan, Longwei Fang, Xue Feng, Enzo Ferrante, Lucas Fidon, Martin Fischer, Andrew P. French, Naomi Fridman, Huan Fu, David Fuentes, Yaozong Gao, Evan Gates, David Gering, Amir Gholami, Willi Gierke, Ben Glocker, Mingming Gong, Sandra González-Villá, T. Grosges, Yuanfang Guan, Sheng Guo, Sudeep Gupta, Woo-Sup Han, Il Song Han, Konstantin Harmuth, Huiguang He, Aura Hernández-Sabaté, Evelyn Herrmann, Naveen Himthani, Winston Hsu, Cheyu Hsu, Xiaojun Hu, Xiaobin Hu, Yan Hu, Yifan Hu, Rui Hua, Teng-Yi Huang, Weilin Huang, Sabine Van Huffel, Quan Huo, Vivek HV, Khan M. Iftekharuddin, Fabian Isensee, Mobarakol Islam, Aaron S. Jackson, Sachin R. Jambawalikar, Andrew Jesson, Weijian Jian, Peter Jin, V Jeya Maria Jose, Alain Jungo, B Kainz, Konstantinos Kamnitsas, Po-Yu Kao, Ayush Karnawat, Thomas Kellermeier, Adel Kermi, Kurt Keutzer, Mohamed Tarek Khadir, Mahendra Khened, Philipp Kickingereder, Geena Kim, Nik King, Haley Knapp, Urspeter Knecht, Lisa Kohli, Deren Kong, Xiangmao Kong, Simon Koppers, Avinash Kori, Ganapathy Krishnamurthi, Egor Krivov, Piyush Kumar, Kaisar Kushibar, Dmitrii Lachinov, Tryphon Lambrou, Joon Lee, Chengen Lee, Yuehchou Lee, M Lee, Szidonia Lefkovits, Laszlo Lefkovits, James Levitt, Tengfei Li, Hongwei Li, Wenqi Li, Hongyang Li, Xiaochuan Li, Yuexiang Li, Heng Li, Zhenye Li, Xiaoyu Li, Zeju Li, XiaoGang Li, Wenqi Li, Zheng-Shen Lin, Fengming Lin, Pietro Lio, Chang Liu, Boqiang Liu, Xiang Liu, Mingyuan Liu, Ju Liu, Luyan Liu, Xavier Llado, Marc Moreno Lopez, Pablo Ribalta Lorenzo, Zhentai Lu, Lin Luo, Zhigang Luo, Jun Ma, Kai Ma, Thomas Mackie, Anant Madabushi, Issam Mahmoudi, Klaus H. Maier-Hein, Pradipta Maji, CP Mammen, Andreas Mang, B. S. Manjunath, Michal Marcinkiewicz, S McDonagh, Stephen McKenna, Richard McKinley, Miriam Mehl, Sachin Mehta, Raghav Mehta, Raphael Meier, Christoph Meinel, Dorit Merhof, Craig Meyer, Robert Miller, Sushmita Mitra, Aliasgar Moiyadi, David Molina-Garcia, Miguel A. B. Monteiro, Grzegorz Mrukwa, Andriy Myronenko, Jakub Nalepa, Thuyen Ngo, Dong Nie, Holly Ning, Chen Niu, Nicholas K Nuechterlein, Eric Oermann, Arlindo Oliveira, Diego D. C. Oliveira, Arnau Oliver, Alexander F. I. Osman, Yu-Nian Ou, Sebastien Ourselin, Nikos Paragios, Moo Sung Park, Brad Paschke, J. Gregory Pauloski, Kamlesh Pawar, Nick Pawlowski, Linmin Pei, Suting Peng, Silvio M. Pereira, Julian Perez-Beteta, Victor M. Perez-Garcia, Simon Pezold, Bao Pham, Ashish Phophalia, Gemma Piella, G. N. Pillai, Marie Piraud, Maxim Pisov, Anmol Popli, Michael P. Pound, Reza Pourreza, Prateek Prasanna, Vesna Prkovska, Tony P. Pridmore, Santi Puch, Élodie Puybareau, Buyue Qian, Xu Qiao, Martin Rajchl, Swapnil Rane, Michael Rebsamen, Hongliang Ren, Xuhua Ren, Karthik Revanuru, Mina Rezaei, Oliver Rippel, Luis Carlos Rivera, Charlotte Robert, Bruce Rosen, Daniel Rueckert, Mohammed Safwan, Mostafa Salem, Joaquim Salvi, Irina Sanchez, Irina Sánchez, Heitor M. Santos, Emmett Sartor, Dawid Schellingerhout, Klaudius Scheufele, Matthew R. Scott, Artur A. Scussel, Sara Sedlar, Juan Pablo Serrano-Rubio, N. Jon Shah, Nameetha Shah, Mazhar Shaikh, B. Uma Shankar, Zeina Shboul, Haipeng Shen, Dinggang Shen, Linlin Shen, Haocheng Shen, Varun Shenoy, Feng Shi, Hyung Eun Shin, Hai Shu, Diana Sima, M Sinclair, Orjan Smedby, James M. Snyder, Mohammadreza Soltaninejad, Guidong Song, Mehul Soni, Jean Stawiaski, Shashank Subramanian, Li Sun, Roger Sun, Jiawei Sun, Kay Sun, Yu Sun, Guoxia Sun, Shuang Sun, Yannick R Suter, Laszlo Szilagyi, Sanjay Talbar, Dacheng Tao, Dacheng Tao, Zhongzhao Teng, Siddhesh Thakur, Meenakshi H Thakur, Sameer Tharakan, Pallavi Tiwari, Guillaume Tochon, Tuan Tran, Yuhsiang M. Tsai, Kuan-Lun Tseng, Tran Anh Tuan, Vadim Turlapov, Nicholas Tustison, Maria Vakalopoulou, Sergi Valverde, Rami Vanguri, Evgeny Vasiliev, Jonathan Ventura, Luis Vera, Tom Vercauteren, C. A. Verrastro, Lasitha Vidyaratne, Veronica Vilaplana, Ajeet Vivekanandan, Guotai Wang, Qian Wang, Chiatse J. Wang, Weichung Wang, Duo Wang, Ruixuan Wang, Yuanyuan Wang, Chunliang Wang, Guotai Wang, Ning Wen, Xin Wen, Leon Weninger, Wolfgang Wick, Shaocheng Wu, Qiang Wu, Yihong Wu, Yong Xia, Yanwu Xu, Xiaowen Xu, Peiyuan Xu, Tsai-Ling Yang, Xiaoping Yang, Hao-Yu Yang, Junlin Yang, Haojin Yang, Guang Yang, Hongdou Yao, Xujiong Ye, Changchang Yin, Brett Young-Moxon, Jinhua Yu, Xiangyu Yue, Songtao Zhang, Angela Zhang, Kun Zhang, Xuejie Zhang, Lichi Zhang, Xiaoyue Zhang, Yazhuo Zhang, Lei Zhang, Jianguo Zhang, Xiang Zhang, Tianhao Zhang, Sicheng Zhao, Yu Zhao, Xiaomei Zhao, Liang Zhao, Yefeng Zheng, Liming Zhong, Chenhong Zhou, Xiaobing Zhou, Fan Zhou, Hongtu Zhu, Jin Zhu, Ying Zhuge, Weiwei Zong, Jayashree Kalpathy-Cramer, Keyvan Farahani, Christos Davatzikos, Koen van Leemput, Bjoern Menze

arXiv:1811.02629v3cs.CVcs.AIcs.LGstat.ML

TL;DR

Glioma heterogeneity makes multimodal MRI segmentation difficult, while inconsistent datasets have limited comparison of machine-learning methods. This study reviews BraTS 2012–2018 across segmentation, progression assessment, and survival prediction, finding that fused segmentation labels consistently ranked first across evaluated tasks and metrics. The interpretation remains bounded by image-based annotations and uncertainty in sub-region delineation.

  • Problem

    Glioma heterogeneity and inconsistent private datasets make multimodal MRI segmentation difficult and hinder comparison of reported strategies.

  • Method

    The study reviews machine-learning methods and evaluation practices across seven BraTS challenge instances, covering segmentation, progression assessment, and overall-survival prediction.

  • Results

    Fused segmentation labels consistently ranked first across whole-tumor, tumor-core, and active-tumor tasks under Dice score and Hausdorff distance.

  • Takeaways & Limitations

    Ensembles of fused segmentation algorithms may be favorable for translating tumor-segmentation methods into clinical practice.

  • Takeaways & Limitations

    BraTS sub-regions are image-based rather than strict biological entities, and boundary delineation remains uncertain.

Abstract

from arXiv · show

Gliomas are the most common primary brain malignancies, with different degrees of aggressiveness, variable prognosis and various heterogeneous histologic sub-regions, i.e., peritumoral edematous/invaded tissue, necrotic core, active and non-enhancing core. This intrinsic heterogeneity is also portrayed in their radio-phenotype, as their sub-regions are depicted by varying intensity profiles disseminated across multi-parametric magnetic resonance imaging (mpMRI) scans, reflecting varying biological properties. Their heterogeneous shape, extent, and location are some of the factors that make these tumors difficult to resect, and in some cases inoperable. The amount of resected tumor is a factor also considered in longitudinal scans, when evaluating the apparent tumor for potential diagnosis of progression. Furthermore, there is mounting evidence that accurate segmentation of the various tumor sub-regions can offer the basis for quantitative image analysis towards prediction of patient overall survival. This study assesses the state-of-the-art machine learning (ML) methods used for brain tumor image analysis in mpMRI scans, during the last seven instances of the International Brain Tumor Segmentation (BraTS) challenge, i.e., 2012-2018. Specifically, we focus on i) evaluating segmentations of the various glioma sub-regions in pre-operative mpMRI scans, ii) assessing potential tumor progression by virtue of longitudinal growth of tumor sub-regions, beyond use of the RECIST/RANO criteria, and iii) predicting the overall survival from pre-operative mpMRI scans of patients that underwent gross total resection. Finally, we investigate the challenge of identifying the best ML algorithms for each of these tasks, considering that apart from being diverse on each instance of the challenge, the multi-institutional mpMRI BraTS dataset has also been a continuously evolving/growing dataset.

1 Introduction

BraTS provides a multi-institutional benchmark for challenging glioma segmentation in pre-operative mpMRI, extending toward progression assessment and overall-survival prediction. The study reviews machine-learning methods across BraTS 2012–2018 and examines how evolving datasets affect identifying effective algorithms.

  • Clinical motivation: BraTS evaluates automated segmentation of heterogeneous glioma sub-regions in multi-institutional pre-operative mpMRI scans.Glioma heterogeneity appears in histology, shape, and multimodal imaging intensity profiles.
  • Study scope: BraTS addresses limited comparability among prior segmentation studies by providing a shared benchmark and publicly available dataset.Earlier private datasets varied in modalities, tumor types, and disease state, complicating comparisons between reported strategies.
  • Clinical motivation: The challenge expanded beyond segmentation to secondary tasks involving longitudinal tumor growth and overall-survival prediction.Survival prediction uses segmentation labels and mpMRI-derived imaging or radiomic features analyzed with machine-learning algorithms.

2 Materials and Methods

Table 1 summarizes the original characteristics of the BraTS dataset.

  • Dataset characteristics: Table 1 summarizes the original characteristics of the BraTS dataset.
  • Dataset characteristics: The dataset characteristics are presented in tabular form.
  • Dataset characteristics: The table is identified as part of the BraTS dataset description.

2.1 BraTS Annotations and Structures

BraTS annotations define active tumor, tumor core, and whole tumor regions from manually segmented scans reviewed by experienced neuroradiologists. The nested structures distinguish enhancing tumor, core components, and surrounding edematous or invaded tissue.

  • Annotation sources: BraTS ground-truth scans were manually segmented by one to four raters using a common protocol and approved by experienced neuroradiologists.A single board-certified neuroradiologist further reviewed final labels for consistency and protocol compliance.
  • Tumor structures: The active tumor is the T1Gd-hyperintense enhancing region, excluding the necrotic center.
  • Tumor structures: The tumor core contains active tumor, necrotic fluid-filled tissue, and non-enhancing solid tumor.
  • Tumor structures: The whole tumor comprises the tumor core plus peritumoral edematous or invaded tissue visible as hyperintensity on T2-FLAIR.

2.2 Annotation Protocol

BraTS standardizes image-based tumor-region annotation across heterogeneous multi-center MRI data using nested delineation of whole tumor, core, and sub-regions. The protocol changed in 2017 by removing NET as a separate label and refining edema criteria, while annotation uncertainty remains substantial.

  • Protocol and imaging: BraTS combines scans from different centers, equipment, and imaging protocols with a standardized annotation protocol.Structural T1, T1Gd, T2, and T2-FLAIR volumes were co-registered to a common template and resampled to 1mm^3.
  • Delineation workflow: The recommended workflow delineates whole tumor first, then tumor core, and finally enhancing and non-enhancing or necrotic components.The protocol begins with abnormal T2 signal and proceeds inward toward core sub-regions.
  • 2012–2016 protocol: The four-region scheme defined AT, NET, NCR, and ED, with imaging-based criteria distinguishing enhancement, necrosis, non-enhancing tumor, and edema.AT is identified on T1Gd, while NET and ED rely substantially on T2-weighted or T2-FLAIR appearance.
  • 2012–2016 protocol: NET could be overestimated because image evidence is often limited, potentially producing institution-dependent ground truth and ranking bias.
  • 2017–present protocol: From BraTS 2017, NET was eliminated and combined with NCR, while certain contralateral and periventricular T2-FLAIR hyperintensities were excluded from ED.
  • 2017–present protocol: Low-grade gliomas often lack strong enhancement or edema, and boundaries between tumor and healthy tissue remain uncertain among expert annotators.

2.3 The BraTS Data Since its Inception

BraTS evolved from a small public benchmark into a larger, multi-institutional dataset with improved splits, longitudinal scans, validation data, and clinically reviewed annotations. Its framework supports segmentation and secondary clinical analyses, with standardized evaluation and survival-prediction data.

  • Dataset growth: BraTS 2012–2013 began with 35 training and 15 testing mpMRI scans, establishing a publicly available dataset and community benchmark.
  • Dataset growth: The 2014–2016 editions substantially increased the dataset, added longitudinal mpMRI scans, and incorporated improved development and evaluation data splits.
  • Latest BraTS data: 477 cases in 2017 and 542 cases in 2018 accompanied the introduction of validation data and further multi-institutional contributions.
  • Challenge tasks: BraTS expanded beyond tumor-subregion segmentation to secondary tasks using segmentation results for further clinical analysis, including progression assessment and overall survival prediction.
  • Latest BraTS data: Since 2017, the datasets included routine clinically acquired 3T pre-operative scans from 19 institutions, with expert-reviewed ground-truth labels and available overall survival data.
  • Evaluation framework: BraTS 2017–2018 rankings aggregated subject-level rankings across three regions and two metrics, with permutation testing used to assess differences between teams.
  • Evaluation framework: Overall survival prediction used matched training, validation, and testing distributions, with patients grouped into short-, intermediate-, and long-survivor categories.

3 Results

Across BraTS 2012–2018, fused segmentations performed strongly, while later challenge rankings showed gradual improvement without persistent dominance by a single approach. BraTS 2018 also revealed compartment-specific robustness and survival-prediction performance above random choice for the top teams.

  • BraTS 2012–2013: Label fusion out-performed all individual methods and matched inter-rater agreement in BraTS 2012–2013.The fused labels consistently ranked first across WT, TC, and AT segmentation under Dice and Hausdorff metrics.
  • BraTS 2017: 48 independent teams participated in BraTS 2017, with 47 submitting segmentation results and 16 submitting survival-prediction results.The first segmentation team was statistically better than the second, while the second was not statistically better than the third or fourth, producing a tie at rank three.
  • BraTS 2018: 61 teams submitted segmentation results and 26 submitted survival-prediction results in BraTS 2018.The patient-wise ranking showed gradual improvement, with no particular dominance across closely ranked approaches.
  • BraTS 2018: Median Dice for the top 54/63 BraTS 2018 teams ranged from 0.74–0.85 for AT, while average Dice ranged from 0.61–0.77.The skew reflected increasing outlier effects on average performance; TC was more robust than AT, and WT had median Dice 0.9 for most teams.
  • BraTS 2018: Median IQRs for 54/63 teams were 1.9 for AT, 4.0 for WT, and 5.4 for TC on the 95% Hausdorff distance.The 95% Hausdorff distance was used to characterize robustness of automated results.
  • BraTS 2018: Top-five BraTS 2018 survival approaches achieved accuracy around 0.6, compared with 0.15–0.55 for the remaining teams and 0.33 for random choice.The survival task was a three-class classification problem.

4 Discussion

The discussion finds different algorithmic strengths across BraTS tasks: deep learning dominates segmentation, whereas traditional machine learning performs better for overall-survival prediction on small datasets. It also emphasizes clinically informed evaluation, ensemble robustness, and future methods that accommodate heterogeneous clinical information.

  • Automated segmentation robustness: Fusion of labels from multiple automated segmentation methods showed greater robustness than expert inter-rater agreement in accuracy and consistency across subjects.Ensembling can reduce outliers and improve precision through consensus segmentation across models.
  • BraTS evaluation: BraTS case-wise ranking accounts for varying patient-case complexity, while statistical-significance testing supports comparisons across challenge instances over seven years.Together, these evaluation features make team ranking more clinically relevant and enable longitudinal analysis of algorithmic improvement.
  • Clinical relevance: Two clinically relevant sub-challenges were added to promote use of segmentation labels for clinical questions, clinical requirements, and potential decision support.These additions expanded BraTS beyond segmentation toward progression assessment and overall-survival prediction.
  • Overall-survival prediction: Overall-survival prediction in BraTS 2017–2018 exposed difficulties for deep learning on small training sets and favored traditional machine learning approaches.The discussion calls for larger training sets and possible synergies between deep learning and traditional machine learning.
  • Future directions: Future research should improve robustness to confounding effects and handle heterogeneous patient information, including radiogenomic and clinical data.The discussion identifies combining deep learning and traditional machine learning as a possible direction as training sets grow.
  • Cross-task algorithmic trends: Deep learning generally outperforms traditional machine learning for segmentation, particularly on Dice scores, but traditional methods remain stronger for overall-survival prediction.The contrast is associated with the larger training sets available for segmentation and smaller training sets for clinical-outcome prediction.

6 Supplementary Material

The supplementary material summarizes BraTS 2017 and 2018 segmentation results using Dice and Hausdorff measures across active-tumor, tumor-core, and whole-tumor compartments. Several Hausdorff figures additionally apply cutoff values for visualization.

  • BraTS 2018: BraTS 2018 Dice summaries cover active-tumor, tumor-core, and whole-tumor compartments.These results are shown in Figures 9–11.
  • BraTS 2018: BraTS 2018 Hausdorff summaries cover active-tumor, tumor-core, and whole-tumor compartments.Figures 13, 15, and 17 include cutoff values for visualization, while Figures 12, 14, and 16 do not.
  • BraTS 2017: BraTS 2017 Dice summaries cover active-tumor, tumor-core, and whole-tumor compartments.These results are shown in Figures 18–20.
  • BraTS 2017: BraTS 2017 Hausdorff summaries cover active-tumor, tumor-core, and whole-tumor compartments.Figures 22, 24, and 26 include cutoff values for visualization, while Figures 21, 23, and 25 do not.
Loading 1811.02629v3…