Source-linked AI summary
GLI-AL: A Multi-Modal Glioma MRI Label Resource with Unified Anatomy-Lesion Labels
Xingyu Xiang, Shuang Hao, Fan Wang, Jianhua Ma, Chunfeng Lian
TL;DR
BraTS-GLI annotations do not systematically represent coexisting WMH and healthy brain tissues in a unified joint-segmentation target. This paper releases 1,251 aligned eight-class labels, with 92.7% of released foreground voxels added outside the original expert lesion masks, enabling WMH-aware joint supervision and evaluation.
Problem
BraTS-GLI labels omit systematic coexisting-abnormality and healthy-tissue representation, limiting their use as unified targets for joint anatomy-lesion segmentation.
Method
The authors construct a controlled-access, labels-only resource of 1,251 aligned eight-class labels merging tumor and comorbid lesions with six healthy-tissue classes.
Results
92.7% of the released foreground, totaling 1.51 billion voxels, is added outside the original expert lesion masks across all 1,251 cases.
Takeaways & Limitations
The resource supports WMH-aware joint supervision and helps researchers distinguish data-quality and sample-size effects on sensitivity to coexisting lesions.
Takeaways & Limitations
Its automatically generated lesion and healthy-tissue labels may contain tool bias, missed detections, limited small-lesion sensitivity, and non-equivalence to manual anatomical annotations.
Abstract
from arXiv · showhide
Existing BraTS-GLI datasets provide a widely used benchmark for adult glioma MRI segmentation, but their task definition focuses on tumor subregions and does not systematically represent coexisting white matter hyperintensities (WMH). In joint segmentation settings, such unlabeled abnormalities introduce task-specific label noise by treating pathological regions as normal tissue. To address this limitation, we introduce BraTS-GLI Anatomy-Lesion, a controlled-access, labels-only derived resource built from the BraTS 2023-GLI training cohort. The resource provides 1,251 unified eight-class anatomy-lesion label sets aligned with the original four-modal MRI cases, including image-repair labels for 116 cases requiring repaired imaging inputs. The cohort is organized into a 394-case purified subset and an 857-case extended subset, with case-level metadata covering label source, image-repair requirements, quality-control status, access conditions, checksums, and release boundaries. Compared with the original BraTS-GLI annotations, the resource substantially expands foreground supervision by incorporating healthy brain tissues and previously unlabeled coexisting abnormalities within a unified label space. A validation study using MedNeXt and T1/FLAIR inputs suggests that WMH-aware supervision preserves healthy-tissue segmentation performance across both in-domain GLI and external WMH datasets, while improving sensitivity to coexisting lesions relative to noisy-control training. The resource is intended for scientific research and supports joint anatomy-lesion supervision, label-noise analysis, and reproducible evaluation. Data are available at https://www.synapse.org/Synapse:syn75210889/wiki/, and code is available at https://github.com/xyx200/brats-gli-anatomy-lesion-code. The data resource DOI is https://doi.org/10.7303/SYN75210889.
1. Background
BraTS datasets are central multi-parametric MRI benchmarks for glioma segmentation, but their tumor-focused annotations omit coexisting abnormalities and healthy tissue structures needed for joint supervision. WMH is common in BraTS-GLI, so treating unlabeled WMH as healthy tissue can introduce misleading supervision.
- Benchmark context: BraTS datasets provide multi-center, pre-operative, multi-parametric MRI with expert tumor-subregion annotations and serve as central public glioma-imaging benchmarks.BraTS 2023-GLI contains four-modal adult-glioma images.
- Annotation limitation: BraTS-GLI annotations cover necrotic tumor core, peritumoral edema, and enhancing tumor, but do not systematically annotate other coexisting brain abnormalities.This limitation becomes important when the same voxel space must represent tumors, healthy tissues, and coexisting abnormalities.
- WMH prevalence: 68.8% of 285 BraTS 2018 training cases contained at least 100 mm3 of WMH, showing that unlabeled WMH is common rather than isolated in glioma cohorts.The cited study added expert WMH annotations to the cases.
- Label noise: In joint segmentation training, unlabeled WMH regions are implicitly treated as healthy tissue, potentially producing a misleading supervision signal.This issue arises because the original task does not systematically represent coexisting WMH.
- Joint supervision gap: Original BraTS-GLI tumor-subregion labels cannot directly serve as joint targets because they omit both coexisting WMH and healthy brain structures from the same label space.The resource gap therefore concerns both abnormality coverage and anatomy coverage.
2. Summary
The resource releases 1,251 WMH-aware, unified anatomy-lesion label sets aligned with BraTS 2023-GLI MRI, enabling joint healthy-tissue and lesion segmentation research. It also provides transparent provenance, quality-control metadata, and repaired-image support for reproducible analyses.
- Research uses: The resource supports joint training and evaluation of healthy tissues and lesions within one label space, including sensitivity analyses by label source and case subset.Its intended applications include medical-image joint segmentation, brain-tumor MRI, label-noise analysis, and AI-ready datasets.
- Resource scope and format: 1,251 four-modal MRI cases receive derived NIfTI labels aligned with the upstream BraTS case-folder layout, without redistributing MRI images.The labels inherit SRI24 template space, 1 mm^3 resolution, 240×240×155 dimensions, and skull-stripped imaging settings.
- Image repair: 105,417 WMH repair voxels occur across repair-label cases, representing 0.89% of accompanying whole-tumor foreground volume and a median 448.5 repaired voxels per case.The accompanying code enables reconstruction of repaired image inputs after users obtain the upstream BraTS MRI.
- Label expansion: The unified labels add foreground supervision outside the original BraTS-GLI expert lesion mask by combining healthy-tissue structures with the final Lesion class.This comparison treats the original mask as pre-existing tumor-task foreground and the released unified label as the expanded representation.
3. Discussion
GLI-AL addresses unlabeled WMH-related label noise through auditable provenance, unified eight-class anatomy-lesion labels, and controlled cohort stratification. Its limitations include screening-source and tool bias, incomplete coverage of other abnormalities, restricted external modalities, and research-only use.
- Auditable provenance, QC status, image-repair entry points, checksums, and fixed release boundaries support case selection and reproducible reporting.The resource handles lesion-label noise at the data-source level while preserving traceability of release decisions.
- Unified eight-class labels enable joint modeling of healthy anatomical structures and lesions within one training target.They also support label-noise sensitivity analyses without switching between incompatible task labels.
- The 394-case purified subset combines expert-negative and model-screened samples rather than exclusively new manual WMH annotations.Extended-subset lesion constraints fuse DeepWMH and LST-AI, introducing potential tool bias, missed detections, and limited sensitivity to small lesions.
- The resource systematically covers WMH comorbid lesions, but microbleeds and lacunar infarcts remain incompletely addressed.Because external WMH validation uses only T1 and FLAIR, the study cannot establish T1ce and T2 value for every downstream task.
- GLI-AL is restricted to scientific research and must not guide clinical diagnosis, treatment decisions, or individual risk assessment.Users should report data tier and label source, disclose automatically generated-label limitations, avoid re-identification, and contact the resource team about errors.
4. Resource Availability
GLI-AL is a controlled-access, labels-only derived resource for scientific research on joint healthy-tissue–lesion supervision and label-noise analysis. Access requires upstream BraTS 2023 data approval followed by a separate GLI-AL application through Synapse.
- Resource scope: GLI-AL provides BraTS-GLI-derived labels and case-level metadata for joint healthy-tissue–lesion supervision involving coexisting WMH and healthy tissue structures.The original tumor-subregion labels cannot directly represent this relationship.
- Repositories and identifiers: The Synapse repository is https://www.synapse.org/Synapse:syn75210889/wiki/, the data resource DOI is https://doi.org/10.7303/SYN75210889, and code is available at https://github.com/xyx200/brats-gli-anatomy-lesion-code.The code repository includes processing scripts and MedNeXt/nnUNet unified-label evaluation adaptation files, but does not redistribute BraTS imaging data or listed third-party resources.
- Access requirements: Users must obtain BraTS 2023 access through Synapse project syn51156910, then apply separately for GLI-AL through the syn75210889 wiki and accept both terms.The derived resource does not replace or bypass the upstream access process.
- Intended use: The resource is intended only for scientific research, including joint segmentation, healthy-tissue and lesion modeling, label-noise analysis, provenance-based training, and external cross-resource validation.These use cases depend on joint labels for healthy tissues and lesions.
- Licensing and conditions: The resource is released under CC BY-NC 4.0 and remains subject to applicable BraTS 2023 post-Challenge Terms and Conditions.Processing scripts use the MIT License, while adapted MedNeXt/nnUNet evaluation files use Apache License, Version 2.0.
- Ethics and data governance: Because GLI-AL contains controlled-access derived labels without new recruitment, scanning, clinical-variable collection, or identifying information, this organization work requires no new ethics approval or exemption process.Upstream BraTS data governance handles consent, ethics approval, de-identification, authorization, and final release decisions.
5. Methods
The resource is constructed from 1,251 BraTS 2023-GLI cases as labels-only unified anatomy-lesion data, with purified and extended subsets preserving case stratification and repair status. Its workflow combines WMH screening or completion, healthy-tissue probability maps, quality control, and lesion-constrained fusion.
- Data source and label construction: 1,251 BraTS 2023-GLI cases yield 1,251 unified segmentation labels and 116 image-repair labels in a labels-only release.Each case corresponds to upstream T1, T1ce, T2, and T2-FLAIR MRI, which are not redistributed; unified labels merge tumor and comorbid lesions into Lesion and add six healthy-tissue classes.
- Subset construction: 394 cases form the purified subset, while the remaining 857 cases form the extended subset with original imaging status retained.The purified subset combines 342 model-negative cases with 52 expert-negative cases; 116 cases use repaired images without being counted as duplicates.
- Extended subset label completion: The extended subset intersects DeepWMH and LST-AI candidate masks, then unites the intersection with existing BraTS lesion masks to form Lesion.This procedure targets coexisting abnormalities while retaining the original four-modal images for the 857 non-purified cases.
- Unified anatomy-lesion labels: TumorSynth probability maps from four MRI modalities are mapped into seven foreground classes, including GM, BG, WM, Lesion, Ven, Cer, and BS.Automatic IQR-based outlier detection and manual review remove low-quality modality maps that cause missing foreground before fusion.
- Unified anatomy-lesion labels: Lesion hard constraints correct modality probabilities before renormalization, uncertainty-based weighting, and final fusion into integer labels.The resulting labels fix the unified lesion constraint while preserving high-confidence healthy-anatomical predictions from valid modalities.
6. Validation
Validation combines multi-level quality control with MedNeXt experiments assessing whether purified, WMH-aware supervision preserves healthy-tissue segmentation while improving interpretation of coexisting-lesion sensitivity. External results support preserved healthy-anatomy performance, but gray-matter discrepancies between label-generation protocols limit interpretation of overall out-of-domain metrics.
- Quality control: Quality control spans expert-negative selection, high-confidence model filtering, intersected lesion constraints, modality outlier removal, and case- and file-level integrity checks.Purified expert-negative cases have strictly zero WMH volume; other purified cases were selected using five-fold MedNeXt predictions, while extended-subset lesion masks intersect DeepWMH and LST-AI outputs.
- Experimental setup: MedNeXt-M experiments used five-fold cross-validation, default architecture settings, AdamW optimization, 1000 epochs, and batch size 2.Evaluation reports case-level mean and standard deviation for DSC and HD95, with explicit handling of empty masks and failed or threshold-exceeding distance computations.
- Validation datasets: 251 GLI cases supported in-domain testing, while external WMH validation used 170 MICCAI 2017 images processed with BraTS-aligned registration, skull stripping, and label-generation procedures.The external set was not part of the released GLI resource, and its healthy-tissue labels were mapped to the same seven foreground classes.
- Validation findings: Healthy-tissue performance remained similar in external zero-shot WMH evaluation, whereas Baseline achieved higher lesion DSC and lower HD95 than Joint-Baseline.This comparison motivates the purified subset as a WMH-aware cleaned reference for separating data-quality effects from sample-size effects on coexisting-lesion sensitivity.
- Limitations: Out-of-domain gray-matter performance was affected by inconsistent boundary definitions between TumorSynth labels used for GLI training and WMH-SynthSeg labels used for WMH testing.Therefore, the relatively low overall out-of-domain GM metric should not be interpreted as a model failure.