Source-linked AI summary
Infant FreeSurfer: An automated segmentation and surface extraction pipeline for T1-weighted neuroimaging data of infants 0-2 years
Lilla Zöllei, Juan Eugenio Iglesias, Yangming Ou, P. Ellen Grant, Bruce Fischl
TL;DR
Infant brain morphometry lacks broadly applicable automated tools because infant MRI involves changing anatomy, motion, and developmental contrast. This paper presents a single-channel T1-weighted pipeline for 0-2-year-olds, combining age-adaptive multi-atlas segmentation with surface extraction and reporting strong overlap against expert brain masks alongside limitations from excluding T2-weighted MRI.
Problem
Automated infant brain morphometry tools lag behind adult methods despite the need to quantify cortical abnormalities associated with neuropsychiatric, neurologic, and developmental disorders.
Method
The paper presents a single-channel T1-weighted pipeline for 0-2-year-olds that combines age-adaptive multi-atlas segmentation with cortical surface extraction.
Results
>90% overlap with expert-delineated brain masks was achieved using the Dice overlap coefficient.
Takeaways & Limitations
The unified procedure supports automated cortical and subcortical segmentation, volumetric outputs, and surface models across the full 0-2-year age range.
Takeaways & Limitations
The pipeline does not accommodate T2-weighted MRI, limiting direct performance comparisons with existing tools, particularly for newborns.
Abstract
from arXiv · showhide
The development of automated tools for brain morphometric analysis in infants has lagged significantly behind analogous tools for adults. This gap reflects the greater challenges in this domain due to: 1) a smaller-scaled region of interest, 2) increased motion corruption, 3) regional changes in geometry due to heterochronous growth, and 4) regional variations in contrast properties corresponding to ongoing myelination and other maturation processes. Nevertheless, there is a great need for automated image-processing tools to quantify differences between infant groups and other individuals, because aberrant cortical morphologic measurements (including volume, thickness, surface area, and curvature) have been associated with neuropsychiatric, neurologic, and developmental disorders in children. In this paper we present an automated segmentation and surface extraction pipeline designed to accommodate clinical MRI studies of infant brains in a population 0-2 year-olds. The algorithm relies on a single channel of T1-weighted MR images to achieve automated segmentation of cortical and subcortical brain areas, producing volumes of subcortical structures and surface models of the cerebral cortex. We evaluated the algorithm both qualitatively and quantitatively using manually labeled datasets, relevant comparator software solutions cited in the literature, and expert evaluations. The computational tools and atlases described in this paper will be distributed to the research community as part of the FreeSurfer image analysis package.
1.1 Atlases
Infant segmentation tools rely on heterogeneous training datasets, making direct comparison difficult. Dataset differences in age, modality, labels, subject characteristics, and encoded information can also introduce bias.
- Training datasets support automated segmentation by encoding regions of interest, intensity information, intensity distributions, or tissue probability maps.
- Direct comparison of infant training datasets and their segmentation performance is challenging because represented ages vary widely.
- Published datasets differ in imaging modality, subject numbers, atlas representation, subject characteristics, label origins, and contained information.
- Many neonatal datasets use prematurely born subjects scanned at term-equivalent ages, which may introduce bias when prior information guides applications.
1.2 Segmentation tools
Existing infant segmentation tools vary substantially in age coverage, imaging inputs, labels, atlas guidance, and surface outputs. Most focus on newborns or discrete ages, while relatively few support single-channel T1-weighted input or cortical surface reconstruction.
- Most postnatal infant segmentation tools target newborns, preterm subjects, discrete ages within the first year, or 2-year-olds rather than the full 0-2-year range.
- Most newborn tools use T2-weighted images because immature brains provide higher cerebral-tissue contrast, while only one known pipeline accommodates a single T1-weighted volume.
- Segmentation labels range from tissue classes such as cortical gray matter, white matter, and cerebrospinal fluid to brainstem, cerebellar, cortical, and subcortical regions.
- Some frameworks use age-specific infant atlases, whereas others extrapolate labels from adult or older-pediatric atlases.
- Only a few infant image-processing packages generate cortical surface models, including adaptations of adult or fetal pipelines such as NEOCIVET.
1.4 Contribution
The contribution is a unified FreeSurfer-based pipeline for clinical T1-weighted infant MRI across ages 0-2 years. It combines infant-specific preprocessing, training data, skull stripping, and segmentation to produce volumetric and surface outputs.
- The pipeline accommodates clinical T1-weighted infant MRI from a population aged 0-2 years and produces cortical and subcortical segmentations, volumes, and surfaces.
- Age-adaptive selection of training subsets enables one procedure to operate across the full target age range without sacrificing accuracy.
- The multi-stage process follows FreeSurfer’s adult reconstruction pipeline while introducing infant-specific image-processing components.
- Infant skull stripping is difficult because of developmental variability, limited cortex-skull separation, lower tissue contrast, and poor fit of adult-oriented tools.
- The double-consensus skullstripping approach combines multiple pediatric datasets with multiple skullstrippers to identify brain regions.
2.2 Volumetric segmentation
Volumetric segmentation uses Bayesian multi-atlas label fusion with infant T1-weighted training data and age- or similarity-based atlas selection. The method accommodates contrast variation and incomplete atlas information across development.
- The framework uses a multi-atlas label-fusion method that applies ground-truth labels from manually labeled training data to new infant brain images.
- The model includes a smooth, non-negative multiplicative bias field represented as the exponential of a linear combination of smooth basis functions.
- Segmentation is formulated as Bayesian inference over the most likely labeling given the image and registered atlas segmentations, using point estimates for model parameters.
- A prior excludes atlas regions lacking sufficient gray-white contrast, allowing a larger and non-uniform training dataset while maximizing usable labels.
- The algorithm uses 26 manually segmented T1-weighted images and selects complete or age- or similarity-based training subsets for each test subject.
- Atlas volumes are spatially normalized to test images with DRAMMS, selected for robustness to noise, field-of-view, appearance, anatomical, and age differences.
2.3 Surface Extraction
Surface extraction tessellates the gray matter–white matter boundary, corrects topology, and deforms surfaces along intensity gradients to place cortical boundaries. Unlike adult processing, infant surface fitting performs better with a relatively heavier weighting on the v
- The pipeline tessellates the gray matter–white matter boundary and applies automated topology correction.
- Surface deformation follows intensity gradients to place the gray matter–white matter and gray matter–cerebrospinal-fluid borders at tissue transitions.
- Infant surface fitting performs more optimally than adult processing when a relatively heavier weight is assigned to the v
2.4 Experiments
The experiments evaluated infant segmentation using manually labeled datasets, independent newborn scans, comparator tools, and qualitative and quantitative measures. The evaluation covered skullstripping, volumetric segmentation, training-set selection, and comparisons with publicly available neonatal pipelines.
- The BCH_0-2yr dataset comprised 26 clinically scanned infants ranging from newborns to 2 years and supported leave-one-out accuracy evaluation.
- An independent dataset contained 17 healthy full-term neonates for comparisons using T1-weighted and T2-weighted imaging.
- The tool achieved >90% Dice overlap with expert-delineated brain masks, comparable to reported adult skullstripping results of 94-96%.
- Comparisons included iBEAT, MANTIS, and the developing Human Connectome Project pipeline using newborn datasets and mutually available labels.
- Manual segmentation variability was below 60% for some structures, including the left amygdala and left and right accumbens, establishing an upper bound for automated performance.
3.1 Qualitative segmentation evaluation
Qualitative evaluations across infants from newborn to 18 months showed high correspondence between automated and manual segmentations despite age-dependent contrast changes. The tool also recovered white matter labels acceptably in several cases where manual gray matter–white matter boundaries were uncertain.
- Across five subjects aged newborn to 18 months, automated and manual segmentations showed high correspondence in gray matter–white matter boundaries and subcortical regions.
- The examples included newborn reverse intensity contrast and older infant images with more adult-like contrast alongside residual unmyelinated areas.
- In subjects aged 2, 3, 5, 6, and 9 months, automated white matter labels were recovered with acceptable accuracy despite uncertain manually defined gray matter–white matter boundaries.
3.2 Quantitative segmentation evaluation
Quantitative evaluation showed that age-matched training subsets generally outperformed the complete training set, with performance varying across age and region. Generalized Dice values reached 0.83 on average and nearly 0.94 at maximum, while thalamus and pons achieved the highest regional scores.
- Age-selected training subsets yielded better overall segmentation performance than the complete training set across the 26 evaluated subjects.
- Generalized Dice values reached 0.83 for age-group averages and nearly 0.94 for age-group maxima.
- The highest measurements occurred in the middle of the studied age range, whereas newborns had the lowest measurements.
- Training-set neighborhood size 5 clustered most often among the best-performing choices across the five age groups.
- Left and right thalamus and pons consistently achieved the highest average Dice scores across evaluated regions.
3.3 Qualitative and quantitative comparisons with other tools
The pipeline was compared with other infant segmentation tools despite the lack of a fair ground-truth comparison, showing generally good but variable correspondence across shared labels and datasets.
- Direct comparison was limited because the tools required different inputs, defined different labels, and lacked manual ground truth for the testing datasets.
- The pipeline and MANTIS were compared on seven commonly identified regions in 17 newborns using corresponding T1w and T2w images.
- The pipeline and iBEAT showed lower overall correspondence, around 0.6, across six commonly defined subcortical labels in 12 newborns.
- On 40 dHCP subjects, the dHCP cortical segmentations seemed slightly more accurate, partly because of differences in T1w and T2w input-image quality.
- Mean Dice overlaps were computed per label, distinguishing 13 total labels from seven more detailed tissue labels overlapping with the pipeline definitions.
3.4 Surface extraction
Surface extraction was evaluated qualitatively, by expert scoring, against dHCP surfaces, and against manually placed control points, with evidence of moderate quality and dataset-dependent differences.
- Because ground-truth surface reconstructions were unavailable, five subjects spanning newborn to 18 months were first assessed qualitatively.
- Against dHCP surfaces, mean absolute distance, sulcal-depth, cortical-thickness, and curvature differences averaged 1.17mm, -0.5, 0.9, and -0.09, respectively.
- On the studied dHCP dataset, the minimal processing pipeline outperformed the proposed surfaces significantly at the 5% significance level using manually drawn control points.
4.1 General comments
The pipeline functioned across the target age range with high reported accuracy, while age-matched training selection and image sharpness were associated with performance; comparisons remain constrained by methodological differences.
- The pipeline was designed for 0–2-year-old clinical T1-weighted MRI and produced automated cortical and subcortical segmentations plus surface extractions from a single channel.
- Overall functioning was consistent across the target age range, with high accuracy measured using Dice and Generalized Dice coefficients.
- Generalized Dice scores were highest for the 2–8-month subset and lowest for newborns; selecting a few atlases close in age to test subjects was optimal in most cases.
- Comparisons with three public algorithms were generally good but varied, with perhaps the best correspondence to dHCP results.
- Higher Tenengrad sharpness was associated with higher Generalized Dice scores and older age at scan.
- The Tenengrad metric for dHCP newborn T1-weighted images was in the same range as that of the newborn training subjects.
- The pipeline and its training dataset were slated for distribution in source and binary formats through FreeSurfer under a modified MIT-style license.
4.2 Limitations
The pipeline has important scope constraints: it does not yet accommodate T2-weighted MRI and requires 1 mm isotropic input resolution. Planned extensions target newborn performance, broader comparisons, and more complete training data.
- Imaging scope: The pipeline does not accommodate T2-weighted MRI, limiting direct performance comparisons with existing tools, particularly for newborns.T2-weighted images typically provide higher contrast-to-noise ratios for infants up to 6 months of age.
- Resolution constraint: All input images are resampled to 1 mm isotropic resolution to match the training datasets.The authors plan to remove this constraint by obtaining higher-resolution manually segmented datasets for training and validation.
- Training data: The current training dataset lacks some gray-matter/white-matter boundary descriptions, which may limit accuracy and consistency.The authors propose increasing the number of training subjects and collecting full boundary segmentations across the entire age range.
- Future extensions: Planned T2-weighted training extensions may improve newborn segmentation accuracy and enable direct comparisons with other tools and multisite datasets.The proposed extensions would use shared datasets such as dHCP and MICCAI challenge datasets.