Source-linked AI summary

Representation learning of human cortical folding to reveal long lasting neurodevelopmental signatures

Julien Laval, Robin Guiavarch, Antoine Dufournet, Racim Menasria, Barthélémy Drabczuk, Cristobal Mendoza, Saeb Tounsi, Chikh Abdelghani Baroud, Merieme Bourenane, Vanessa Troiani, William Snyder, Marisa A Patti, Mylène Moyal, Marion Plaze, Arnaud Cachia, Federica Santacroce, Giorgia Committeri, Claire Cury, Kevin De Matos, Olivier Colliot, Zhong Yi Sun, Clara Fischer, Vincent Frouin, Pietro Gori, Denis Rivière, Joël Chavas, Jean-François Mangin

arXiv:2609.05438v1q-bio.QMcs.CVcs.LG

TL;DR

Current neuroimaging foundation models may underrepresent cortical folding, despite its early emergence, lifelong stability, and relevance to neurodevelopment. Champollion learns local, interpretable folding representations from structural MRI and consistently captures folding patterns better than tested alternatives, while revealing genetic and clinically relevant localized signatures.

  • Problem

    Cortical folding is an early and stable neurodevelopmental signal, but whether generic neuroimaging representations capture its variability remains unclear.

  • Method

    Champollion uses self-supervised learning with ROI-specific lightweight convolutional encoders to learn local, interpretable cortical-folding representations from structural MRI.

  • Results

    Champollion consistently captured folding patterns better than tested medical and general-purpose foundation models across cortical regions and external datasets.

  • Takeaways & Limitations

    The representations expose genetically grounded folding information and localized signatures associated with incomplete hippocampal inversion, prematurity, and maternal smoking.

  • Takeaways & Limitations

    Multivariate genetic comparisons were conducted on selected ROIs, with significance assessment explicitly accounting for the number of phenotypes.

Abstract

from arXiv · show

The human brain folds in utero, primarily during late gestation. Shortly after birth, cortical folding patterns are established and remain stable thereafter, making them promising early neurodevelopmental markers. Yet it is unclear whether representations given by current neuroimaging foundation models capture cortical folding variability. Here, we introduce Champollion, a self-supervised learning framework that learns interpretable local representations of cortical folding from structural MRI. Optimized on representative folding-related tasks, Champollion accurately captures known folding patterns across cortical regions and external datasets. In a comprehensive benchmark, it consistently outperforms neuroimaging and general-purpose foundation models. Furthermore, Champollion reveals richer genetic associations than conventional morphometric descriptors and identifies localized folding signatures associated with incomplete hippocampal inversion, prematurity, and maternal smoking. These results establish cortical folding as a rich and largely untapped source of neurodevelopmental information and illustrate how pre-processing and architectural inductive biases can recover biologically meaningful signals overlooked by current generalist foundation models.

2 Introduction

Cortical folding emerges early, remains stable after birth, and varies across individuals in genetically patterned ways. These properties motivate testing whether self-supervised neuroimaging representations capture folding variability and introducing Champollion to learn interpretable local representations.

  • Motivation: Individual differences in sulcal shape are structured rather than random, with developmental gene-expression gradients matching future sulci and gyri.
  • Motivation: Cortical folding reaches near-adult levels shortly after birth, while fold topology and shape remain stable thereafter.This stability contrasts with scalar descriptors such as sulcal depth and gyrification index, which continue evolving during development and aging.
  • Motivation: Folding patterns have been associated with behavioral traits and clinical phenotypes, including inhibitory control and psychiatric disorders.Reported examples include anterior cingulate and inferior frontal asymmetries and differing orbitofrontal sulcal distributions in patients and controls.
  • Research gap: Self-supervised learning extracts transferable imaging features without predefined anatomical measurements, but it remains unclear whether generic foundation-model representations capture cortical folding.
  • Approach: Champollion addresses this gap by learning low-dimensional, interpretable folding representations from regional MRI crops.The framework is trained on approximately 42,000 UK Biobank samples after preprocessing MRIs to extract cortical folds.
  • Contributions: Across folding-pattern and downstream analyses, Champollion outperformed handcrafted descriptors, prior self-supervised methods, and vision foundation models.The analyses also examined genetic associations and folding patterns related to incomplete hippocampal inversion, prematurity, and maternal smoking exposure.

3 Results

Champollion learns local, interpretable cortical-folding representations from MRI and consistently outperforms existing alternatives across folding-related tasks. Its representations preserve biologically meaningful, regionally distinct information, including genetic associations, sex-related patterns, twin similarity, and localized neurodevelopmental signatures.

  • 3.1: Champollion extracts aligned cortical skeleton tiles from 56 sulcus-centered ROIs and encodes each ROI with a dedicated lightweight model.A small annotated dataset defines anatomical masks enabling label-free regional extraction, while dedicated decoders support interpretation.
  • 3.1: 20.9% higher R in central sulcus regression resulted from jointly using positional preservation, mild positional augmentation, and spatially consistent masking.Preserving sulcal fundi added a 7.6% ROC-AUC increase on the orbitofrontal cortex task compared with standard cutout.
  • 3.2: Decoder reconstructions recovered central sulcal structures across optimized and non-optimized ROIs, while latent traversals represented sulcal translation and curvature.For the right central sulcus, the first principal component primarily encoded rostro-caudal translation and the second encoded curvature variation.
  • 3.2: Champollion consistently outperformed all zero-shot baselines across tasks, whereas foundation models largely overlooked cortical folding information.A VBM preprocessing pipeline that introduced nonlinear deformations completely erased folding information; β-VAE performance was task-dependent.
  • 3.2: 43 independent genetic loci versus 10 for classical morphometry were identified in left SFinf–BROCA–SPeCinf, while right ScCal–SLi yielded 122 versus 9 loci.Champollion encompassed and exceeded classical morphometry’s genetic associations in both examined ROIs, despite weak-to-moderate correlations in one region.
  • 3.3: Regional latent spaces showed less than 12% pairwise similarity across non-overlapping ROIs, supporting largely independent local folding variability.This regional independence supports learning folding representations locally rather than relying on a single globally shared representation.
  • 3.4: 97% AUC sex prediction transferred to HCP with less than 1% AUC loss, while concatenated regional representations placed 80% of monozygotic twins as nearest neighbors.The first five nearest neighbors contained 92% of twin pairs.
  • 3.5: Localized predictive signals included the central-sulcus knob pattern, smoother Isomap-related variation at R2 = 60%, and signatures associated with incomplete hippocampal inversion, prematurity, and maternal smoking.Across these applications, predictive information remained spatially confined and latent traversals produced coherent shape variations, including a smoking-related signal recovered across age-separated cohorts.

4 Discussion

Champollion captures localized, interpretable cortical-folding information that is underrepresented in general neuroimaging representations. Its associations with genetic similarity and developmental or exposure-related phenotypes support cortical folding as a stable source of neurodevelopmental information.

  • Implications: Champollion consistently captured folding patterns better than tested medical and general-purpose foundation models across cortical regions and external datasets.The discussion attributes this gap partly to the domain mismatch between foundation-model training data and localized cortical-folding information.
  • Implications: Regional Champollion representations revealed largely independent folding patterns, supporting local representations for whole-brain variability.The framework combines native interpretability with localized encoding of folding information.
  • Biological relevance: Champollion representations showed stronger genetic associations than classical morphometric descriptors and matched monozygotic twins through latent proximity.This supports the interpretation that the learned features reflect heritable aspects of cortical folding.
  • Biological relevance: Latent traversals identified localized folding signatures associated with incomplete hippocampal inversion, prematurity, and maternal smoking.Reported changes included collateral-sulcus torque reduction, superior-temporal-sulcus flattening, and occipital-fold reorientation.
  • Robustness: The maternal-smoking pattern replicated across UK Biobank and ABCD cohorts, suggesting a stable long-term signature and robustness to site and population variability.The cohorts differed substantially in age and demographic context.
  • Scope: The exploratory associations extend beyond conventional morphometric measures, but the authors relate them to prior literature without presenting them as definitive biological associations.For prematurity, the discussion connects the temporal-lobe finding to prior work while noting that the result concerns subtler folding alterations than sulcal depth.

5 Methods

The methods pipeline extracts and regionally represents cortical folds from structural MRI while addressing large images, limited samples, inter-individual variability, and acquisition biases. It uses skeletonized folds, affine alignment, anatomically corresponding ROIs, and local self-supervised encoders.

  • Overview: The study learns self-supervised representations of cortical-folding variability from structural MRI, covering the pipeline, optimization, and downstream analyses.The methods are organized around preprocessing, model optimization, and downstream evaluation.
  • Preprocessing: Cortical-folding preprocessing addresses large semantically dense volumes, limited samples, high inter-individual variability, and age or site acquisition biases.Standard voxel- and surface-based registration can reduce variability but inevitably distort cortical folds.
  • Preprocessing: Folds are extracted as one-voxel-wide skeletons tracing the cortical surface medial axis, removing fold-opening information that changes with aging.The BrainVisa Morphologist pipeline produces a skeletonized negative cast of the white-matter surface.
  • Preprocessing: Skeletons are affinely aligned to MNI space and resampled at 2 mm isotropic resolution while remaining one voxel wide.Affine alignment standardizes brain orientation and size before regional extraction.
  • Regional tiling: The aligned cortical skeleton is tiled into anatomically corresponding ROIs so local representations can focus on one or a few neighboring sulci.The tiling relies on spatial normalization rather than assuming that folds are identical across individuals.
  • Regional tiling: The pipeline defines 56 overlapping ROIs and treats left- and right-hemisphere regions separately because corresponding sulcal regions need not be bilaterally symmetric.The ROIs cover the whole brain and may contain flexible combinations of neighboring sulci.
  • Self-supervised learning: Champollion uses joint-embedding self-supervision with augmentations that encode tolerance to small positional variation and attention to sulcal fundi.These choices inject folding-specific prior knowledge into representation learning.

5.3 Data augmentations for SSL

The augmentation policy adapts joint-embedding self-supervision to binary, spatially aligned cortical skeletons. It preserves meaningful spatial context while regularizing residual misalignment and emphasizing fundus voxels.

  • Design rationale: Augmentation choices encode prior knowledge about the data and can strongly influence which features are represented.This motivates designing policies specifically for cortical skeletons rather than importing natural-image defaults unchanged.
  • Design rationale: Natural-image augmentations such as blur, color jitter, grayscale conversion, and flipping are poorly matched to binary, textureless, aligned cortical skeletons.Cropping is the main natural-image analogue retained, but it must respect spatial normalization.
  • Spatial design: The backbone preserves absolute local position because fold features have different meanings in anterior and posterior locations after global alignment.This positional sensitivity motivates replacing translation-equivariant global pooling with a flatten-and-dense design.
  • Spatial design: Mild translations and small rotations regularize residual local misalignment while Barlow Twins enforces invariance to these perturbations.The augmentations complement a backbone that remains sensitive to spatial location.
  • Masking: Cut-out masks a random bounding box without recentering or rescaling, preserving spatial positions; cut-in complements it by addressing variable ROI boundaries.These operations replace standard crop-resize augmentation.
  • Fundus focus: Cut-in and cut-out selectively preserve fundus voxels to encourage attention to semantically meaningful traces of folding development.Fundus preservation is sampled at the simple-surface level with a tunable probability.
  • Complete policy: The final policy applies rotations up to 18° per axis and translations of ±1 voxel, followed by cut-in or cut-out masking with probability 0.8.Cut-in and cut-out are exclusive and use proportionally sized bounding boxes.

5.4 Backbone

Champollion uses a compact custom 3D CNN whose representation preserves spatial location through flattening rather than global pooling. A dense projection produces a 32-dimensional latent space, with a separate MLP used for self-supervised loss computation.

  • Architecture: The custom backbone is a 12-layer 3D CNN with approximately 2 million convolutional parameters and three downsampling stages.It uses 7 × 7 × 7 kernels initially, 3 × 3 × 3 kernels thereafter, and progressively increases filters from 1 to 128.
  • Representation: Feature maps are flattened and projected into a 32-dimensional latent space, binding local features to their spatial locations.Flattening replaces global pooling and removes translation equivariance.
  • Optimization: A three-layer nonlinear MLP expands the representation to 128 dimensions for Barlow Twins loss computation.The 32-dimensional latent representation remains the compact feature space before loss projection.
  • Training: Training uses batch normalization, LeakyReLU activations, 5% dropout, a 4 ∗10^-4 learning rate, and 80 epochs.These are implementation settings applied after the convolutional layers and during optimization.
  • Regional modeling: Each ROI has its own backbone, yielding approximately 100 million convolutional parameters across the 56 cortical regions.The regional models collectively cover the cortex while retaining region-specific processing.

5.5 Optimization

Champollion is optimized through folding-specific linear-probing tasks spanning multiple cortical regions and variability types. The design uses shared optimization choices, selected hemisphere-specific settings, and stratified evaluation procedures.

  • Optimization strategy: Linear probing evaluates representations on folding-specific classification and regression tasks rather than global traits such as sex, age, or pathology.This focuses optimization on cortical folding patterns instead of residual information related to broader biological attributes.
  • Optimization strategy: Region-specific hyperparameter optimization was avoided because sparse annotations would make it impractical and could encourage local overfitting.The shared strategy preserves flexibility to define new ROIs without retraining optimization separately for each region.
  • Task design: Four downstream tasks span sulcus presence, interruption types, and continuous shape regression across distinct cortical regions and external datasets.The tasks were selected to cover topological and non-topological variability while encouraging robustness to scanner, resolution, and population shifts.
  • Optimization strategy: A single hemisphere was optimized for each task to reduce computation, using the left orbitofrontal cortex and right intraparietal sulcus for their respective task settings.The orbitofrontal choice reflected more balanced class distributions, whereas the intraparietal choice followed preliminary evidence of more challenging predictions on the right.
  • Evaluation procedure: Held-out testing used a 20% subject split followed by 5-fold stratified cross-validation accounting for labels, sex, age, acquisition site, and sibling grouping in HCP.Classification used LogisticRegression and regression used ElasticNet, with their regularization settings tuned over predefined grids.

5.6 Evaluation of the baselines

Champollion was compared with handcrafted, self-supervised, and general-purpose foundation-model baselines on the same cortical-folding tasks. The benchmark varied model families, inputs, preprocessing, and representation extraction, then adapted the strongest zero-shot foundation model to cortical skeletons.

  • Evaluation design: The benchmark used identical data splits and linear models for Champollion and every baseline.Ridge regression and RidgeClassifier were used for very large foundation-model latent spaces to reduce computation time.
  • Baseline families: Baselines included a cortical-folding β-VAE retrained on UK Biobank data and foundation models for natural images, point clouds, 3D medical representation, and segmentation.The compared foundation models included DINOv3, Point-M2AE, 3DINO-ViT, BrainSegFounder, SAM-Med3D, and VISTA3D.
  • Evaluation design: More than 1,500 configurations tested ROI-centered MRI crops and cortical skeletons with multiple preprocessing and representation-extraction settings.These choices were intended to make zero-shot comparisons fair across models with different input domains and capabilities.
  • Model adaptation: 3DINO-ViT, the best zero-shot foundation model, was adapted to cortical skeletons using continual LoRA pretraining, sparsity-aware masking, and Champollion-style positional augmentations.The adaptation targeted the binary and sparse structure of cortical skeleton inputs.

5.7 Multivariate genetics associations

The study compares Champollion’s 32-dimensional regional representations with conventional sulcal morphometric phenotypes in multivariate genetic analyses. The comparison standardizes regional definitions and adjusts phenotypes for demographic, imaging, anatomical, and genetic covariates before applying MOSTest.

  • Phenotype construction: Each ROI was represented either by Champollion’s 32-dimensional latent space or by tabular sulcal morphometric measures.The morphometric measures comprised geodesic length, surface area, mean depth, and maximum depth.
  • Phenotype construction: Sulcal morphometric features were concatenated according to the same regional definitions used for Champollion.Each region therefore combined measures from all sulci composing that ROI.
  • Preprocessing: Both latent and tabular phenotypes were residualized for sex, age, age^2, age-by-sex interactions, imaging center, total intracranial volume, and the first 20 genetic principal components.These adjustments were specified for the UK Biobank cohort.
  • Genetic association analysis: MOSTest combined summary statistics from univariate GWAS across phenotype dimensions while accounting for their correlation structure.The test was applied separately to Champollion representations and tabular morphometric phenotypes.
  • ROI selection: Two ROIs were selected to balance subject count and phenotype dimensionality while contrasting weak-to-moderate versus strong correlations between morphometry and Champollion’s latent space.This design addressed dimensional imbalance between the 32-dimensional representation and classical morphometric features.

5.8 Exploratory analysis

The exploratory analysis tests whether regional Champollion representations predict biological phenotypes and supports interpretation through latent traversal. It uses harmonization, residualization, cross-validated classification, permutation testing, and a variance-based interpretability constraint.

  • Exploratory analysis: Exploratory classification used regional latent spaces after site harmonization and residualization on sex, age, and socioeconomic status when available.The framework generated brain-wide predictive maps before interpreting the most informative regions through latent traversal.
  • Exploratory analysis: The analyzed phenotypes included pronounced versus absent left incomplete hippocampal inversion, very prematurity versus full term, and maternal smoking exposure during pregnancy.The IHI task used QTIM, prematurity used ABCD, and smoking exposure used UK Biobank with ABCD replication.
  • Statistical analysis: Predictive performance was quantified with ROC-AUC using L2-penalized logistic regression, balanced class weights, and 5-fold cross-validation for regularization selection.The regularization parameter was selected from a predefined grid within each cross-validation fold.
  • Statistical analysis: Permutation testing assessed associations conditional on confounders by permuting only residualized predictors in the training set.The procedure followed the ter Braak scheme and used a regularization coefficient selected from non-permuted cross-validation.
  • Interpretation: Latent-traversal directions were selected with a minimum 2% explained global variance to preserve visual interpretability while reducing classifier performance by less than 2%.The direction was chosen from performance curves relating regularization strength to explained variance ratio.
  • Stability analysis: Prediction-direction stability was assessed by bootstrap resampling through mean pairwise Pearson correlations across runs.Prematurity was selected as the least stable exploratory task for a representative visual-interpretation analysis.

6 Data availability

The study uses data from UK Biobank, ABCD, QTIM, and the Human Connectome Project Young Adult database.

  • UK Biobank data were used under Application Number 64984.
  • ABCD data supported the prematurity study and come from a multisite, longitudinal cohort of children followed into early adulthood.
  • Publicly available QTIM MRI data were used for incomplete hippocampal inversion classification.
  • Human Connectome Project Young Adult data were used for the twins study.

9 Author contributions

The authors divided the work across research conception, model development, analyses, evaluation, visualization, and code development.

  • J. Laval, P. Gori, D. Rivière, J. Chavas, and J.-F. Mangin conceived the research idea and experiments.
  • J. Laval wrote the manuscript, designed and optimized the model, and conducted latent-space analyses and interpretation.
  • R. Guiavarch evaluated foundation models, while A. Dufournet conducted genetic experiments and designed the decoder.
  • Other contributors developed the brain-wide association pipeline, unified code, and visualizations, and contributed to analysis and data-related work.

11 Extended Data

The extended data describe the datasets, Champollion architecture, decoder, morphometric analyses, genetic associations, and exploratory analyses.

  • Datasets: UK Biobank provided SSL pretraining data, whereas all other listed datasets were used only for downstream analyses.
  • Architecture: Champollion embeds local cortical-skeleton crops with a 12-layer convolutional network into a 32-dimensional latent space.
  • Decoder: The decoder comprises six convolutional layers and is trained for 10 epochs with Binary Cross Entropy loss, a 5 × 10−4 learning rate, and batch size 32.
  • Morphometric analyses: Figure 9 reports R² for predicting each BrainVisa-labeled sulcal morphometric descriptor from the Champollion latent space using the smallest containing ROI.
  • Genetic analyses: Figure 10 compares MOSTest genetic associations for two ROIs encoded by Champollion against associations from four classical sulcal morphometric measurements.
  • Terminology: The acronym table defines region names composed from BrainVISA abbreviations for sulci, fissures, and cortical areas.
  • Exploratory analyses: The extended analyses also include maternal-smoking classification and latent traversal of directions obtained from the highest-scoring ROI.
Loading 2609.05438v1…