Source-linked AI summary
Taming foundation model with invariance-oriented pre-training for broad-spectrum EEG analysis across signal-level, brain-state, and brain-health tasks
Yulong Dou, Han Wu, Guo Chen, Fangmao Ju, Zhiming Cui, Dinggang Shen
TL;DR
EEG foundation models need representations that transfer across subjects, recording conditions, and diverse analysis tasks rather than relying solely on locally predictable signal content. INCEPT addresses this gap with invariance-oriented contextual pre-training, achieving the most consistent performance among recent EEG foundation models across ten datasets spanning signal-level, brain-state, and brain-health tasks. In objective ablations, invariance learning yields a 15.5% average relative gain across 12 metric-dataset pairs over masked modeling alone.
Problem
EEG variability across subjects, devices, montages, task contexts, and artifacts creates a need for representations that transfer across broad-spectrum analysis tasks.
Method
INCEPT combines masked contextual learning with cross-view invariance learning to retain stable subject-sensitive information while reducing reliance on nuisance variability.
Results
INCEPT delivers the most consistent performance among recent EEG foundation models across ten datasets, while invariance learning improves over masked modeling by 15.5% on average across 12 metric-dataset pairs.
Takeaways & Limitations
INCEPT provides a reusable representational starting point for signal-level assessment, brain-state decoding, and brain-health evaluation.
Takeaways & Limitations
Benchmark performance does not by itself establish clinical utility because retrospective classification datasets rarely capture expert interpretation, longitudinal outcomes, or workflow-level effects.
Abstract
from arXiv · showhide
Electroencephalography (EEG) is a widely used window into human brain function, but most EEG models remain tied to a one-dataset-one-model supervised paradigm. Recent EEG foundation models offer a route toward reusable representations, but most remain reconstruction-centered, assuming that EEG content predictable from local context is necessarily transferable neural information. Here we present INCEPT, an invariance-oriented EEG foundation model trained on over 11,000 hours of unlabelled clinical EEG. Rather than prioritizing signal recovery alone, INCEPT learns representation-level stability across correlated EEG observations, separating stable neural structure and essential subject-sensitive information from the nuisance variability that dominates scalp recordings while preserving subject-, state- and condition-discriminative information. We evaluate INCEPT on a broad-spectrum benchmark of ten datasets spanning three levels of post-acquisition EEG analysis: signal-level assessment, brain-state decoding, and brain-health evaluation. INCEPT ranks first among recent EEG foundation models on 26 of 30 linear-probing metrics and 24 of 30 fine-tuning metrics, and also surpasses strong task-specific specialist encoders across diverse downstream settings. Objective ablations and representation analyses further show that invariance-oriented pre-training improves transfer and organizes subject-sensitive neural representations beyond reconstruction alone. These results establish invariance learning as a promising principle for building reusable EEG foundation models.
1 Introduction
EEG supports broad-spectrum analysis from signal quality to brain-state and brain-health evaluation, but scalp recordings contain substantial nuisance variability that limits reusable modelling. INCEPT addresses this challenge with invariance-oriented contextual pre-training designed to learn transferable whole-brain representations from unlabelled EEG and is evaluated across ten downstream datasets.
- Motivation: EEG analysis spans signal-level assessment, brain-state decoding, and brain-health evaluation, progressing from recording quality to transient states and clinically relevant variation.Signal-level assessment considers quality, abnormality, and reliability; brain-state decoding infers physiological or behavioural states; brain-health evaluation examines clinically relevant variation.
- Motivation: Scalp EEG varies substantially across subjects, devices, montages, task contexts, and recording conditions because measurements combine neural activity with anatomy, referencing, placement, and artifacts.Although EEG is non-invasive, inexpensive, portable, and temporally precise, its low signal-to-noise ratio complicates reusable modelling.
- Motivation: Most deep learning EEG approaches train task-specific supervised models from scratch on a single dataset, protocol, and label space, encouraging features entangled with acquisition and cohort factors.These models can perform well when training and test conditions are tightly matched, but their learned features may not transfer broadly.
- Contribution: INCEPT combines masked contextual learning with invariance-oriented pre-training to capture local temporal, spectral, and cross-channel dependencies while retaining essential subject-sensitive information stable across observations.The model is designed to learn transferable whole-brain representations from unlabelled EEG.
- Evaluation: INCEPT is evaluated across ten downstream datasets spanning signal-level EEG assessment, brain-state decoding, and brain-health evaluation under linear probing and fine-tuning.The benchmark includes abnormal EEG detection, artifact recognition, emotion recognition, motor imagery, sleep staging, mental-stress recognition, depression-related assessment, neurodegenerative disease evaluation, and seizure detection.
2 Results
INCEPT supports EEG analysis across signal-level assessment, brain-state decoding, and brain-health evaluation, with frozen representations often transferring strongly and fine-tuning helping task-specific boundaries. Results show particularly strong performance for artifact recognition, high-density affective decoding, sleep staging, cognitive-demand assessment, and neurodegenerative classification.
- Signal-level assessment: INCEPT supports complementary signal-level assessment: frozen representations favor clinical abnormality screening, while fine-tuning improves artifact-specific recognition.On TUAR, fine-tuning raises INCEPT to 61.45% balanced accuracy, 69.32% weighted F1 and 52.62% kappa, whereas TUAB performance decreases after fine-tuning.
- Brain-state decoding: 31.0%, 32.4% and 74.9% are INCEPT’s linear-probing improvements over the strongest task-specific encoder on SEED-V balanced accuracy, weighted F1 and kappa.INCEPT reaches 46.91% balanced accuracy, 45.74% weighted F1 and 33.00% kappa, while full fine-tuning remains highest across all three metrics.
- Brain-state decoding: INCEPT remains the strongest foundation model under both protocols for ISRUC-S1 sleep staging despite the six-channel montage shift.SPaRCNet remains a competitive supervised baseline, reaching 80.59% weighted F1.
- Brain-health evaluation: 95.51% balanced accuracy, 99.47% AUROC and 99.50% AUC-PR are INCEPT’s linear-probing results on Mumtaz2016, exceeding the strongest task-specific encoder across all three metrics.The passage characterizes this as a frozen-representation pattern with performance near ceiling.
- Brain-health evaluation: 14.8% and 27.0% are INCEPT’s linear-probing improvements over the average of other frozen EFMs on MentalArithmetic AUROC and AUC-PR.INCEPT reaches 86.17% AUROC, 72.34% AUC-PR and 62.08% balanced accuracy; fine-tuning lowers balanced accuracy to 59.72%.
- Brain-health evaluation: 65.48%, 67.75% and 52.29% are INCEPT’s full-fine-tuning results on ADFTD balanced accuracy, AUROC and AUC-PR, preserving its lead.ADFTD evaluates three-class neurodegenerative disease classification from resting-state EEG and shows both a frozen-readout advantage and a small additional fine-tuning gain.
3 Discussion and conclusion
INCEPT is positioned as a reusable EEG foundation backbone that supports frozen linear probing and full fine-tuning across signal-level, brain-state, and brain-health tasks. The discussion emphasizes that invariance-oriented pre-training should preserve subject-sensitive neural organization while addressing reconstruction’s limitations, benchmark-to-clinic gaps, and insufficient standardization.
- Reusable foundation backbone: INCEPT representations remain readable when frozen and adaptable when fine-tuned across new montages, temporal contexts, and label semantics.Linear probing tests information accessibility from a frozen backbone, whereas fine-tuning tests reshaping for new downstream requirements.
- Invariance-oriented representation: Transferable EEG representations should preserve subject-sensitive physiology and recording structure rather than remove all subject-specific information.The discussion identifies individual physiology, electrode geometry, oscillatory traits, functional coupling, and recording-specific baselines as components of observed EEG organization.
- Invariance-oriented representation: Reconstruction alone is incomplete because predictable contextual structure can include artifacts, reference effects, montage regularities, device characteristics, and acquisition-site patterns.Masked contextual modelling remains useful for learning local temporal, spectral, and cross-channel dependencies.
- Limitations and future needs: Benchmark performance does not by itself establish clinical utility, because retrospective labels rarely capture expert interpretation, longitudinal outcomes, treatment response, disease progression, uncertainty, or workflow effects.The passage distinguishes the value of controlled benchmark transfer evaluation from evidence of clinical usefulness.
- Limitations and future needs: EEG foundation-model progress is constrained by inconsistent montages, sampling rates, references, preprocessing, annotations, hardware, cohorts, and task designs.The discussion calls for shared preprocessing practices, reporting standards, metadata, quality-control procedures, and evaluation protocols.
- Reusable foundation backbone: INCEPT provides a common representational starting point for broad-spectrum post-acquisition EEG analysis instead of separate dataset-specific models.It spans signal-level assessment, brain-state decoding, and brain-health evaluation.
4 Methods
INCEPT is a self-supervised EEG pre-training framework that uses dynamically sampled, multi-scale correlated views to learn invariance across neural observations. Its pipeline combines contextual view objectives, signal-and-topology-aware token embeddings, and large-scale clinical EEG pre-training.
- Dynamic spatiotemporal sampling: INCEPT dynamically samples macro-level, micro-level, and masked macro-level views from each EEG segment to expose global context, local dynamics, and missing-context inference.Macro-level views preserve broad brain-state context, micro-level views capture localized neural dynamics, and masked macro-level views introduce block-wise missing regions.
- Dynamic spatiotemporal sampling: Dynamic spatial and temporal extents expose variable electrode coverage and recording duration, reducing dependence on a specific montage or segment length.The sampler constructs views with variable spatial and temporal ranges rather than fixed-size crops.
- Dynamic spatiotemporal sampling: Random block-wise masking removes contiguous patch regions across electrodes and time while preserving view size, requiring inference of missing neural context.The masking is applied to each macro-level view, and the visible patches provide context for reconstructing the missing neural information.
- Neural token embedding: Neural token embeddings combine time-domain and FFT-derived spectral features with spherical-harmonics and Cartesian electrode-position information.The spatial encoding preserves montage-aware geometry and supports variable electrode subsets across datasets.
- Self-supervised pre-training: Pre-training uses multi-scale inputs to align representations across spatiotemporal scales, recover missing contextual information, and maintain a well-distributed latent space.The model is pre-trained on approximately 27,077 hours from 69,672 EDF files covering 14,987 patients and 26,846 recording sessions in TUEG.
Code availability
The authors will release the study’s code upon acceptance, including preprocessing, INCEPT implementation, pre-training, and downstream evaluation components.
- Code availability: Code will be released upon acceptance, with availability during peer review if editors or reviewers require it.The public repository will include preprocessing scripts for all datasets, the INCEPT architecture, self-supervised pre-training, and downstream evaluation code.
Appendix A Dataset information · A.1 Dataset Overview
The study pre-trains on one large-scale unlabeled clinical EEG corpus and evaluates INCEPT across ten heterogeneous downstream datasets. These datasets span clinical, cognitive, sleep, motor-imagery, and affective scenarios, testing transfer across varied acquisition and label conditions.
- A.1 Dataset Overview: The benchmark covers abnormal EEG detection, depression-related classification, neurodegenerative disease classification, mental stress detection, seizure detection, motor imagery classification, sleep staging, and emotion recognition.Together, these tasks represent clinical, cognitive, sleep, motor-imagery, and affective EEG scenarios.
- A.1 Dataset Overview: The datasets vary substantially in cohort composition, acquisition setting, electrode montage, channel density, recording duration, and label structure.This heterogeneity broadens the conditions under which INCEPT is evaluated.
- A.1 Dataset Overview: The heterogeneous benchmark is designed to assess whether INCEPT learns transferable EEG representations rather than representations optimized for a single dataset.The evaluation therefore spans multiple EEG scenarios and dataset characteristics.
- A.1 Dataset Overview: Table 1 reports task-level signal settings and available demographic metadata for the pre-training corpus and all downstream evaluation datasets.The table covers both the pre-training corpus and the complete downstream evaluation set.
- A.1 Dataset Overview: Unavailable demographic information is marked with “/” because public EEG datasets differ in metadata completeness.For datasets with subject-level metadata, age is reported by sex or diagnostic group according to the original metadata organization.
A.2 Dataset Details
The study uses TUEG as an unlabeled clinical pre-training corpus and evaluates EEG representations across signal-quality, brain-health, stress, seizure, sleep, motor-imagery, and emotion tasks. All downstream inputs share resampling and channel-wise normalization, while dataset-specific protocols are retained when required.
- Pre-training corpus: TUEG provides the unlabeled self-supervised pre-training corpus and contains recordings from 14,987 subjects in routine hospital settings.The available metadata reports 48.8% male participants and mean ages of 49.3 ± 20.1 years for males and 50.1 ± 20.5 years for females.
- Signal-level tasks: TUAB supports binary abnormal-EEG detection using normal-versus-abnormal labels from clinical recordings originally sampled at 256 Hz.Signals are band-pass filtered between 0.3 and 75 Hz, notch-filtered at 60 Hz, and divided into non-overlapping 10-s windows.
- Signal-level tasks: TUAR supports segment-level artifact recognition, retaining background EEG and three artifact categories after excluding chewing and shivering because they have few usable samples.The original annotations include eye movement, chewing, shivering, muscle artifact, electrode-related artifact, and background EEG.
- Brain-health and brain-state tasks: The benchmark includes clinical and cognitive-state datasets for depression, neurodegenerative disease, mental stress, seizures, sleep stages, and motor imagery.Mumtaz2016 targets binary depression-related classification; ADFTD uses three classes; MentalArithmetic detects rest versus mental arithmetic; Siena detects seizures; ISRUC-S1 uses five sleep stages; and PhysioNet-MI uses four motor-imagery classes.
- Brain-state tasks: FACED and SEED-V provide nine-class and five-class emotion-recognition tasks, respectively, with 123 and 16 subjects.FACED covers anger, disgust, fear, sadness, neutral emotion, amusement, inspiration, joy, and tenderness; SEED-V covers happy, sad, neutral, disgust, and fear.
- Shared preprocessing: All downstream EEG signals are resampled to 250 Hz and channel-wise z-score normalized after segmentation, while dataset-specific filtering, segmentation, and splits are retained when required.The shared preprocessing is consistent with the temporal resolution used during self-supervised pre-training.
Appendix B Experimental details … B.3 Evaluation metrics
Appendix B details the supervised and foundation-model baselines, downstream adaptation protocols, and task-appropriate evaluation metrics used to assess INCEPT. The evaluation distinguishes frozen-representation accessibility from joint adaptation and uses metrics suited to binary or multi-class labels.
- B.1 Baseline models: INCEPT is compared with task-specific supervised EEG encoders trained separately on each labelled dataset and EEG foundation models pre-trained on unlabelled recordings before adaptation.The supervised baselines represent the one-dataset-one-model paradigm, whereas foundation-model baselines represent large-scale pre-training.
- B.1 Baseline models: EEGNet, ST-Transformer, EEGConformer, and SPaRCNet serve as task-specific supervised encoders using convolutional, attention-based, hybrid, or DenseNet-style temporal architectures.EEGConformer and SPaRCNet are trained from scratch on each downstream dataset.
- B.1 Baseline models: CBraMod, CSBrain, and CodeBrain are reconstruction-centered EEG foundation models that learn representations through masked reconstruction or masked token prediction.Their designs incorporate time-frequency features, cross-scale spatiotemporal structure, or separate temporal and frequency-domain codebooks.
- B.2 Adaptation protocols: Linear probing freezes the pre-trained encoder and trains a lightweight readout, directly testing whether downstream-relevant information is separable in the learned representation space.For ISRUC-S1 sleep staging, frozen epoch representations are passed to a lightweight sequence head for epoch-wise prediction.
- B.2 Adaptation protocols: Full fine-tuning jointly optimizes the pre-trained encoder and classification head, using separate learning rates with a smaller rate for the backbone and a larger rate for the new task head.The classification head generally uses fully connected layers with nonlinear activation and dropout.
- B.2 Adaptation protocols: Downstream models use AdamW, cross-entropy with label smoothing, gradient clipping, cosine learning-rate scheduling, validation-based checkpoint selection, and five random seeds reported as mean ± standard deviation.Final evaluation uses the best validation checkpoint on the held-out test set.
- B.3 Evaluation metrics: Binary tasks use balanced accuracy, AUROC, and AUC-PR, while multi-class tasks use balanced accuracy, weighted F1 score, and Cohen’s kappa coefficient.Balanced accuracy reduces majority-class dominance; weighted F1 incorporates class support; Cohen’s kappa corrects agreement for chance; AUC-PR emphasizes positive identification under low false-positive burden and rare positive classes.
Appendix C Supplementary results
Appendix C presents supplementary downstream results organized by each dataset’s relationship to the 19-channel 10-20 montage used during pre-training. It provides the complete metric set while distinguishing montage-matched from montage-shifted datasets.
- The supplementary analysis reorganizes downstream results by the relationship between each dataset’s montage and the 19-channel 10-20 pre-training configuration.
- It reports the complete metric set for each downstream dataset.
- The analysis separates datasets into montage-matched and montage-shifted groups.
C.1 Results on montage-matched downstream datasets
On montage-matched datasets, task-specific supervised encoders remain competitive in several settings, while INCEPT shows strong, stable performance with a frozen encoder and remains competitive after fine-tuning. Its strongest result is near-ceiling performance on Mumtaz2016 across three binary-classification metrics under both evaluation regimes.
- Overall performance: Task-specific supervised encoders remain competitive on montage-matched datasets, particularly TUAB and Mumtaz2016.These datasets have electrode layouts close to the 19-channel 10-20 configuration used during pre-training.
- Linear probing: EEG foundation models show more stable behaviour under linear probing, with INCEPT performing strongly using a frozen encoder.The trend is most evident on TUAB, TUAR, Mumtaz2016 and ADFTD.
- Linear probing: INCEPT remains competitive across all metrics on TUAB, TUAR, Mumtaz2016 and ADFTD under linear probing.The passage specifically identifies these datasets as showing the clearest trend for INCEPT’s frozen-encoder performance.
- Mumtaz2016: On Mumtaz2016, INCEPT reaches near-ceiling performance across three binary-classification metrics under both linear probing and full fine-tuning.The three metrics are not specified in the supplied passage.
C.2 Results on montage-shifted downstream datasets … D.3.1 (i) Depression- and stress-related EEG assessment
Across montage-shifted and downstream EEG tasks, INCEPT’s invariance-oriented pre-training supports adaptable, stable transfer across altered montages, signal-quality assessments, brain-state decoding, and brain-health evaluation. Its advantages include stronger fine-tuned performance, improved reproducibility, and competitive frozen representations across diverse datasets.
- C.2 Results on montage-shifted downstream datasets: Montage-shifted datasets show a clearer separation between frozen readout and supervised adaptation, with full fine-tuning generally outperforming linear probing.These datasets vary in electrode layout, channel density, segment duration, or task context.
- D.1.1 (i) Clinical abnormality assessment from routine EEG: 88.0% lower seed-to-seed standard deviation in balanced accuracy is achieved by EEG foundation models than task-specific supervised encoders on TUAB under linear probing.Within the foundation-model regime, INCEPT improves balanced accuracy by 3.9% relative to the average task-specific supervised encoder and by 2.3% relative to the average of other foundation models.
- D.1.2 (ii) Artifact-type recognition for EEG quality control: INCEPT achieves the strongest TUAR balanced accuracy and Cohen’s kappa among EEG foundation models under linear probing, while weighted F1 remains comparable to the best foundation-model result.Relative to the average task-specific supervised encoder, INCEPT improves balanced accuracy by 17.9%, weighted F1 by 9.6%, and kappa by 26.5%.
- D.2.1 (i) Affective-state decoding across low- and high-density emotion EEG: 55.8% lower weighted-F1 standard deviation is achieved by EEG foundation models than task-specific supervised encoders on FACED under linear probing.Among foundation models, INCEPT improves weighted F1 by 16.8% relative to the average of the other EEG foundation models.
- D.2.1 (i) Affective-state decoding across low- and high-density emotion EEG: 57.6% higher weighted F1 than the average task-specific supervised encoder is achieved by INCEPT on SEED-V under linear probing, with full fine-tuning preserving the advantage.SEED-V has shorter segments, higher channel density, and lower overall agreement across models.
- D.2.2 (ii) Sensorimotor-state decoding from high-density motor-imagery EEG: 13.8% higher weighted F1 and 27.8% higher kappa than other EEG foundation models are achieved by INCEPT on PhysioNet-MI under linear probing.Only INCEPT reaches performance comparable to task-specific supervised encoders in this setting.
- D.2.3 (iii) Sleep-state staging from sparse clinical sleep montages: 11.1% higher average weighted F1 than task-specific supervised encoders is achieved by EEG foundation models on ISRUC-S1, while seed-to-seed standard deviation falls by 83.6%.INCEPT further improves balanced accuracy by 3.5%, weighted F1 by 2.1%, and kappa by 2.8% relative to other EEG foundation models under linear probing.
- D.3.1 (i) Depression- and stress-related EEG assessment: On brain-health datasets, foundation models improve transfer and reproducibility: on Mumtaz2016, average foundation models gain balanced accuracy by 6.0%, AUROC by 1.6%, and AUC-PR by 0.6%.On MentalArithmetic, INCEPT shows the clearest frozen transfer among foundation models, while score-level separability and final class-decision performance differ.
D.3.2 (ii) Neurodegenerative disease evaluation from resting-state EEG · D.3.3 (iii) Seizure-related neurological assessment under extended clinical montages · Appendix E Scaling trends
INCEPT delivers stronger and more stable transfer for neurodegenerative-disease evaluation and excels in AUROC for seizure-related assessment under extended clinical montages. Scaling experiments vary unlabelled-EEG volume and trainable parameter count to test whether representation quality improves with pre-training resources and model capacity.
- D.3.2 (ii) Neurodegenerative disease evaluation from resting-state EEG: 19.3% balanced-accuracy, 23.3% weighted-F1, and 50.8% kappa improvements distinguish INCEPT from the average task-specific supervised encoder on ADFTD under linear probing.INCEPT provides stronger and more stable transfer than task-specific supervised encoders, whose performance varies substantially across architectures.
- D.3.2 (ii) Neurodegenerative disease evaluation from resting-state EEG: 19.1% balanced-accuracy, 19.5% weighted-F1, and 47.6% kappa improvements distinguish INCEPT from the average of other EEG foundation models on ADFTD under linear probing.Against the strongest existing EEG foundation model, INCEPT improves balanced accuracy by 10.6%, weighted F1 by 10.0%, and kappa by 24.6%.
- D.3.2 (ii) Neurodegenerative disease evaluation from resting-state EEG: INCEPT shows a small but consistent full-fine-tuning increase on ADFTD, while reconstruction-centered models exhibit larger standard deviations and INCEPT remains below 1.0 across all three metrics.CBraMod and CSBrain decrease after fine-tuning, CodeBrain improves modestly, and INCEPT remains stable.
- D.3.3 (iii) Seizure-related neurological assessment under extended clinical montages: 70.3% AUC-PR, 16.3% AUROC, and 13.9% balanced-accuracy improvements give the average of four EEG foundation models an advantage over task-specific supervised encoders on Siena under linear probing.Under full fine-tuning, the corresponding improvements are 71.1%, 19.5%, and 14.7%.
- D.3.3 (iii) Seizure-related neurological assessment under extended clinical montages: INCEPT achieves the strongest AUROC on Siena under both linear probing and full fine-tuning, while CBraMod achieves higher balanced accuracy and AUC-PR.Under linear probing, INCEPT improves AUROC by 6.9% over the other-foundation-model average and by 5.2% over the strongest existing EEG foundation model; after fine-tuning, the improvement is 4.3% relative to the other-model average.
- Appendix E Scaling trends: Scaling experiments vary unlabelled-EEG pre-training amount and trainable model parameters, evaluating resulting models on two representative downstream datasets under the same full-fine-tuning protocol.The stated purpose is to examine whether INCEPT exhibits scaling behaviour in EEG foundation pre-training.
E.1 Increasing pre-training data size improves downstream transfer · E.2 Increasing model size improves downstream transfer
Scaling both pre-training data and model capacity generally improves INCEPT’s downstream transfer. Data scaling is especially beneficial under acquisition and brain-state shifts, while model scaling improves performance across evaluated datasets.
- E.1 Increasing pre-training data size improves downstream transfer: INCEPT’s data-scaling study sampled 10%, 20%, 30%, 40%, 60%, 80% and 100% of an approximately 11,000-hour corpus while holding architecture, optimization and fine-tuning fixed.The 100% setting matched the full INCEPT pre-training configuration.
- E.1 Increasing pre-training data size improves downstream transfer: Data scaling was evaluated with downstream balanced accuracy under full fine-tuning, using scores normalized to 100% pre-training data and means across five random seeds.The figure also reports mean ± s.d. and individual seed results.
- E.1 Increasing pre-training data size improves downstream transfer: Downstream performance generally improved with more pre-training data, with the clearest effect on FACED despite differences in montage, electrode density and affective brain-state context.This pattern indicates that larger unlabelled EEG corpora provide more transferable representation-learning information.
- E.1 Increasing pre-training data size improves downstream transfer: Data scaling was more modest on Mumtaz2016, whose closer match to the clinical 10-20 pre-training recordings may let smaller subsets capture relevant acquisition geometry and resting-state structure.Cohort size, label noise and depression-related classification difficulty may also limit the observed improvement.
- E.2 Increasing model size improves downstream transfer: INCEPT model variants ranged from 9.61 million to 202.26 million parameters, varying input embedding dimension and Transformer-layer count while using the full corpus and identical evaluation protocols.Variants were evaluated on Mumtaz2016 and FACED under full fine-tuning.
- E.2 Increasing model size improves downstream transfer: Downstream performance generally improved as model size increased on both Mumtaz2016 and FACED, suggesting larger variants better exploit the same unlabelled EEG corpus across acquisition settings.Potentially relevant capacities include long-range temporal dependencies, cross-channel interactions and global segment-level structure.
- E.2 Increasing model size improves downstream transfer: Overall, increasing unlabelled EEG data improved representation quality, particularly under montage and brain-state shift, while increasing model capacity further enhanced downstream transfer.These results frame EEG foundation modelling as a scalable representation-learning problem and support invariance-oriented pre-training as an organizing principle.
Appendix F Ablation studies
Appendix F evaluates how INCEPT’s components contribute to downstream transfer through staged ablations on montage-matched and montage-shifted EEG datasets. The results indicate consistent, cumulative gains as multi-band augmentation, dynamic multi-view sampling, and dual-domain embedding are introduced.
- Experimental setup: The ablation compares INCEPT variants on Mumtaz2016, which matches the 19-channel 10-20 pre-training montage, and FACED, which tests a 32-channel low-density 10-10 montage shift.This design contrasts a closely aligned evaluation with a more challenging montage-shifted scenario.
- Experimental setup: All variants use the same unlabeled EEG corpus and downstream full fine-tuning protocol, isolating the contribution of staged architectural and training components.The basic baseline uses raw EEG without augmentation, time-domain patch embedding only, Cartesian electrode coordinates, and deterministic multi-scale views.
- Component contributions: Multi-band PCA-based augmentation perturbs signals in a low-dimensional principal subspace to improve robustness against amplitude and spectral variability.Without this module, the model learns directly from raw normalized EEG segments.
- Component contributions: The ablation shows consistent and cumulative improvement as components are introduced, with dynamic multi-view sampling and dual-domain signal embedding further improving transfer.Dynamic views address downstream differences in channel number and segment duration, while dual-domain embedding jointly models time- and frequency-domain information.