Source-linked AI summary

iMINDBench: iEEG Multi-Institution Neural Decoding Benchmark

Geeling Chau, Saba Hashemi, Yonghyeon Gwon, Eshani Patel, Jan DeWitt, Christopher Wang, Andrii Zahorodnii, Sabera J Talukder, Danny Dongyeop Han, Chun Kee Chung, Maryam M Shanechi, Yisong Yue

arXiv:2609.18104v1cs.LG

TL;DR

iEEG decoding lacks reliable cross-task and cross-institution evaluation because datasets, splits, and preprocessing vary. iMINDBench standardizes these factors across fifteen tasks and three naturalistic datasets, finding that pretrained systems generally beat within-track baselines but strong spectral baselines remain competitive and broader supervision yields limited, task-dependent gains. The benchmark therefore exposes how strongly model advantages depend on representation and evaluation setting.

  • Problem

    iEEG decoding lacks consistent evidence of generalization across tasks and institutions because datasets, evaluation splits, and preprocessing pipelines vary.

  • Method

    iMINDBench evaluates fifteen decoding tasks across three naturalistic movie-watching datasets using fixed splits, standardized preprocessing tracks, and controlled supervised-data scaling domains.

  • Results

    Pretrained systems generally outperform baselines within their preprocessing tracks, while strong Multi-STFT baselines remain competitive and broader supervision produces smaller, task-dependent gains than within-session data.

  • Takeaways & Limitations

    Progress requires models that improve over strong spectral preprocessing baselines and exploit heterogeneous cross-subject and cross-institution data more effectively.

  • Takeaways & Limitations

    The evaluation covers three naturalistic movie-watching datasets, assumes target-session calibration labels, and evaluates broader scaling with a single pretrained model.

Abstract

from arXiv · show

Intracranial electroencephalography (iEEG) is widely used to record electrical activity directly from electrodes inside the human brain, making it an attractive modality for neural decoding. However, progress in iEEG decoding, especially toward general-purpose foundation models, remains difficult to measure reliably: datasets are task- or institution-specific, limiting evidence of generalization across tasks and recording environments, and preprocessing choices can strongly influence performance, making model improvements difficult to distinguish from preprocessing gains. Thus, we introduce iMINDBench, an iEEG Multi-Institution Neural Decoding Benchmark that evaluates models on a shared suite of fifteen decoding tasks across three naturalistic movie-watching datasets. The benchmark additionally defines standardized preprocessing tracks and fixed evaluation splits to support consistent model comparisons. Using iMINDBench, we find that the evaluated pretrained systems generally outperform baselines within their respective preprocessing tracks, while strong spectral baselines remain competitive across institutional datasets. In our scaling study, adding up to 25 times more supervised data from other subjects or institutions yields only small or task-dependent gains over within-session training. Together, these findings highlight the need for iEEG models that improve on strong preprocessing baselines and make more effective use of data across subjects and institutions. Project website: https://imindbench.github.io/

1 Introduction

iMINDBench addresses inconsistent datasets, tasks, preprocessing, and splits by providing a shared multi-institution benchmark for fifteen naturalistic iEEG decoding tasks. Its evaluations show that pretrained systems usually beat within-track baselines, but strong spectral baselines remain competitive and broader supervision gives limited, task-dependent gains.

  • Benchmark contribution: iMINDBench spans fifteen language, auditory, and visual tasks across naturalistic movie-watching datasets from three institutions, with fixed splits and specified preprocessing tracks.Results are reported separately by dataset and preprocessing track to expose variation across institutions and input representations.
  • Benchmark findings: Pretrained systems generally outperform non-pretrained baselines within their respective preprocessing tracks on Main units, while strong Multi-STFT baselines remain competitive across institutional datasets.This establishes engineered spectral preprocessing as an important reference for evaluating learned representations.
  • Benchmark findings: Adding subjects from the same dataset is broadly beneficial, but gains are smaller than those from additional target-session data and weaker or more task-dependent across institutions.The scaling study tests whether aligned tasks and neural-signal metadata translate broader supervision into target-session improvements.
  • Significance: The benchmark enables systematic evaluation of whether model benefits generalize across tasks, institutions, and preprocessing tracks or remain specific to the evaluation setting.Its controls are intended to separate modeling advances from changes in evaluation setup.

2 Related Work

Prior iEEG resources support either multi-institution clinical harmonization or naturalistic decoding, but typically remain specialized by task or institution. Foundation-model evaluations also vary in inputs, preprocessing, training exposure, and split units, motivating shared evaluation conditions.

  • Existing iEEG benchmarks: Multi-institution iEEG benchmarks harmonize clinical recordings across sites but primarily target seizure, pathology, high-frequency oscillation, or related epileptology tasks.These resources address institutional integration without providing broad naturalistic decoding coverage.
  • Existing iEEG benchmarks: Naturalistic decoding resources provide interpretable tasks, but evaluations often remain tied to a single dataset or institutional setting.This limits direct evidence about generalization across recording environments.
  • Relationship to EEG benchmarks: Scalp EEG benchmarks and foundation-model evaluations do not transfer cleanly to iEEG because cortical recordings retain high-frequency activity and lack standardized electrode montages.Clinically determined electrode placement breaks assumptions behind many conventional EEG pipelines.
  • Foundation-model evaluation: iEEG and EEG foundation models differ in waveform, spectrogram, token, embedding, spatial-metadata, corpus, and split choices, leaving evaluations heterogeneous and limited.These differences make shared conditions necessary for interpreting model comparisons.

3 Benchmark Design

iMINDBench combines three heterogeneous movie-watching iEEG datasets with aligned tasks, metadata, fixed splits, scaling domains, decodability subsets, and standardized preprocessing tracks. The design preserves institutional variation while making model comparisons more consistent across datasets, tasks, and input representations.

  • Datasets and harmonization: The benchmark combines Neuroprobe/Brain Treebank, BYD, and Pippi, whose recording duration, electrode coverage, subject/session structure, and institutional practices differ.Neuroprobe provides larger within-session training sets, BYD a different institutional and stimulus context, and Pippi a smaller independent dataset.
  • Tasks and domains: Fifteen aligned binary classification tasks cover language, auditory, and visual domains, pairing each example with a 1-second neural window and stimulus label.Examples include speech presence, sentence onset, surprisal, pitch, volume, optical flow, brightness, and face count.
  • Splits and domains: Two-fold within-session splits use largely contiguous movie portions, fit preprocessing statistics on training data only, and average final scores across folds.Each dataset–subject/session–task combination is treated as an evaluation unit.
  • Splits and domains: Scaling domains add supervised data progressively from the target session, other units in the same dataset, and other institutions while preserving target units and evaluation splits.These regimes test whether broader supervision improves decoding beyond target-session calibration labels.
  • Unit decodability: Main, Challenge, and All subsets distinguish units decodable by baseline within-session models from units near chance under screening decoders.Challenge units are defined by both screening scores being at or below 0.60, but this does not establish absent task-relevant neural signal.
  • Preprocessing and reporting: Multi-STFT and Waveform tracks standardize distinct neural inputs, with the spectral track using multi-resolution frequency representations and the waveform track retaining less-processed signals.Reporting includes dataset, model family, preprocessing track, pretraining exposure, and mean ROC-AUC, with Overall defined as the unweighted mean of three dataset scores.

4 Benchmark Results

Across the benchmark, pretrained systems generally outperform non-pretrained baselines within matched preprocessing tracks, but strong spectral baselines remain competitive. Preprocessing choices materially affect baseline performance, while broader supervised data provide smaller and task-dependent gains than within-session scaling.

  • 4.1 Decoding Performance: PopT-v2 achieves 0.645 Overall ROC-AUC versus 0.631 for the strongest non-pretrained Multi-STFT baseline, while Multi-STFT remains competitive with pretrained waveform systems.
  • 4.1 Decoding Performance: Waveform performance depends on architecture: HTNet beats generic baselines across datasets, while BaRISTA and DIVER-1 improve further with dataset-dependent relative strengths.
  • 4.2 Preprocessing: Global Norm improves ROC-AUC by more than 0.10 in some settings, whereas omitting the 0.5-Hz high-pass filter reduces mean ROC-AUC from 0.612 to 0.589 across all datasets.
  • 4.2 Preprocessing: Multi-STFT is retained because single-STFT optima vary across datasets and tasks, despite comparable aggregate performance.
  • 4.3 Data Scaling: Within-session scaling increases mean test ROC-AUC by approximately 0.025 per training-data doubling, while broader scaling yields smaller, task-dependent gains.

5 Release and Extensibility

iMINDBench will release the evaluation infrastructure needed for reproducible comparisons and extend it through a hosted leaderboard and future benchmark additions.

  • The release includes evaluation code, preprocessing configurations, baseline implementations, dataset preparation pipelines, and benchmark split metadata.
  • A hosted website and leaderboard will compare submissions across datasets, tasks, preprocessing tracks, scaling domains, and unit decodability subsets.
  • Future releases may add datasets, tasks, preprocessors, models, or hidden-test evaluation while preserving archived benchmark versions.

6 Discussion

The discussion emphasizes that preprocessing and institutional heterogeneity shape both model comparisons and the usefulness of additional supervision. Learned representations must therefore be evaluated against strong spectral baselines and made robust to heterogeneous recordings.

  • Strong Multi-STFT baselines remain competitive with pretrained waveform systems, especially when downstream training sets are small.
  • Matched preprocessing is essential because normalization and filtering choices can otherwise make apparent model advantages reflect the input pipeline.
  • Within-session scaling produces the largest gains, whereas broader supervision yields smaller, task-dependent improvements across heterogeneous institutions.
  • Residual differences in referencing, electrode geometry, anatomical coverage, amplifiers, and environmental noise may make pooled institutional data difficult to exploit.

7 Limitations and Future Work

The benchmark’s scope and conclusions are bounded by naturalistic movie datasets, target-session labels, limited scaling-model coverage, and study-specific tuning choices.

  • The evaluation covers three naturalistic movie-watching datasets and assumes access to target-session calibration labels.Broader generalization would require other stimulus paradigms and evaluation without target-session labels.
  • Adding new datasets can require manual alignment of electrode locations, brain-area labels, other metadata, and dataset-specific stimulus processing.Automated alignment and validation are proposed to reduce curation effort while preserving protocol consistency.
  • Scaling experiments evaluate one pretrained model with joint supervised fine-tuning across subjects, while other pretrained systems use linear heads over frozen output embeddings.Additional models are needed to test how broadly the observed scaling pattern holds.
  • Model comparisons depend on the training recipes and tuning budgets used, and further optimization could change relative performance.The study reports a human-guided AutoResearch feasibility study in which tuning a pretraining recipe improved downstream decoding.

8 Conclusion

iMINDBench provides a shared framework for evaluating iEEG decoding across tasks, institutions, and preprocessing tracks. Its results show that model advantages depend on representation and setting, while additional cross-subject or cross-institution supervision does not reliably improve target-session decoding.

  • Model advantages depend on the input representation and evaluation setting, while additional supervised data from other subjects and institutions do not reliably improve target-session decoding.The conclusion motivates models that extract useful features under limited supervision and learn from heterogeneous recordings.
  • iMINDBench uses matched-preprocessing comparisons and cross-institution evaluations to measure progress toward general-purpose iEEG decoding.The framework spans tasks, institutions, and preprocessing tracks.
  • Future progress requires representations robust or adaptable to differences in anatomy, acquisition, and recording conditions.This consequence is stated within the benchmark’s supported evaluation scope.

A Dataset Harmonization and Acquisition Metadata

iMINDBench harmonizes decoding labels and electrode metadata across independently collected movie datasets without treating their acquisition conditions as uniform.

  • Cross-institution evaluation aligns decoding labels and iEEG electrode metadata as two benchmark-specific dimensions.The alignment supports common decoding targets and spatial metadata across heterogeneous source datasets.
  • Task and temporal alignment: Task and label alignment applies the Neuroprobe/Brain Treebank stimulus-feature pipeline to BYD and Pippi videos where source annotations support it.Task labels, thresholds, and lexical/nonverbal row construction follow specified protocols.
  • Electrode metadata alignment: Electrode locations are transformed into Brain Treebank’s left-posterior-inferior coordinate space, while brain-area labels are harmonized to the Destrieux parcellation.The procedure addresses differing coordinate systems and anatomical labels.
  • Remaining acquisition differences: Remaining acquisition and stimulus differences are summarized using metadata reported in the source dataset papers.
  • Remaining acquisition differences: The datasets retain differences in stimulus duration, language, electrode type and coverage, sampling, reference conventions, and coordinate reporting.BYD and Pippi focus on macroelectrode and clinical sEEG recordings, respectively.

B Decoding Tasks

The benchmark defines aligned binary decoding tasks and evaluates them with fixed-track scorecards, decodability subsets, and task- and model-family-specific analyses.

  • Scalar binary tasks use bottom-versus-top quartiles in Neuroprobe and bottom-versus-top terciles in BYD and Pippi to improve class support.Lexical and surprisal tasks are restricted to word rows.
  • BYD and Pippi visual and auditory task tables combine lexical word rows with sampled nonverbal rows from gaps longer than 2.0 s.Nonverbal rows provide the negative class for speech and sentence-onset tasks.
  • All aligned tasks use binary classification on 1-second neural windows at word onset or sampled nonverbal-interval onset, with rebalanced classes.
  • Table 4 reports within-session ROC-AUC on all supported units, grouped by fixed-track and model-native model families, with best and second-best entries marked.
  • Table 5 reports the corresponding within-session ROC-AUC scorecard on Challenge units.
  • Figure 5 compares fold-averaged ROC-AUC for Main and Challenge units using STFT screening performance and a validation screening score.Dashed and solid lines indicate the Main-unit threshold and chance performance, respectively.
  • Figure 6 breaks out Main-unit mean test ROC-AUC by task and model family, with dataset means and SEM across three dataset means.
  • Figure 8 evaluates within-session sample efficiency as target-session training data increase from one-sixteenth to the full set, retaining complete-case units.Panel titles report unique supported subjects, and dashed lines mark chance performance.

G Task-Level Scaling Trajectories

Figure 9 presents task-level scaling trajectories for PopT-v2 across three datasets, comparing within-dataset and multi-dataset training against within-session evaluation. The broader section also documents preprocessing tracks and model-specific input adaptations relevant to interpreting these comparisons.

  • Task-Level Scaling Trajectories: Figure 9 reports each task’s ROC-AUC change relative to within-session evaluation across within-dataset and multi-dataset training.Within-session values are fixed at zero; positive values indicate improvement, and rows correspond to Neuroprobe, BYD, and Pippi.
  • Model Coverage: MVPFormer and Brant were excluded from the main comparison until their temporal input adaptations were validated against one-second benchmark windows.Their released configurations use longer temporal contexts or patches than the benchmark target window.

I.2 Benchmark Runtime Estimates

The benchmark reports estimated preprocessing and model-fitting times, while the interactive leaderboard supports decomposed comparisons across datasets, tasks, tracks, and scaling domains. An AutoResearch feasibility experiment additionally links lower pretraining validation loss with improved within-session decoding.

  • Runtime Estimates: Timing records are normalized to four concurrent workers and are estimates rather than measured end-to-end wall-clock times.The assumed hardware is two 32-core AMD EPYC 7513 processors for preprocessing and CPU fitting, plus four 80-GB NVIDIA A100 GPUs for neural fitting.
  • Runtime Estimates: Estimated preprocessing and model-fitting times are reported separately for the benchmark workflows.The preprocessing estimate is based on HTNet runs whose cache was reused by waveform Logistic, MLP, and CNN models; fitting estimates exclude subject loading and preprocessing.
  • Benchmark Viewer: The leaderboard decomposes comparisons across datasets, tasks, preprocessing tracks, scaling domains, and unit-decodability subsets.This supports examining runtime and performance results by evaluation dimension.
  • AutoResearch Feasibility: ΔROC-AUC = 0.0066 for the best AutoResearch-selected PopT-v2 variant versus the original 1M-step single-STFT baseline on Neuroprobe within-session decoding.The 95% confidence interval is 0.0021–0.0111 with p = 0.0049.
Loading 2609.18104v1…