Source-linked AI summary

A foundation model of vision, audition, and language for in-silico neuroscience

Stéphane d'Ascoli, Jérémy Rapin, Yohann Benchetrit, Teon Brooks, Katelyn Begany, Joséphine Raugel, Hubert Banville, Jean-Rémi King

arXiv:2605.04326v1q-bio.NCcs.LG

TL;DR

Cognitive neuroscience lacks a unified model that connects findings from specialized experimental paradigms. TRIBE v2 integrates audio, video, and language representations to predict high-resolution fMRI across diverse datasets, subjects, and conditions. The results support its use as a predictive and interpretable framework for in silico neuroscience, while its scope remains bounded by fMRI resolution, input modalities, and heterogeneous preprocessing.

  • Problem

    Specialized neuroscience models have produced detailed findings, but a unified account of brain function remains difficult to synthesize.

  • Method

    TRIBE v2 maps audio, video, and language stimuli to high-resolution fMRI using pretrained AI embeddings and a unified dataset spanning over 1,000 hours across 720 subjects.

  • Results

    TRIBE v2 predicts cortical responses across diverse naturalistic and experimental conditions, generalizes to unseen participants, and supports in silico hypothesis testing within a unified framework.

  • Takeaways & Limitations

    The results support unified predictive foundation models as a framework for studying brain and cognitive functions and interpreting neural function through intervention.

  • Takeaways & Limitations

    TRIBE v2 is constrained by fMRI’s spatio-temporal resolution, currently omits several primary sensory modalities, and combines datasets with heterogeneous preprocessing pipelines.

Abstract

from arXiv · show

Cognitive neuroscience is fragmented into specialized models, each tailored to specific experimental paradigms, hence preventing a unified model of cognition in the human brain. Here, we introduce TRIBE v2, a tri-modal (video, audio and language) foundation model capable of predicting human brain activity in a variety of naturalistic and experimental conditions. Leveraging a unified dataset of over 1,000 hours of fMRI across 720 subjects, we demonstrate that our model accurately predicts high-resolution brain responses for novel stimuli, tasks and subjects, superseding traditional linear encoding models, delivering several-fold improvements in accuracy. Critically, TRIBE v2 enables in silico experimentation: tested on seminal visual and neuro-linguistic paradigms, it recovers a variety of results established by decades of empirical research. Finally, by extracting interpretable latent features, TRIBE v2 reveals the fine-grained topography of multisensory integration. These results establish artificial intelligence as a unifying framework for exploring the functional organization of the human brain.

1 Introduction

Neuroscience has generated detailed findings through specialized models, but these results remain difficult to synthesize into a unified account of brain function. TRIBE v2 addresses this gap with a tri-modal foundation model designed to integrate, predict, generalize, and interpret brain responses across conditions.

  • 1 Introduction: Specialized studies have mapped functions such as motion, faces, and written language to distinct neural substrates, but their results remain fragmented.This fragmentation makes it difficult to synthesize how neuronal assemblies represent and integrate information into a coherent model of the surrounding world.
  • 1 Introduction: Pretrained vision, language, and audio models offer a route to predicting brain responses because their latent representations can align with the representational geometry of the primate brain.Prior work has demonstrated direct prediction of brain responses to natural images and videos using such representations.
  • 1 Introduction: A foundation model of human brain function must integrate whole-brain responses across experimental conditions, achieve strong predictive performance, generalize to novel conditions, and remain interpretable.These criteria are framed as integration, performance, generalization, and interpretability.
  • 1 Introduction: TRIBE v2 is a tri-modal foundation model that predicts high-resolution fMRI from audio, video, and language stimuli using pretrained AI embeddings.The approach is illustrated with naturalistic and experimental stimuli, including movies, podcasts, flashed objects, and isolated words.
  • 1 Introduction: Across more than 1,000 hours of fMRI recordings from 720 subjects, TRIBE v2 is evaluated for cortical prediction across diverse naturalistic and experimental conditions and for in silico hypothesis testing.The authors present the model as a unified framework intended to accelerate neuroscientific discovery.

2 Results

TRIBE v2 predicts naturalistic brain responses across cortical and subcortical regions, generalizes to unseen subjects, and outperforms a linear baseline. It also reproduces established visual and language neuroscience findings while revealing modality-specific and multisensory organization.

  • Encoding performance across naturalistic tasks: TRIBE v2 predicts responses above chance across wide cortical and subcortical regions, with task-dependent peaks in temporal, visual, and multimodal cortical areas.Subcortical predictions remain significant in most areas but are generally two to three folds lower than cortical scores.
  • Comparison to baselines: Across all datasets, TRIBE v2 significantly outperforms the optimized linear baseline, with q(FDR) < 10^-4.The baseline uses the same pretrained embeddings, isolating the comparison to the model architecture and nonlinear integration.
  • Scaling: TRIBE v2’s encoding accuracy increases log-linearly with training data on Courtois NeuroMod without reaching a plateau.An earlier iteration also achieved first place among 263 teams in the Algonauts 2025 competition.
  • Generalization to new subjects: On unseen subjects, TRIBE v2 predicts group-averaged responses zero-shot, reaching Rgroup near 0.4 on HCP, two-fold above the median subject’s group-predictivity.Finetuning on at most one hour per participant further improves subject-specific encoding by two- to four-fold over a linear encoder trained from scratch.
  • In-silico experiments: vision: In visual localizers, TRIBE v2 reproduces expected hemodynamic timing and recovers established areas including FFA, PPA, EBA, and VWFA.Predicted contrast maps qualitatively match measurements from the original experiments and show significant spatial correlation.
  • In-silico experiments: language: In language paradigms, TRIBE v2 recovers core language areas, pain-related TPJ and MTG responses, and left-hemisphere lateralization for linguistic contrasts.These results support qualitative agreement between model predictions and classic visual and neurolinguistic experiments.
  • Interpretability and multimodality: ICA components resemble well-studied functional networks, while multimodal encoding improves most around the temporal-parietal-occipital junction, with gains up to 50%.Text, audio, and video dominate different cortical regions, indicating complementary modality contributions and modeled interactions.

3 Discussion

TRIBE v2 is presented as a unified, predictive framework for integrating brain responses across people and experimental conditions. Its current scope remains bounded by fMRI resolution, limited sensory inputs, and modeling of the brain as a passive observer.

  • Discussion: TRIBE v2 supports a shift from fragmented cognitive-task mapping toward unified foundation models of brain and cognitive functions.The model aligns AI representations with human-brain responses across a broad fMRI repertoire.
  • Discussion: The model’s observed log-linear scaling of encoding accuracy suggests that prediction of human brain activity has not reached a ceiling.
  • Limitations: fMRI resolution prevents TRIBE v2 from capturing millisecond neuronal dynamics.
  • Limitations: Current inputs omit olfaction, balance, somatosensation, and their integration, while the model does not represent active behavior, development, or clinical pathology.
  • Discussion: TRIBE v2 provides a platform for interpreting neural function through intervention and testing theories against accumulated neuroimaging data.

5 Methods

The methods transform video, audio, and transcript inputs into synchronized pretrained-model embeddings, combine them across modalities, and use them to predict cortical and subcortical fMRI responses. The pipeline preserves temporal information while compressing modality-specific representations for efficient modeling.

  • 5 Methods: TRIBE v2 predicts BOLD signals at 20,484 cortical vertices and 8,802 subcortical voxels from naturalistic stimuli.
  • 5 Methods: The model accepts video, audio, and transcripts, extracting embeddings from frozen pretrained text, audio, and video models.The frozen feature extractors are intended to support robustness to out-of-distribution data.
  • 5 Methods: A trainable transformer aggregates the multimodal embedding time series before a subject-conditioned layer predicts regional brain responses.
  • 5 Methods: Text, audio, and video features are aligned on a 2 Hz time grid before modality-specific representations are combined.Text embeddings are temporally aligned with audio and video features; audio and video representations are resampled or constructed at 2 Hz.
  • 5 Methods: Each modality is compressed to a shared dimension of 384, yielding multimodal embeddings of dimension 1,152 for the transformer encoder.The three modality representations are concatenated after layer normalization.

5.3 Model

The model processes 100-second multimodal windows with a subject-aware transformer and includes modality and subject dropout to support incomplete inputs and zero-shot prediction for unseen subjects.

  • 5.3 Model: TRIBE v2 applies an 8-layer, 8-head transformer to 100-second windows and resamples outputs from 2 Hz to the 1 Hz fMRI rate.
  • 5.3 Model: Modality dropout randomly masks inputs with probability 0.3 to reduce reliance on any single modality and support missing-modality predictions.
  • 5.3 Model: A subject-conditional linear layer projects transformer outputs to cortical or subcortical targets while modeling subject-specific responses.
  • 5.3 Model: Subject dropout with probability 0.1 trains a special unseen-subject layer for zero-shot group-response prediction.

5.4 Training and validation

TRIBE is trained against fMRI responses with mean-squared error and validated using stimulus-disjoint splits and Pearson encoding scores. Fine-tuning adapts the model to new subjects, with low-rank factorization controlling the size of new subject blocks.

  • 5.4 Training and validation: TRIBE is optimized with mean-squared error using AdamW, a warmed-up learning rate of 10^-4, cosine decay, and validation-based early stopping.
  • 5.4 Training and validation: Validation prevents stimulus overlap between training and held-out data, and encoding score is the mean Pearson correlation across subjects and parcels.
  • 5.4 Training and validation: Fine-tuning runs for one epoch with all TRIBE parameters unfrozen and the same training hyperparameters.
  • 5.4 Training and validation: New-subject blocks are initialized from the average subject block and factorized with rank 128 to reduce parameter size.The factorization uses SVD, with the resulting low-rank component used as the new subject block.
  • 5.4 Training and validation: The baseline replaces the transformer with a linear convolution spanning 9 TRs and offset by 5 seconds.This temporal receptive field aggregates stimulus information across time.

5.6 Statistics

TRIBE v2 is evaluated across curated deep training datasets and broad testing datasets, with statistical procedures used to assess prediction reliability and model performance.

  • Statistics: Statistical significance in figure 2 is assessed by comparing estimated correlations with a null distribution based on independent Gaussian random vectors.To address fMRI autocorrelation, only one TR every 60 seconds is retained for this test.
  • Statistics: Model performances in figures 2 and 3 are compared using paired t-tests with FDR correction across subjects.
  • Datasets: Eight fMRI datasets are curated for training and testing, spanning naturalistic speech, video, and multimodal stimuli.Training uses four deep datasets, while testing uses four broad datasets.
  • Datasets: Training combines four deep datasets totaling 3, 8, 10, and 4 subjects, with 35, 85, 61, and 265 hours respectively.The datasets cover silent videos, speech, videos with sound but no speech, and multimodal videos.
  • Datasets: Testing combines four broad datasets totaling 433 subjects and 326 hours for speech listening, plus 262 subjects and 338 hours for movie watching.These datasets expose many participants to relatively small volumes of naturalistic stimuli.

5.8 fMRI preprocessing

The preprocessing workflow harmonizes heterogeneous fMRI datasets, maps signals into cortical and subcortical spaces, standardizes sampling, and offsets signals for the hemodynamic lag.

  • Workflow: Because datasets use diverse preprocessing pipelines and formats, the study adopts a harmonized workflow rather than a single identical fMRIPrep workflow.Missing metadata and non-BIDS-compliant organization made identical processing infeasible.
  • Signal extraction: Cortical signals are projected from volumetric data onto the fsaverage surface using template-specific FreeSurfer surfaces and 3 mm-radius ball sampling.
  • Signal extraction: Subcortical extraction yields 8,802 voxels spanning eight Harvard-Oxford atlas regions.The regions include the hippocampus, amygdala, thalamus, caudate, putamen, pallidum, accumbens, and lateral ventricles.
  • Standardization: Timeseries are z-scored per session, detrended, and linearly resampled to f_fMRI = 1 Hz across datasets.Detrending prevents slow drifts from becoming spurious predictors, especially with the model’s long context window.
  • Hemodynamic alignment: To account for the roughly 5-second hemodynamic delay, TRIBE predicts fMRI window [0,T] from stimuli window [-5,T-5].

5.9 In-silico experiments

TRIBE v2 is used in unseen-subject mode to reproduce controlled visual and language experiments, generating contrast maps through task-specific analysis procedures.

  • Protocol: The in-silico replications use TRIBE in unseen-subject mode to reproduce visual and language tasks from the IBC dataset.
  • Visual experiments: Visual FaceBody and Visu images are randomized, presented for one second, and separated by eight seconds by converting them into static videos.
  • Language experiments: Language replications cover Bang, Audio, EmotionalPain, and RSVP, with stimuli presented as speech, audio segments, or sentences.Bang contrasts speech-containing and speech-free segments from the Algonauts dataset because the original movie was unavailable.
  • Contrast maps: Visual contrast maps select predicted responses five seconds after image onset and subtract the mean response for the other categories.
  • Contrast maps: Language contrast maps use a GLM whose regressors are formed by convolving predicted BOLD responses with the canonical hemodynamic response function.

5.10 Independent component analysis

The analysis decomposes TRIBE’s unseen-subject mapping from latent features to cortical space with ICA, then compares the resulting components with functional cortical maps.

  • ICA decomposition: ICA is applied to TRIBE’s unseen-subject layer, a tensor of shape (Dmodel, Ntargets) mapping latent features to each subject’s cortical space.
  • ICA decomposition: FastICA extracts five components, each represented as a vector of Ntargets values that can be plotted on the cortical surface.The component order has no meaning because ICA, unlike PCA, does not impose an ordered ranking.
  • Functional characterization: The components are functionally characterized against five NeuroSynth maps labeled primary auditory, language, motion, default network, and visual.
  • Functional characterization: Each NeuroSynth map is resampled from volumetric to cortical space and compared with ICA components using Pearson correlation across cortical vertices.

A Related works

Existing brain-encoding approaches are limited by linear mappings, subject/task/region specificity, and unimodal inputs. TRIBE v2 extends this line of work with broader coverage, unseen-subject evaluation, and generalization to controlled stimuli.

  • Limitations of current encoding models: Existing encoding models typically use ridge regression, assuming a linear relationship between AI representations and brain representations.
  • Limitations of current encoding models: Subject-, task-, and brain-area-specific models cannot learn co-occurring patterns in large heterogeneous datasets.
  • Limitations of current encoding models: Unimodal encoding approaches cannot capture multisensory integration, including cross-modal interactions in primary sensory areas.
  • Multimodal encoders: Multimodal encoding studies report gains over unimodal transformers but still rely on linear mappings from transformer activations to brain responses.
  • Comparison to TRIBE v1: Compared with TRIBE v1, TRIBE v2 predicts 20k cortical vertices plus subcortical regions, trains across 25 participants and four tasks, and tests on 695 unseen participants.
  • Comparison to TRIBE v1: TRIBE v2 generalizes to controlled images from the IBC protocol, enabling in-silico experimentation that was not tested in TRIBE v1.
Loading 2605.04326v1…