Source-linked AI summary
The 2026 PNPL Competition: Word Classification and Efficient Cross-Subject Generalisation in LibriBrain100
Francesco Mantegna, Gereon Elvers, Dulhan Jayalath, Gilad Landau, Tasha Kim, Miran Özdogan, Luisa Kurth, Teyun Kwon, SungJun Cho, Benjamin Ballyk, Alex Fung, Anna Greer, Pratik Somaiya, Christian Herff, Yorguin Mantilla Ramos, Hamza Abdelhedi, Karim Jerbi, Greg Farquhar, Brendan Shillingford, Mark Woolrich, Oiwi Parker Jones
TL;DR
The competition addresses the need for non-invasive speech decoding to generalise beyond heavily sampled individuals and uses LibriBrain100 to benchmark this challenge. It introduces standardised word classification with Deep and Broad tracks, targeting within-subject performance and efficient cross-subject adaptation. The Broad track progressively limits subject-specific fine-tuning and includes zero-shot evaluation, while the competition uses balanced accuracy and vocabulary-aware evaluation.
Problem
Non-invasive speech decoding needs evaluation of cross-subject generalisation and comparable word-classification performance beyond single-subject, heavily sampled settings.
Method
The competition uses LibriBrain100 and two tracks: Deep for densely sampled within-subject decoding and Broad for adaptation to new subjects with limited fine-tuning data.
Results
The Broad track evaluates progressively data-limited adaptation, including zero-shot cross-subject generalisation, while word classification is measured with BAcc@10, BAcc@1, and OVMI.
Takeaways & Limitations
Word classification offers a standardised intermediate benchmark for measuring representation learning, cross-subject generalisation, and decoding from non-invasive neural data.
Takeaways & Limitations
The baseline study leaves joint training on Subject 0 and Subjects 1–32, which may improve generalisation, for future exploration.
Abstract
from arXiv · showhide
The ambition of the 2025 PNPL competition (Landau et al., 2025) was to launch a multi-year curriculum for non-invasive speech decoding. Designed to progress from foundational tasks toward the linguistic complexity required for a practical brain-computer interface (BCI), it set the stage with speech detection and phoneme classification tasks. Winning submissions reached F1-macro scores of 95.6% and 73.6% on the respective tasks (Elvers et al., 2026), highly significant advances. This success was built on the LibriBrain dataset (Özdogan et al., 2025), the largest within-subject MEG dataset recorded at the time with ${\sim}50$ hours of data for one subject. However, while within-subject scale drives strong decoding performance, a practical BCI must generalise to new users from minutes of data, not hours. The 2026 PNPL competition responds to this challenge with LibriBrain100 (Mantegna et al., 2026), an extended LibriBrain dataset with 32 additional subjects (${\sim}40$ minutes each) plus even more within-subject data (${\sim}80$ hours). Advancing the curriculum of tasks to focus on word classification, two complementary tracks are presented in this competition: the Deep track targets within-subject word classification at scale, aiming at the best possible performance; the Broad track targets cross-subject generalisation, progressively reducing the amount of subject-specific fine-tuning data from ${\sim}40$ to ${\sim}20$ to ${\sim}10$ minutes, a duration that falls within a clinically feasible range and brings us a step closer to a non-invasive BCI capable of restoring communication to people living with profound paralysis.
1 Competition description
The 2026 PNPL competition advances non-invasive speech decoding through standardised word classification on LibriBrain100, combining deep within-subject data with systematic cross-subject evaluation. Its Deep and Broad tracks target peak performance and data-efficient adaptation to new individuals.
- Evaluation: The competition standardises word evaluation with a custom 50-word vocabulary, the Moses 50 vocabulary, and OVMI to account for accuracy and vocabulary coverage.These methods address the comparability problems caused by differing target vocabularies across studies.
- Tracks: The Broad track evaluates adaptation to new individuals with progressively limited subject-specific data, including a zero-shot regime.This track is designed to assess data-efficient cross-subject generalisation under clinically relevant constraints.
- Task: Word classification provides a structured intermediate task between sub-lexical decoding and open-vocabulary brain-to-text.It directly probes meaningful linguistic units while remaining suitable for clear benchmarking and broad participation.
- Competition resource: LibriBrain100 combines approximately 80 hours from one subject with data from 32 additional subjects, enabling within-subject and cross-subject benchmarking.This dataset supports the competition’s two-track design by pairing unusually deep recordings with a multi-subject cohort.
- Tracks: The Deep track evaluates word classification in a single densely sampled subject and seeks the best possible decoding performance.Participants may train on any data they wish for Subject 0.
- Evaluation: BAcc@10 is the competition metric, with 20% expected from uniform random guessing for a 50-word vocabulary, while BAcc@1 helps distinguish ranking quality.BAcc@10 counts whether the correct word appears among the top ten predictions but does not distinguish first from tenth place.
2 Organisation
The competition is designed for broad, accessible participation while separating within-subject and cross-subject word-classification objectives into two tracks. Standardised infrastructure, public communication, and automated evaluation support reproducible submissions.
- Competition schedule: Participants can submit to either or both tracks during the three-month competition period from 15 July to 15 October.Progress is reflected on the appropriate public leaderboard for each track.
- Competition tracks: The competition offers Deep and Broad tracks for word classification over a fixed vocabulary, ranked separately.The Deep track focuses on one deeply sampled subject, whereas the Broad track evaluates many held-out subjects.
- Accessibility: Participation is open to everyone, without requiring domain-specific knowledge or specialised hardware.The stated rules aim to encourage open participation while protecting evaluation integrity.
- Permitted materials: Participants may use external public datasets, self-collected data, and other permitted training materials.The permitted-materials policy is described as applying to either track.
- Communication and community: The competition supports communication through a Discord server, website, GitHub, public leaderboards, and planned blog posts.These channels provide information, resources, updates, and discussion throughout the competition.
3 Resources
The competition provides a Python library and public reference resources to make data access and model development accessible within standard deep-learning workflows.
- Software and reference resources: The pnpl Python library downloads the data and integrates it with standard deep-learning frameworks.Reference model code and checkpoints are publicly available on GitHub.
- Software and reference resources: Tutorial code runs in Google Colab, providing browser-based access to some GPU resources.This lowers infrastructure requirements for getting started.
Organising team
The organising team comprises researchers and students from PNPL, the University of Oxford, the Université de Montréal, Mila, Maastricht University, Google DeepMind, and related institutions.
- University of Oxford and PNPL: PNPL and the University of Oxford contribute researchers and students across engineering, machine learning, neuroscience, and computational neuroscience.The listed members include Oiwi Parker Jones, Mark Woolrich, Francesco Mantegna, Gereon Elvers, Dulhan Jayalath, Gilad Landau, and others.
- External academic affiliations: Christian Herff is an Associate Professor and heads the Neural Interfacing Lab at Maastricht University.
- Université de Montréal and Mila: The organising team includes researchers affiliated with the Université de Montréal and Mila, including Karim Jerbi, Yorguin Mantilla Ramos, and Hamza Abdelhedi.Their roles span psychology, Neuro-AI, and graduate research.
- Industry affiliations: Greg Farquhar and Brendan Shillingford are Staff Research Scientists at Google DeepMind.
A Dataset Comparison
LibriBrain100 is compared with public MEG speech datasets by cohort breadth, single-subject depth, and total recording time. Its distinctive position combines a competitive subject count with substantially greater recording scale and depth than most existing datasets.
- Dataset dimensions: LibriBrain100 combines a competitive number of subjects with substantially more total recording time and greater single-subject depth than most existing datasets.The comparison uses number of subjects, maximum hours per subject, and bubble area proportional to total recording hours.
- Figure encoding: Figure 2 uses logarithmic axes for subject count and maximum hours per subject, with each bubble representing one dataset.Bubble area encodes total recording hours.
- Dataset expansion: The arrow highlights the expansion from LibriBrain to LibriBrain100.
B Retrieval Vocabulary
The competition evaluates word classification with two separate 50-word vocabularies and uses competition-vocabulary BAcc@10 as its primary leaderboard metric. The vocabularies support complementary goals: reliable evaluation and comparison with assistive-communication word sets used in invasive studies.
- Only BAcc@10 on the competition vocabulary determines competition leaderboards and prize decisions.
- Competition vocabulary: The competition vocabulary was selected for frequent occurrence and coverage of multiple grammatical categories, supporting low-variance evaluation and meaningful short utterances.
- Moses 50: Moses 50 is a patient-informed 50-word list emphasizing clinical, daily-living, social, and affective communication needs.
- Moses 50: Reporting Moses 50 enables direct comparison with invasive decoding studies, although some words are inadequately attested in LibriBrain100.
- The two vocabularies share 13 words and together cover 87 distinct word types while supporting short assistive and conversational utterances.
C Stimulus Materials
Subject 0 heard material spanning literary, acoustic-phonetic, articulatory, and podcast sources, while Subjects 1–32 heard LibriBrain validation and test recordings.
- Subject 0 listened to the complete Sherlock Holmes canon, TIMIT, MOCHA-TIMIT, and 30 podcast stories from The Moth.
- Subjects 1–32 listened to the validation and test recordings from LibriBrain.
D Code Samples
The recommended participation workflow uses the pnpl Python library, which provides dataset access, PyTorch dataloaders, task configurations, and Kaggle submission utilities.
- The pnpl library provides ready-to-use PyTorch dataloaders and automatically downloads LibriBrain100 from Hugging Face when needed.
- The code imports LibriBrain100 and the WordClassification task configuration from pnpl.
- Each dataset item returns x as channel-by-time input data and y as an integer word-class identifier.
- The library includes utilities to write prediction files and submit them directly to the Kaggle competition.