# 2026 PNPL 竞赛：LibriBrain100 上的词汇分类与高效跨被试泛化

- 来源：HuggingFace Daily Papers（社区热门论文）
- 发布时间：2026-09-03 08:00
- AIHOT 分数：41
- AIHOT 链接：https://aihot.news/items/cmtrhifyt03inrotnoubtk08j
- 原文链接：https://arxiv.org/abs/2609.03231

## AI 摘要

2026 PNPL 竞赛聚焦非侵入式语音解码，基于扩展数据集 LibriBrain100（新增 32 名被试，每人约 40 分钟数据，单被试数据增至约 80 小时）推进词汇分类任务。竞赛设两条赛道：Deep track 面向大规模被试内词汇分类，Broad track 则针对跨被试泛化，将被试专属微调数据从约 40 分钟逐步缩减至约 10 分钟，向临床可行的非侵入式脑机接口迈进。

## 正文

, Department of Engineering Science, University of Oxford, UK

francesco@robots.ox.ac.uk

Gereon Elvers

PNPL

, Department of Engineering Science, University of Oxford, UK

gereon@robots.ox.ac.uk

Dulhan Jayalath

PNPL

, Department of Engineering Science, University of Oxford, UK

dulhan@robots.ox.ac.uk

Gilad Landau

PNPL

, Department of Engineering Science, University of Oxford, UK

oiwi@robots.ox.ac.uk

Tasha Kim

PNPL

, Department of Engineering Science, University of Oxford, UK

Miran Özdogan

PNPL

, Department of Engineering Science, University of Oxford, UK

Luisa Kurth

PNPL

, Department of Engineering Science, University of Oxford, UK

Teyun Kwon

PNPL

, Department of Engineering Science, University of Oxford, UK

SungJun Cho

PNPL

, Department of Engineering Science, University of Oxford, UK

OHBA, Oxford Centre for Integrative Neuroimaging, University of Oxford, UK

Benjamin Ballyk

PNPL

, Department of Engineering Science, University of Oxford, UK

Alex Fung

PNPL

, Department of Engineering Science, University of Oxford, UK

FMRIB, Oxford Centre for Integrative Neuroimaging, University of Oxford, UK

Anna Greer

PNPL

, Department of Engineering Science, University of Oxford, UK

Pratik Somaiya

PNPL

, Department of Engineering Science, University of Oxford, UK

Christian Herff

Maastricht University, The Netherlands

Yorguin Mantilla Ramos

UNIQUE, Université de Montréal, Canada

Hamza Abdelhedi

UNIQUE, Université de Montréal, Canada

Karim Jerbi

UNIQUE, Université de Montréal, Canada

Mila–Quebec AI Institute, Canada

Greg Farquhar

Google DeepMind, UK

Joint first authors

Brendan Shillingford

Google DeepMind, UK

Joint first authors

Mark Woolrich

OHBA, Oxford Centre for Integrative Neuroimaging, University of Oxford, UK

Oiwi Parker Jones

PNPL

, Department of Engineering Science, University of Oxford, UK

Abstract

The ambition of the 2025 PNPL competition (Landau et al., 2025) was to launch a multi-year curriculum for non-invasive speech decoding. Designed to progress from foundational tasks toward the linguistic complexity required for a practical brain-computer interface (BCI), it set the stage with speech detection and phoneme classification tasks. Winning submissions reached F1-macro scores of 95.6% and 73.6% on the respective tasks (Elvers et al., 2026), highly significant advances. This success was built on the LibriBrain dataset (Özdogan et al., 2025), the largest within-subject MEG dataset recorded at the time with ∼50 hours of data for one subject. However, while within-subject scale drives strong decoding performance, a practical BCI must generalise to new users from minutes of data, not hours.

The 2026 PNPL competition responds to this challenge with LibriBrain100 (Mantegna et al., 2026), an extended LibriBrain dataset with 32 additional subjects (∼40 minutes each) plus even more within-subject data (∼80 hours). Advancing the curriculum of tasks to focus on word classification, two complementary tracks are presented in this competition: the Deep track targets within-subject word classification at scale, aiming at the best possible performance; the Broad track targets cross-subject generalisation, progressively reducing the amount of subject-specific fine-tuning data from ∼40 to ∼20 to ∼10 minutes, a duration that falls within a clinically feasible range and brings us a step closer to a non-invasive BCI capable of restoring communication to people living with profound paralysis.

1 Competition description

Figure 1: Subject-wise data splits. (a) Schematic illustration of the dataset. A single extensively recorded subject is shown alongside other subjects with progressively reduced amounts of available training data. (b) Quantitative representation of the data splits. A broken y-axis is used to accommodate the large difference between sub-0 and the other subjects, while preserving the visibility of differences in training data proportions. Vertical dashed lines separate the subject groups, and labels indicate the proportion of training data available within each group. A decreasing curve illustrates the overall reduction in available training data across subjects.

1.1 Background and impact

Invasive brain-computer interfaces (BCIs) for speech have advanced at a remarkable pace. Since the first demonstration of connected-speech decoding from surgically implanted electrodes in a paralysed individual (Moses et al., 2021), vocabulary sizes have rapidly grown from 50 words to over 125,000 words (Willett et al., 2023) while word-error rates (WERs) have shrunk to 2.5% (Card et al., 2024)—a superhuman score well below the ∼5–10% reported for human transcription on standard automatic speech recognition (ASR) benchmarks (Xiong et al., 2016). Yet invasive BCIs face fundamental barriers to widespread deployment: brain surgery carries inherent risks, per-patient data collection does not easily scale, and the populations who might benefit most are often those least able to undergo elective neurosurgery; non-invasive alternatives are therefore the real prize.

Among non-invasive neuroimaging modalities, magnetoencephalography (MEG) stands out for its combination of millisecond temporal resolution and typical spatial precision around 5–10 mm (Hämäläinen et al., 1993)—with the capacity to go even lower to 2–4 mm (Barratt et al., 2018)—putting it in the range of invasive modalities like ECoG but without the surgical risks. Recent years have seen genuine advances in MEG-based speech decoding (Défossez et al., 2023; d’Ascoli et al., 2025; Jayalath et al., 2025b), but progress has been harder to interpret and harder to compare across studies than in the invasive setting. A key reason is the lack of shared infrastructure: common datasets, fixed evaluation splits, standard metrics, and baseline implementations that allow the community to build cumulatively on each other’s work.

The PNPL competition series is motivated by the view that closing this gap will require not only better models, but better community infrastructure. The 2025 PNPL competition (Landau et al., 2025) was designed with this goal in mind. Building on the LibriBrain dataset (Özdogan et al., 2025)—the largest within-subject MEG dataset recorded at the time, with ∼50 hours of data from a single participant—it introduced common benchmarks for speech detection and phoneme classification, together with a Python library for automatic data download and loading, standardised train/validation/test splits for replicability, reference models, a public leaderboard, tutorial code, and a dedicated competition website. These tasks were deliberately chosen to be relatively simple: speech detection reduces to binary classification over time samples, while phoneme classification is a well-defined supervised learning problem with manageable output structure. This was part of a broader curricular design—by starting with accessible tasks, we aimed to lower the barrier to entry for researchers from machine learning who might not yet have the domain knowledge to confidently navigate human brain data, while laying the groundwork for more ambitious forms of neural speech decoding. Winning submissions reached F1-macro scores of 95.6% and 73.6% on the respective tasks, driving significant advances in the field (Elvers et al., 2026). Over the summer of 2025, the 2025 PNPL competition attracted 155 registered teams from 15 countries and generated 6,041 submissions. By April 2026, the LibriBrain dataset had reached over 16,000 downloads — indicating a clear demand for the materials we provided, even after the 2025 competition ended.

The 2026 PNPL competition represents a logical next step, with significantly more data and larger ambitions. A central reason this next step is now feasible is the scale of the expanded LibriBrain100 dataset (Mantegna et al., 2026). Compared with existing public MEG speech datasets, LibriBrain100 occupies a distinctive position: it combines substantially greater within-subject depth with a larger multi-subject cohort, supporting both high-performance within-subject modelling and systematic evaluation of cross-subject generalisation (visualised in Appendix A). This combination of depth and breadth is essential for the two-track design of the competition: the Deep track exploits the unusually large amount of data available for a single subject, while the Broad track tests whether models trained across subjects can adapt efficiently to new individuals. The task of word classification is a natural progression in the curriculum: it remains sufficiently structured to support clear benchmarking and broad participation, but it is substantially closer to the longer-term goal of full brain-to-text (B2T) (Herff et al., 2015), which is analogous to speech-to-text but with brain signals as inputs. Whereas speech detection and phoneme classification primarily target lower-level acoustic structure in the signal, word classification brings lexical and semantic information into scope and begins to probe whether models can recover meaningful units of speech from MEG data. As a stepping stone toward a working non-invasive B2T, it is well-motivated: word classification was a key component of the pipeline in the first invasive BCI for a paralysed individual (Moses et al., 2021), and remains a useful intermediate target given the apparent difficulty of non-invasive B2T (Jo et al., 2024) — though we have finally begun to see the first WER scores for non-invasive speech B2T that are significantly better than chance (Jayalath et al., 2025a). Although Brain2Qwerty models report interesting results for decoding sentences from MEG acquired while subjects overtly type on a keyboard (Lévy et al., 2026; Zhang et al., 2026), we view this as distinct from speech BCIs: there is currently no clear path for typing-based decoding to be used by patients unable to move their fingers.

Word classification has recently attracted growing attention in the non-invasive decoding literature (d’Ascoli et al., 2025; Jayalath et al., 2025a; Jayalath and Parker Jones, 2026). Yet without standard evaluation protocols, results remain difficult to compare across studies. A key source of incomparability is the choice of target vocabulary. Two studies that report the same evaluation metric on the most frequent few hundred words in their respective datasets may appear comparable (d’Ascoli et al., 2025; Jayalath et al., 2025a, e.g.,), but this is complicated if the target words differ or if their frequencies vary between datasets (Jayalath et al., 2026). The choice of words also affects what can be communicated. For all of these reasons, it is important to report the target vocabulary explicitly. Official competition rankings are evaluated on a 50-word competition vocabulary designed to leverage high-frequency words in the training corpora and to support expressive communication (Appendix B).

To make it easier to compare competition results with prior results in the literature, we further track performance on a second 50-word vocabulary, the Moses 50 vocabulary, which has been used many times in invasive studies (Moses et al., 2021; Willett et al., 2023, e.g.,). Finally, we report an information-theoretic measure designed to compare results with different target vocabularies: open-vocabulary mutual information (OVMI) (Jayalath et al., 2026). OVMI accounts for both decoding performance and vocabulary coverage (Section 1.5).

1.2 Novelty

The 2025 PNPL competition was the first dedicated to language decoding from non-invasive brain data (Landau et al., 2025). Building on its success, the 2026 competition introduces several elements that are new both to the PNPL series and to the broader field.

Word classification as a standardised competition task. Although word classification has recently attracted attention in the non-invasive decoding literature (d’Ascoli et al., 2025; Jayalath et al., 2025a; Jayalath and Parker Jones, 2026), standardised evaluation is harder here than for tasks with fixed symbol sets such as phoneme classification. Unlike phonemes, words form an effectively open set, and the choice of vocabulary—not merely its size—affects measured performance (Jayalath et al., 2026). Custom vocabularies, such as the most-frequent K words in a corpus (d’Ascoli et al., 2025), maximise training examples and reported accuracy scores. On the other hand, results for custom vocabularies are not comparable across studies and high-frequency words are often dominated by function words of limited communicative utility. Standardised cross-study vocabularies such as the Moses 50 (Moses et al., 2021) enable comparisons (including with invasive BCIs), but may be poorly attested in a given corpus. OVMI (Jayalath et al., 2026) addresses both concerns, jointly accounting for decoding accuracy and corpus coverage. Because they achieve different goals, we use all three methods of evaluation (Section 1.5) — a practice we recommend to the field. The primary metric for both the leaderboard and final ranking is the custom, dataset-tailored top-10 balanced accuracy.

Cross-subject generalisation as a competition target. The 2025 competition and the LibriBrain dataset focused exclusively on a single participant. Notably, all recent state-of-the-art invasive systems have likewise been trained and evaluated on a single individual (Moses et al., 2021; Metzger et al., 2023; Willett et al., 2023; Card et al., 2024), meaning cross-subject generalisation has received little attention in the field as a whole. The 2026 competition is the first to explicitly target it, introducing a dedicated Broad track that evaluates how well models can adapt to new individuals with progressively limited data. This mirrors the constraints of clinical deployment and provides a novel evaluation paradigm for data-efficient generalisation. Uniquely, the Broad track also includes holdout subjects for whom no subject-specific training data is released at all, enabling evaluation of zero-shot cross-subject generalisation — the most clinically relevant regime.

Contrast with invasive brain-to-text competitions.

Building on high-performance but invasive speech neuroprosthesis systems (Willett et al., 2023; Card et al., 2024), the Brain-to-Text Benchmark ’24 and its successor competition target the decoding of full transcripts from intracortical recordings (Willett et al., 2024). Our goal is to build the corresponding non-invasive benchmark for MEG. Since open-vocabulary B2T from speech-related neural activity remains out of reach for current non-invasive systems (Jayalath et al., 2025a, though see), we instead take a curricular approach and focus on word classification. This gives the community a concrete intermediate task on which to measure progress in representation learning, cross-subject generalisation, and decoding under the constraints of non-invasive neural data.

LibriBrain100 as competition resource. The competition is the first to be built around a large-scale, multi-subject MEG dataset. LibriBrain100 (Mantegna et al., 2026) combines the deepest within-subject MEG dataset recorded to date (∼80 hours) with data from 32 additional subjects, enabling both within-subject and cross-subject benchmarking within a single competition.

1.3 Data

For this year’s competition, we primarily use the LibriBrain100 dataset (Mantegna et al., 2026), which comprises non-invasive MEG recordings from 33 subjects listening to naturalistic speech (stimuli details in Appendix C). Subject 0—the single participant from the original LibriBrain dataset (Özdogan et al., 2025)—is extended to ∼80 hours of recordings, the deepest within-subject MEG dataset recorded to date. For an additional 32 subjects, ∼40 minutes of MEG data were collected, but for the competition we have initially released tiered amounts: the full ∼40 minutes for 12 subjects, ∼20 minutes for 10, and ∼10 minutes for a further 10—a regime within the range of what will often be clinically feasible (Figure 1). The motivation for this is to challenge competition participants to predict with decreasing amounts of subject-specific data. For a final group of 8 subjects not included in LibriBrain100 (subjects 33–40), there are no subject-specific training data, requiring zero-shot cross-subject generalisation: the most demanding but clinically useful regime. We will release the full data for subjects 1–32 after the competition.

The LibriBrain100 data come with standard train, validation, and test splits for reproducibility. Evaluation in the competition is conducted exclusively on an additional competition holdout set, which is drawn from sources not revealed to participants in advance. For the competition, we release the MEG recordings for this holdout set but withhold the labels. Submissions consist of predictions on these recordings, uploaded to our evaluation platform. The holdout is divided into two disjoint partitions: one used to update the public leaderboard throughout the competition, and one reserved for final ranking of submissions. This design mirrors the 2025 competition and is intended to prevent overfitting to the holdout distribution while still giving participants timely feedback on their progress.

MEG data and labels are saved in HDF5 and TSV formats, respectively. The data are also available on Hugging Face in both serialised (HDF5) and raw (FIF) formats, across two repositories: libribrain contains the original LibriBrain recordings, while libribrain2 contains the additional data comprising LibriBrain100. Note that raw FIF files have not been preprocessed and include head movement artefacts and other noise; they are also substantially larger than the serialised versions.

To make interacting with the dataset easy, we provide an updated Python library that automatically downloads and loads data for PyTorch. It can be installed with a single command: pip install pnpl. We recommend using the library to load serialised data, as it handles downloads automatically and retrieves only the partitions requested (Appendix D); no knowledge of underlying repo structures is required.

1.4 Tasks and applications

The task in this competition is word classification from non-invasive brain recordings. Given a window of MEG data x∈ℝ306×T (sensors × time samples), the goal is to predict the word y∈𝒱 being heard by the participant, where 𝒱 is a fixed retrieval vocabulary (Mantegna et al., 2026, see). We use a custom 50-word vocabulary designed for reliable evaluation on LibriBrain100 (see Appendix B). Relevant neural information for this task may include both phonetic representations (reflecting acoustic and articulatory processing in auditory and motor cortices) and lexical semantic representations which cover most of the cortex (Huth et al., 2016). Given the distributed nature of semantic representations in particular, MEG’s whole-brain coverage is a distinct advantage over surgically implanted arrays, which sample only a limited (surgically accessed) cortical region. This makes word classification a richer decoding target than either speech detection or phoneme classification alone.

Word classification occupies an important position in the curriculum of non-invasive speech decoding. It is more structured than open-vocabulary brain-to-text, which has only recently yielded word error rates better than chance from non-invasive recordings (Jayalath et al., 2025a), making it a reliable and reproducible benchmark. At the same time, it goes substantially beyond the tasks of the 2025 competition (Landau et al., 2025): speech detection and phoneme classification target sub-lexical structure, but word classification directly probes the recovery of meaningful linguistic units. It is also practically motivated. Word classification, combined with speech detection and a language model, formed the core pipeline of the first invasive speech BCI for a paralysed individual (Moses et al., 2021). Solving word classification from non-invasive data, together with the speech detection foundation laid in 2025, would therefore constitute a meaningful step toward a non-invasive analogue.

The competition hosts two complementary word classification tasks that run concurrently:

(a) D​e​e​p track: This track evaluates word classification in a single, densely-sampled subject. Participants may train on any data they wish. The goal is the best possible decoding performance on the densely-sampled individual (Subject 0), to push the limits of non-invasive word classification.

(b) B​r​o​a​d track: This track evaluates cross-subject generalisation. The challenge is to produce accurate word classification predictions for individuals with limited subject-specific data: for 12 subjects, ∼40 minutes of subject-specific data is available; for 10 subjects, ∼20 minutes; and for a further 10 subjects, ∼10 minutes. The ∼10 minute regime is of particular interest for clinical applications, where extended patient data collection is often infeasible. Holdout data is provided for a final set of 8 subjects with no released subject-specific training data, enabling evaluation of zero-shot generalisation. To perform well on this track, submissions must generalise across all data regimes—motivating the emphasis on data-efficiency in the competition title.

1.5 Metrics

To measure success in the word classification tasks we use the Top-10 Balanced Accuracy with a fixed retrieval vocabulary of K=50:

BAcc​@​10=1K​∑k=1KRecall​@​10k.

This metric has been used many times for decoding words in the recent decoding literature, with retrieval vocabularies ranging from K=50 to K=250 (d’Ascoli et al., 2025; Jayalath and Parker Jones, 2026). For the competition, we track both a custom 50-word vocabulary and the Moses 50 vocabulary (Appendix B). Recall​@​10k denotes the Top-10 recall for class k:

Recall@10k=1Nk∑i:yi=k𝕀[yi∈{y^i,1,…,y^i,10}],

where yi is the true label, y^i,r the class with the r-th highest predicted score, and Nk the number of evaluation examples with true label k. We use a balanced metric so that each word contributes equally, regardless of how frequently it appears in the evaluation set (Thölke et al., 2023, see, e.g.,). In practice, BAcc​@​10 scores range from 0 to 1 and can be reported as percentages. For a vocabulary of K=50 words, uniform random guessing without replacement yields a BAcc​@​10 of 20% in expectation, since the correct class has probability 10/50 of appearing in a random top-10 set. A score of 100% means that the correct word is always included in the model’s top-10 predictions. A limitation of this metric is that it does not distinguish between cases where the correct word is ranked first or tenth; therefore, we also compute the Top-1 Balanced Accuracy (BAcc​@​1) to help break ties.

Finally, as an auxiliary metric, we report Open-Vocabulary Mutual Information (OVMI) (Jayalath et al., 2026). OVMI is designed to support comparison across decoders with different target vocabularies. It is an information-theoretic quantity that measures the mutual information between a user’s intent and the output of a decoding model relative to a reference communication distribution. Concretely, OVMI⁡(S)=C⁡(S)​I​(X;Y∣X∈S), where C⁡(S) is the lexical coverage of the retrieval vocabulary S under a reference distribution p and I⁡(X;Y∣X∈S) is the in-vocabulary MI between the intended word X and the decoder output Y. Conceptually, p acts like a prior over the words a user is likely to intend. For the competition, we use SUBTLEX-UK as p (van Heuven et al., 2014) and report OVMI in bits per word perceived under this distribution. OVMI is 0 under uniform random guessing; higher OVMI scores are better.

1.6 Baselines, code, and material provided

To provide participants with baseline scores that are representative of the state-of-the-art and straightforward to implement and improve upon, we provide two reference models—one for each track.

The Deep track uses an implementation of the supervised word decoding model of d’Ascoli et al. (2025) as reference. This model achieves strong within-subject word classification performance and provides a competitive baseline for participants targeting the best possible performance on Subject 0.

The Broad track uses MEG-XL (Jayalath and Parker Jones, 2026) as a reference model. MEG-XL is a self-supervised foundation model pre-trained on ∼300 hours of MEG from 800 subjects, intended to be fine-tuned for word classification. Crucially, it is designed to generalise well with small amounts of subject-specific fine-tuning data, where it outperforms the d’Ascoli et al. (2025) model, making it particularly well-suited to the data-efficient cross-subject generalisation setting of the Broad track. Code and checkpoints are available through GitHub11 1 https://github.com/neural-processing-lab/MEG-XL, providing participants with a competitive starting point that they can directly improve upon. The choice of reference models reflects recent advances in the field: with large amounts of within-subject data, supervised models outperform self-supervised ones (Jayalath and Parker Jones, 2026), while with limited per-subject data, pre-trained priors of self-supervised models like MEG-XL confer a substantial advantage (Mantegna et al., 2026).

For our baseline results, we train or fine-tune each reference model following the procedure described in their respective papers and report results on the competition test splits (Table 1). For the Deep track, we use all the available training data for Subject 0. For the Broad track, across Subjects 1–32, we use half of the data from the Sherlock 1 Session 11 recording for fine-tuning and the other half for validation. We suspect joint training on Subject 0 and Subjects 1–32 will improve generalisation on the latter subjects further, but leave this and other avenues of exploration to competition participants.

媒体内容 · 前往原文查看

Competition Vocabulary Moses Vocabulary

Method BAcc​@​1 BAcc​@​10 OVMI BAcc​@​1 BAcc​@​10 OVMI

Deep track

d’Ascoli (reference) 25.60% 73.23% .220 38.84% 83.60% .157

Random chance 2.00% 20.00% .000 2.00% 20.00% .000

Broad track

MEG-XL (reference) 6.20% 42.96% .014 10.21% 52.23% .011

d’Ascoli 5.14% 33.13% .009 6.10% 49.56% .006

Random chance 2.00% 20.00% .000 2.00% 20.00% .000

Table 1: Scores for reference models and random baselines. The column in (grey) highlights the competition evaluation metric (BAcc​@​10 on the competition vocabulary). All scores are on each track’s test data from the checkpoint with the best validation performance on the competition metric.

As an entry point to the library and competition, we provide three (3) tutorial notebooks in the form of interactive Google Colab notebooks, optimised for usage with a free GPU:

First tutorial: Data loading, introduction to LibriBrain100 and word classification task. Contains code for fine-tuning the reference model on Subject 0.

Second tutorial: Loading subject-specific data, fine-tuning the reference model for cross-subject generalisation, more details on the Broad track.

Third tutorial: Prediction generation on the holdout data, submission to the leaderboard.

1.7 Competition website

The competition website22 2 https://libribrain.com/editions/2026/ serves as a central hub, aggregating the starter kit, documentation, tutorial notebooks, leaderboards, and links to the Kaggle submission platform that accepts prediction files in CSV format. To maximise ease of onboarding, all tutorials are provided as interactive Colab notebooks that run in the browser, requiring no local installation beyond: pip install pnpl.

2 Organisation

2.1 Protocol

Participating is designed to be simple and accessible. To participate, one can install the pnpl Python library (pip install pnpl), which automatically downloads the dataset and provides a standard PyTorch DataLoader. Once data is loaded, participants can develop solutions and generate predictions on the holdout data. The pnpl library can be used to write these predictions to a CSV file, which can then be uploaded to the submission platform (Appendix D). Solutions are evaluated automatically according to the metrics in Section 1.5, and rankings updated continuously on the public leaderboard.

Both the D​e​e​p and B​r​o​a​d tracks run concurrently for three months (15 July – 15 October). Participants can submit to either or both tracks with progress reflected on the appropriate leaderboard.

2.2 Rules and engagement

Goal. The competition aims to be open and accessible to a broad community of researchers. This includes participants that may have no prior experience in analysing neural data. Our rules are intended to encourage open participation, while protecting integrity of the evaluation process.

Competition rules.

Participation is open to everyone. Domain-specific knowledge or specialised hardware is not required for fully participating in the competition.

Competition tracks. The competition is organised around the LibriBrain100 dataset and consists of two distinct tracks, both targeting word classification using balanced accuracy over a fixed vocabulary: (a) D​e​e​p track: word classification within a single, deeply-sampled subject, and (b) B​r​o​a​d track: word classification across many held-out subjects.

Note, the two tracks are ranked separately, with (i) a leaderboard and (ii) prizes separately tracked for each. Participants may submit to either or both tracks.

Permitted materials. Participants may use any training data for either track. This includes (i) external, publicly available datasets, (ii) self-collected or synthetically-generated data, or (iii) pre-trained models available from public repositories on Hugging Face (https://huggingface.co) or GitHub (https://github.com/). Participants must independently ensure their data use complies with the appropriate licensing and ethical requirements.

Winners. Final rankings are determined against an independent holdout dataset to prevent data leakage. We discourage any attempt to infer held-out labels outside of the official evaluation procedure. The organisers reserve the right to make final decisions on edge cases.

Prize distribution. The highest-ranked three (3) submissions in each track are to be considered for prizes, provided they beat the performance of reference model baselines described in Section 1.6. Each team is eligible to win at most one prize across the competition. If the same team places in the prize-winning positions on both tracks, they will be listed on all relevant leaderboards, but prizes for the additional track pass to the next eligible team. In the unlikely event of a tie, the prize will be split between the tied teams.

The organisers will contact all prize contenders for additional verification of their submissions. Participants may be asked to share the training code, model checkpoints, or descriptions of their approach. If neither correctness of the approach nor compliance with the competition rules can be verified, the submission may be excluded from prize eligibility.

The organising team encourages participants to share their code in the spirit of open science, including direct contributions (e.g., data loaders, reusable baselines, documentation) to the pnpl GitHub repository (https://github.com/neural-processing-lab/pnpl).

Communication and community. To facilitate open and real-time communication, the competition uses the same Discord server established during last year’s competition33 3 https://libribrain.com/links/discord.

2.3 Schedule and readiness

All essential competition materials are online. For prizes, the equivalent of $5,000 USD has been raised.

15 July, 2026: Submissions to both tracks officially open. All necessary materials are released, including documentation and unlabelled evaluation data for each track.

15 July – 15 October, 2026: Competition progress is tracked on the public leaderboards. The organising team provides support through the competition website, GitHub, and public Discord server. Participants may submit to either track during the competition period.

15 October, 2026: Submissions to both tracks officially close.

15 October – 30 November, 2026: The organising team reviews and verifies the highest-ranking submissions. Contenders for all prizes will be contacted to further confirm submission details.

1 December, 2026: The confirmed winners are announced through the competition website. The organising team independently conducts post-competition analysis of the results. Winners will be connected with sponsors to coordinate prizes.

2.4 Competition promotion and incentives

We are promoting the competition through social networks and academic mailing lists (e.g., NeurIPS affinity groups). We launched a dedicated website to share information, resources, and updates regarding the competition. This includes live leaderboards. During the competition, the organising team plans to release blog posts about the tasks and data to foster discussion.

3 Resources

The pnpl Python library is available to download the data and integrate it seamlessly into standard deep learning frameworks. Reference model code and checkpoints are publicly available on GitHub. The tutorial code runs on Google Colab in the browser, providing some GPU access to get started.

Organising team

Francesco Mantegna is a Postdoctoral Researcher in PNPL, Department of Engineering Science, University of Oxford. He received his PhD from NYU under the supervision of David Poeppel.

Gereon Elvers is an Encode: AI for Science Fellow in PNPL, University of Oxford. He was previously a master’s student, visiting PNPL from the Technical University of Munich (TUM).

Dulhan Jayalath is a DPhil student in Machine Learning at the University of Oxford, supervised by Oiwi Parker Jones as part of PNPL and funded by an Amazon Web Services (AWS) studentship as part of the EPSRC Centre for Doctoral Training in Autonomous Intelligent Machines and Systems (AIMS).

Gilad Landau is a DPhil student in Engineering and member of PNPL, supervised by Oiwi Parker Jones at the University of Oxford.

Tasha Kim is a DPhil student in Engineering and member of PNPL, supervised by Oiwi Parker Jones at the University of Oxford.

Miran Özdogan is a DPhil student in Computer Science and member of PNPL, supervised by Oiwi Parker Jones and Michael Bronstein, at the University of Oxford.

Luisa Kurth is a DPhil student in PNPL and the EPSRC Centre for Doctoral Training in Autonomous Intelligent Machines and Systems (AIMS), University of Oxford.

Teyun Kwon is a DPhil student in PNPL and the EPSRC Centre for Doctoral Training in Autonomous Intelligent Machines and Systems (AIMS), University of Oxford.

SungJun Cho is a DPhil student in Neuroscience supervised by Mark Woolrich and Oiwi Parker Jones at the University of Oxford.

Benjamin Ballyk is a DPhil student in Engineering and member of PNPL, supervised by Oiwi Parker Jones at the University of Oxford.

Alex Fung is a DPhil student in Neuroscience and member of PNPL, supervised by Alex Green, Saad Jbabdi, and Oiwi Parker Jones at the University of Oxford.

Anna Greer is a DPhil student in the EPSRC Centre for Doctoral Training in Autonomous Intelligent Machines and Systems (AIMS), University of Oxford. She contributed during a rotation mini-project in PNPL.

Pratik Somaiya is a Software Engineer at the Oxford Robotics Institute.

Christian Herff is an Associate Professor and head of the Neural Interfacing Lab in the Department of Neurosurgery at Maastricht University.

Yorguin Mantilla Ramos is a master’s student at the Université de Montréal and Graduate Research Assistant at Mila.

Hamza Abdelhedi is a Biomedical Engineering PhD student at the Université de Montréal, supervised by Karim Jerbi and specialising in Neuro-AI.

Karim Jerbi is Professor in the Psychology Department of the Université de Montréal and Associate Professor at Mila. He heads UNIQUE, the Quebec-wide Neuro-AI research center, and is also Canada Research Chair in Computational Neuroscience and Cognitive Neuroimaging.

Greg Farquhar is a Staff Research Scientist at Google DeepMind.

Brendan Shillingford is a Staff Research Scientist at Google DeepMind.

Mark Woolrich is Professor of Computational Neuroscience at the University of Oxford. He is Head of Analysis and Associate Director of the Oxford Centre for Human Brain Activity (OHBA).

Oiwi Parker Jones heads the Parker Jones Neural Processing Lab (PNPL) in the Department of Engineering Science, University of Oxford. He is also a Fellow in Computer Science at Jesus College Oxford, a Principal Investigator at the Oxford Robotics Institute, and an Honorary Fellow in the Nuffield Department of Clinical Neurosciences.

Acknowledgments

Many thanks to Pillar VC for sponsoring this year’s competition. We gratefully acknowledge the use of the University of Oxford Advanced Research Computing (ARC) facility (http://dx.doi.org/10.5281/zenodo.22558), Hartree Centre resources, and the NVIDIA Corporation for donating additional GPUs. PNPL is supported by the MRC (MR/X00757X/1), Royal Society (RG\R1\241267), NSF (2314493), NFRF (NFRFT-2022-00241), SSHRC (895-2023-1022), and ARIA (SCNI-SE01-P004).

References

Armeni et al. (2022) K. Armeni, U. Güçlü, M. van Gerven, and J.-M. Schoffelen A 10-hour within-participant magnetoencephalography narrative dataset to test models of language comprehension. Scientific Data 9, pp. 278. Cited by: Appendix A.

Barratt et al. (2018) E. L. Barratt, S. T. Francis, P. G. Morris, and M. J. Brookes Mapping the topological organisation of beta oscillations in motor cortex using MEG. NeuroImage 181, pp. 831–844. External Links: Document, Link Cited by: §1.1.

Card et al. (2024) N. S. Card, M. Wairagkar, C. Iacobacci, X. Hou, T. Singer-Clark, F. R. Willett, E. M. Kunz, C. Fan, M. Vahdati Nia, D. R. Deo, A. Srinivasan, E. Y. Choi, M. F. Glasser, L. R. Hochberg, J. M. Henderson, K. Shahlaie, S. D. Stavisky, and D. M. Brandman An accurate and rapidly calibrating speech neuroprosthesis. New England Journal of Medicine 391 (7), pp. 609–618. Cited by: §1.1, §1.2, §1.2.

Défossez et al. (2023) A. Défossez, C. Caucheteux, J. Rapin, O. Kabeli, and J. King Decoding speech perception from non-invasive brain recordings. Nature Machine Intelligence 5 (10), pp. 1097–1107. External Links: Document, Link Cited by: §1.1.

Doyle (1888) A. C. Doyle A study in scarlet. Ward, Lock & Co. Cited by: Appendix C.

Doyle (1890) A. C. Doyle The sign of the four. Spencer Blackett. Cited by: Appendix C.

Doyle (1892) A. C. Doyle The adventures of sherlock holmes. George Newnes. Cited by: Appendix C.

Doyle (1893) A. C. Doyle The memoirs of sherlock holmes. George Newnes. Cited by: Appendix C.

Doyle (1902) A. C. Doyle The hound of the baskervilles. George Newnes. Cited by: Appendix C.

Doyle (1905) A. C. Doyle The return of sherlock holmes. George Newnes. Cited by: Appendix C.

Doyle (1915) A. C. Doyle The valley of fear. George H. Doran Company. Cited by: Appendix C.

Doyle (1917) A. C. Doyle His last bow. John Murray. Cited by: Appendix C.

Doyle (1927) A. C. Doyle The case-book of sherlock holmes. John Murray. Cited by: Appendix C.

d’Ascoli et al. (2025) S. d’Ascoli, C. Bel, J. Rapin, H. Banville, Y. Benchetrit, C. Pallier, and J. King Towards decoding individual words from non-invasive brain recordings. Nature Communications 16, pp. 10521. External Links: Document, Link Cited by: Appendix A, Appendix B, §1.1, §1.1, §1.2, §1.5, §1.6, §1.6.

Elvers et al. (2026) G. Elvers, G. Landau, F. Mantegna, M. Özdogan, T. Kim, T. Kwon, S. Cho, B. Ballyk, L. Kurth, D. Jayalath, P. Somaiya, X. de Zuazo, B. Shillingford, G. Farquhar, M. Jiang, K. Jerbi, H. Abdelhedi, Y. Mantilla Ramos, C. Gulcehre, M. Woolrich, N. Voets, and O. Parker Jones Benchmarking non-invasive speech BCIs: lessons learned from the 2025 PNPL competition. Journal of Machine Learning Research. Note: In press Cited by: §1.1, Abstract.

Garofolo et al. (1993) J. S. Garofolo, L. F. Lamel, W. M. Fisher, J. G. Fiscus, D. S. Pallett, N. L. Dahlgren, and V. Zue TIMIT acoustic-phonetic continuous speech corpus. Note: Linguistic Data Consortiumhttps://catalog.ldc.upenn.edu/LDC93S1 Cited by: Appendix C.

Gwilliams et al. (2023) L. Gwilliams, G. Flick, A. Marantz, L. Pylkkanen, D. Poeppel, and J. King Introducing MEG-MASC: a high-quality magneto-encephalography dataset for evaluating natural speech processing. Scientific Data 10 (1), pp. 862. External Links: Document, Link Cited by: Appendix A.

Hämäläinen et al. (1993) M. Hämäläinen, R. Hari, R. J. Ilmoniemi, J. Knuutila, and O. V. Lounasmaa Magnetoencephalography—theory, instrumentation, and applications to noninvasive studies of the working human brain. Reviews of Modern Physics 65 (2), pp. 413–497. Cited by: §1.1.

Herff et al. (2015) C. Herff, D. Heger, A. de Pesters, D. Telaar, P. Brunner, G. Schalk, and T. Schultz Brain-to-text: decoding spoken phrases from phone representations in the brain. Frontiers in Neuroscience 9 (217), pp. 1–11. Cited by: §1.1.

Huth et al. (2016) A. G. Huth, W. A. de Heer, T. L. Griffiths, F. E. Theunissen, and J. L. Gallant Natural speech reveals the semantic maps that tile human cerebral cortex. Nature 532, pp. 453–458. Cited by: §1.4.

Jayalath et al. (2026) D. Jayalath, B. Ballyk, and O. Parker Jones A common measure of communication for speech brain–computer interfaces. Note: Manuscript in preparation Cited by: §1.1, §1.1, §1.2, §1.5.

Jayalath et al. (2025a) D. Jayalath, G. Landau, and O. Parker Jones Unlocking non-invasive brain-to-text. International Conference on Machine Learning (ICML), Workshop on Generative AI and Biology. Note: arXiv preprint arXiv:2505.13446 Cited by: §1.1, §1.1, §1.2, §1.2, §1.4.

Jayalath et al. (2025b) D. Jayalath, G. Landau, B. Shillingford, M. W. Woolrich, and O. Parker Jones The Brain’s Bitter Lesson: scaling speech decoding with self-supervised learning. International Conference on Machine Learning (ICML). Note: arXiv preprint arXiv:2406.04328 Cited by: §1.1.

Jayalath and Parker Jones (2026) D. Jayalath and O. Parker Jones MEG-XL: data-efficient brain-to-text via long-context pre-training. International Conference on Machine Learning (ICML). Note: arXiv preprint arXiv:2602.02494 Cited by: Appendix B, §1.1, §1.2, §1.5, §1.6.

Jo et al. (2024) H. Jo, Y. Yang, J. Han, Y. Duan, H. Xiong, and W. H. Lee Are EEG-to-text models working?. arXiv preprint. Note: https://arxiv.org/abs/2405.06459 External Links: 2405.06459, Link Cited by: §1.1.

Landau et al. (2025) G. Landau, M. Özdogan, G. Elvers, F. Mantegna, P. Somaiya, D. Jayalath, L. Kurth, T. Kwon, B. Shillingford, G. Farquhar, M. Jiang, K. Jerbi, H. Abdelhedi, Y. Mantilla Ramos, C. Gulcehre, M. Woolrich, N. Voets, and O. Parker Jones The 2025 PNPL competition: speech detection and phoneme classification in the LibriBrain dataset. NeurIPS, Competition Track. Note: arXiv preprint arXiv:2506.10165 Cited by: §1.1, §1.2, §1.4, Abstract.

Lévy et al. (2026) J. Lévy, M. Zhang, S. Pinet, J. Rapin, H. Banville, S. d’Ascoli, and J. King Noninvasive decoding of typed sentences from human brain activity. Nature Neuroscience. External Links: Document Cited by: §1.1.

Mantegna et al. (2026) F. Mantegna, D. Jayalath, G. Elvers, T. Kim, B. Ballyk, A. Fung, S. Cho, T. Kwon, L. Kurth, M. Özdogan, G. Landau, P. Somaiya, N. Voets, M. Woolrich, and O. Parker Jones LibriBrain100: one hundred hours of broad and deep MEG data for neural speech decoding at scale. arXiv preprint arXiv:2608.25204. Cited by: Appendix A, Appendix C, §1.1, §1.2, §1.3, §1.4, §1.6, Abstract.

Metzger et al. (2023) S. L. Metzger, K. T. Littlejohn, A. B. Silva, D. A. Moses, M. P. Seaton, R. Wang, M. E. Dougherty, J. R. Liu, P. Wu, M. A. Berger, I. Zhuravleva, A. Tu-Chan, K. Ganguly, G. K. Anumanchipalli, and E. F. Chang A high-performance neuroprosthesis for speech decoding and avatar control. Nature 620, pp. 1037–1046. Cited by: §1.2.

Moses et al. (2021) D. A. Moses, S. L. Metzger, J. R. Liu, G. K. Anumanchipalli, J. G. Makin, P. F. Sun, J. Chartier, M. E. Dougherty, P. M. Liu, G. M. Abrams, A. Tu-Chan, K. Ganguly, and E. F. Chang Neuroprosthesis for decoding speech in a paralyzed person with anarthria. New England Journal of Medicine 385 (3), pp. 217–227. External Links: Document, Link Cited by: Table 2, Table 2, Appendix B, §1.1, §1.1, §1.1, §1.2, §1.2, §1.4.

Özdogan et al. (2025) M. Özdogan, G. Landau, G. Elvers, D. Jayalath, P. Somaiya, F. Mantegna, M. Woolrich, and O. Parker Jones LibriBrain: over 50 hours of within-subject MEG to improve speech decoding methods at scale. NeurIPS, Datasets and Benchmarks Track. Cited by: Appendix A, Appendix C, §1.1, §1.3, Abstract.

Schoffelen et al. (2019) J. Schoffelen, R. Oostenveld, N. H. L. Lam, J. Uddén, A. Hultén, and P. Hagoort A 204-subject multimodal neuroimaging dataset to study language processing. Scientific Data 6 (17), pp. 1–13. Cited by: Appendix A.

Tang et al. (2023) J. Tang, A. LeBel, S. Jain, and A. G. Huth Semantic reconstruction of continuous language from non-invasive brain recordings. Nature Neuroscience 26 (5), pp. 858–866. External Links: Document, Link Cited by: Appendix C.

Thölke et al. (2023) P. Thölke, Y. Mantilla-Ramos, H. Abdelhedi, C. Maschke, A. Dehgan, Y. Harel, A. Kemtur, L. Mekki Berrada, M. Sahraoui, T. Young, A. Bellemare Pépin, C. El Khantour, M. Landry, A. Pascarella, V. Hadid, E. Combrisson, J. O’Byrne, and K. Jerbi Class imbalance should not throw you off balance: choosing the right classifiers and performance metrics for brain decoding with imbalanced data. NeuroImage 277, pp. 120253. External Links: Document Cited by: §1.5.

van Heuven et al. (2014) W. J. B. van Heuven, P. Mandera, E. Keuleers, and M. Brysbaert SUBTLEX-UK: a new and improved word frequency database for British English. Quarterly Journal of Experimental Psychology 67, pp. 1176 – 1190. Cited by: §1.5.

Willett et al. (2023) F. R. Willett, E. M. Kunz, C. Fan, D. T. Avansino, G. H. Wilson, E. Y. Choi, F. Kamdar, M. F. Glasser, L. R. Hochberg, S. Druckmann, K. V. Shenoy, and J. M. Henderson A high-performance speech neuroprosthesis. Nature 620, pp. 1031–1036. External Links: Document Cited by: Appendix B, §1.1, §1.1, §1.2, §1.2.

Willett et al. (2024) F. R. Willett, J. Li, T. Le, C. Fan, M. Chen, E. Shlizerman, Y. Chen, X. Zheng, T. S. Okubo, T. Benster, H. D. Lee, M. Kounga, E. K. Buchanan, D. Zoltowski, S. W. Linderman, and J. M. Henderson Brain-to-Text Benchmark ’24: lessons learned. arXiv preprint arXiv:2412.17227. External Links: Link, 2412.17227 Cited by: §1.2.

Wrench (1999) A. Wrench The MOCHA-TIMIT articulatory database. Note: https://www.cstr.ed.ac.uk/research/projects/artic/mocha.htmlCreated November 1999; accessed 2026-02-07 Cited by: Appendix C.

Xiong et al. (2016) W. Xiong, J. Droppo, X. Huang, F. Seide, M. Seltzer, A. Stolcke, D. Yu, and G. Zweig Achieving human parity in conversational speech recognition. arXiv preprint arXiv:1610.05256. External Links: Link Cited by: §1.1.

Zhang et al. (2026) M. Zhang, J. Lévy, C. Rommel, J. Rapin, C. Bel, J. Bonnaire, D. Nieto, P. Bourdillon, S. Pinet, S. d’Ascoli, T. Moreau, and J. King Accurate decoding of natural sentences from non-invasive brain recordings. arXiv preprint arXiv:2608.18114. Cited by: §1.1.

Appendices / Supplemental Materials

Appendix A Dataset Comparison

Figure 2 compares LibriBrain100 against other public MEG speech datasets along three dimensions: cohort breadth, single-subject depth, and total recording time. Each point corresponds to one dataset, with the x-axis showing the number of subjects and the y-axis showing the maximum number of recording hours available for any single subject. Bubble area is proportional to the total number of recording hours in the dataset. The arrow highlights the step from LibriBrain to LibriBrain100: LibriBrain100 substantially increases the total recording time and number of subjects while preserving the unusually deep single-subject regime that made LibriBrain useful for within-subject speech decoding. The public MEG speech datasets compared to LibriBrain100 (Mantegna et al., 2026) are LibriBrain (Özdogan et al., 2025), Armeni (Armeni et al., 2022), Le Petit Prince (d’Ascoli et al., 2025), MEG-MASC (Gwilliams et al., 2023), and MOUS (Schoffelen et al., 2019).

媒体内容 · 前往原文查看

Figure 2: Dataset comparison. Comparison of public MEG speech datasets by number of subjects and maximum hours per subject. Both axes are logarithmically scaled. Each bubble represents one dataset, with bubble area proportional to total recording hours. LibriBrain100 occupies a distinctive position: it combines a competitive number of subjects with substantially more total recording time and substantially greater single-subject depth than most existing datasets. The arrow indicates the expansion from LibriBrain to LibriBrain100.

Appendix B Retrieval Vocabulary

Table 2 lists the two 50-word retrieval vocabularies used for evaluating word classification in this competition. Both metrics defined in Section 1.5 (top-10 balanced accuracy and OVMI) are computed independently on each vocabulary and reported separately. Only BAcc​@​10 on the competition vocabulary is the primary metric used for competition leaderboards and prize decisions.

媒体内容 · 前往原文查看

(a) Competition vocabulary (50 words) # Word # Word # Word # Word # Word 1 a 11 be 21 him 31 on 41 they 2 all 12 but 22 i 32 our 42 think 3 always 13 can 23 in 33 out 43 this 4 am 14 do 24 is 34 people 44 time 5 an 15 for 25 it 35 really 45 to 6 and 16 good 26 it’s 36 she 46 very 7 any 17 had 27 my 37 so 47 was 8 are 18 has 28 new 38 that 48 we 9 as 19 have 29 not 39 the 49 were 10 at 20 he 30 of 40 there 50 will

(b) Moses 50 vocabulary (Moses et al., 2021) # Word # Word # Word # Word # Word 1 am 11 faith 21 here 31 need 41 that 2 are 12 family 22 hope 32 no 42 they 3 bad 13 feel 23 how 33 not 43 thirsty 4 bring 14 glasses 24 hungry 34 nurse 44 tired 5 clean 15 going 25 i 35 okay 45 up 6 closer 16 good 26 is 36 outside 46 very 7 comfortable 17 goodbye 27 it 37 please 47 what 8 coming 18 have 28 like 38 right 48 where 9 computer 19 hello 29 music 39 success 49 yes 10 do 20 help 30 my 40 tell 50 you

Table 2: The two 50-word retrieval vocabularies used in the word-classification task, listed alphabetically with index values for reference. (a) The competition vocabulary, selected from words that have sufficient occurrences for low-variance evaluation. (b) The Moses 50 vocabulary (Moses et al., 2021), developed for assistive communication and used in invasive decoding studies.

Competition vocabulary. The first vocabulary is a 50-word set selected to achieve low-variance performance estimation. Two criteria guided selection. First, words were required to occur frequently enough to provide sufficient examples for both training and held-out evaluation. Second, the set was constructed to span multiple grammatical categories—function words (determiners, pronouns, auxiliaries, prepositions, conjunctions) together with common content words (e.g. time, good, think, people, new)—so that the vocabulary supports the composition of practically meaningful utterances. This design preserves continuity with recent non-invasive word-decoding work, where similar high-frequency vocabularies are standard (d’Ascoli et al., 2025; Jayalath and Parker Jones, 2026).

Moses 50. The second vocabulary is the 50-word list of Moses et al. (2021), developed with patient input for assistive communication and subsequently adopted in invasive decoding studies (Willett et al., 2023). In contrast to the competition vocabulary, Moses 50 emphasises content words tied to clinical and daily-living needs (hungry, thirsty, tired, glasses, nurse, music) and social or affective expressions (hello, goodbye, please, hope). Reporting results on Moses 50 enables direct comparison between non-invasive systems evaluated here and invasive systems evaluated against the same vocabulary for the first time. But even with the size of the LibriBrain100 corpus, not all of the Moses 50 words are adequately attested in it. Evaluating on the Moses 50 data alone would underrepresent what is possible and the words in it are not the only words we want to decode.

Despite their modest size, both vocabularies support a useful range of short utterances. From Moses 50: assistive expressions such as “i feel tired”, “please bring my glasses”, “i am not comfortable”, and “tell my family”. From the competition vocabulary: short conversational forms such as “i think so”, “we have time”, “i am very good”, “i am not good at all”, “can i have this”, and “we can do this”. The two vocabularies share 13 words, so their union covers 87 distinct word types.

Appendix C Stimulus Materials

Subject 0 listened to the complete Sherlock Holmes canon (Doyle, 1888; Doyle, 1890; Doyle, 1892; Doyle, 1893; Doyle, 1902; Doyle, 1905; Doyle, 1915; Doyle, 1917; Doyle, 1927), the TIMIT acoustic-phonetic corpus (Garofolo et al., 1993), the MOCHA-TIMIT articulatory database (Wrench, 1999), and 30 podcast stories from The Moth, a subset of those used in Tang et al. (2023). Subjects 1–32 listened to the validation and test recordings from LibriBrain (Özdogan et al., 2025). Further details for stimuli, recording protocols, and preprocessing can be found in the LibriBrain100 dataset paper (Mantegna et al., 2026).

Appendix D Code Samples

The recommended way to participate in the competition is through our custom Python library, pnpl (pip install pnpl), which integrates ready-to-use PyTorch dataloaders, including automatic dataset downloads from Hugging Face, and task configurations:

pnpl.datasets

import

LibriBrain100

pnpl.tasks

import

WordClassification

ds

LibriBrain100(

data_path

"./data/LibriBrain100"

#datadownloadedifmissingfrompath

WordClassification(tmin

,tmax

partition

"train"

x,y

ds[

#x:(channels,time),y:integerwordclassid

In addition, it includes a small utility for submitting predictions to the Kaggle competition.

pnpl.competition

import

write_submission,submit_to_kaggle

write_submission(

"submission.csv"

indices

indices,

#(N,)

primary_probs

primary,

#(N,50)—CompetitionVocabulary

secondary_probs

secondary,

#(N,50)—MosesVocabulary

submit_to_kaggle(

"submission.csv"
