Abstract
Non-invasive speech decoding remains constrained by the low signal-to-noise ratio of neural recordings, which makes fine-grained reconstruction of phonemes or individual words difficult. Motivated by neuroscientific evidence that high-level semantic representations are distributed across cortical regions and evolve over slower temporal scales, we hypothesize that semantic content may provide a more suitable target for non-invasive decoding than low-level acoustic or lexical features. We introduce Brain2Semantics2Text, a method that reconstructs text through an intermediate semantic embedding space. Our model maps sentence-level MEG responses into a semantic manifold and then inverts the predicted embeddings into natural language. This semantic bottleneck enables recovery of high-level meaning without word-level alignment. We describe the core principles of the approach, its implementation, and the strategies used to mitigate the challenges of learning a reliable neural-to-semantic mapping. Finally, we compare against prior non-invasive Brain2Text methods and show improved sentence-level results.
1 Introduction
Speech decoding BCIs have long been a sought-after goal in both neuroscience and healthcare. Recent progress in invasive Brain2Text systems has demonstrated the feasibility of translating neural activity into language (Anumanchipalli et al., 2019; Moses et al., 2021; Willett et al., 2023; Card et al., 2024). However, these invasive approaches require surgical implantation of intracranial electrodes, creating a strong incentive to develop non-invasive speech decoding solutions that are safer, more accessible, and easier to deploy.
Extending this paradigm to non-invasive systems remains a major challenge, largely due to their inherently lower signal-to-noise ratios. Lower signal fidelity makes it difficult to reliably decode fine-grained linguistic units such as phonemes or individual words. In contrast, higher-order contextual semantic representations are spatially distributed across the cortex (Huth et al., 2016), exhibit substantial redundancy, and evolve over slower temporal scales (Gwilliams and others, 2025). These properties may make semantic representations particularly amenable to decoding from non-invasive neural recordings, whose spatial and temporal characteristics are better matched to distributed, slowly evolving signals.
Recently, a growing body of work in both the speech decoding literature and the neuroscience of language has converged on the view that speech comprehension and production rely on a hierarchical organization of neural representations (Gwilliams and others, 2025). Lower levels of this hierarchy are dominated by auditory and articulatory representations closely tied to the acoustic structure of speech, while progressively higher levels abstract away from these surface properties, giving rise to representations that are less dependent on specific phonetic or lexical features and more closely associated with the meaning of speech (Gwilliams et al. (2025), Goldstein et al. (2025)).
Within this framework, the success of invasive approaches can be largely attributed to their ability to directly access high-fidelity neural signals from the auditory and articulatory components of speech processing, which occupy the lower levels of the cortical hierarchy (often supplemented by post-hoc language models that guide generation; (Willett et al., 2024)). The same hierarchical view also motivates directly targeting higher-level semantic representations as a distinct and parallel neural signal. If such representations can be reliably decoded, their defining properties—slow temporal dynamics, distributed cortical organization, and representational redundancy—are better matched to the spatial and temporal characteristics of non-invasive modalities such as fMRI, MEG, and EEG.
This work targets semantic representations in brain activity, focusing on higher-level stages of the speech-processing hierarchy. Brain2Semantics2Text maps sentence-length MEG responses during heard speech into a pretrained semantic embedding space, from which text is subsequently reconstructed. By constraining neural decoding to pass through this semantic bottleneck, the method aims to shift the objective toward higher-level speech representations that carry information about sentence-level meaning.
Several aspects distinguish our method from previous work on Brain2Text:
Semantic embedding inversion. We build on recent advances in semantic embedding inversion (Morris et al., 2023), which enables the reconstruction of text from semantic embeddings either as an intrinsic property of newer embedding models (Duquenne et al., 2023) or via general inversion techniques applicable to arbitrary pre-trained semantic spaces (Jha et al., 2025). This allows us to frame speech decoding as semantic reconstruction rather than word or phoneme prediction.
Sentence-level semantic decoding. Our method operates directly at the sentence level, targeting compositional semantic representations near the top of the speech-processing hierarchy. This formulation shifts the decoding problem away from exact word or phoneme recovery and toward reconstruction of the intended semantic content. As a result, it avoids dependence on precise word-level alignment and closed-vocabulary supervision, both of which are difficult to assume in realistic settings. Although full-sentence decoding from MEG is ambitious, sentence-level semantic decoding is well matched to the distributed and temporally extended nature of high-level language representations, making it a promising direction for future non-invasive communication systems.
Semantic decoding from MEG. Unlike most prior semantic decoding work, which relies on the high spatial resolution of fMRI (Tang et al., 2023), we leverage the largest single-subject heard-speech MEG dataset of its kind to date. Decoding semantic-level information from MEG enables the joint exploitation of slow, distributed semantic signals and local, high-frequency neural activity within the same recordings, opening new possibilities for improved speech decoding.
2 Related Work
Semantic decoding.
Semantic representations have previously been used as an intermediate target for non-invasive language decoding. Pereira et al. (2018) demonstrated that fMRI responses to sentences could be mapped into semantic embedding spaces, providing early evidence that distributed neural activity can be aligned with sentence-level meaning. More recently, Tang et al. (2023) reconstructed continuous perceived and imagined language from fMRI by mapping distributed cortical responses into semantic representations and using these representations to constrain language generation. These approaches exploit the high spatial resolution of fMRI to recover distributed semantic information, but sacrifice the temporal resolution available in electrophysiological recordings.
Wang et al. (2023) extended semantic reconstruction to MEG, demonstrating that semantic information can also be recovered from temporally resolved non-invasive recordings. Their approach, however, operates at the word level, reconstructing a temporally aligned sequence of contextual word embeddings that is subsequently used to generate continuous text. In contrast, our method treats the entire sentence as the unit of decoding, mapping sentence-length MEG responses directly into a single pre-trained sentence-level semantic embedding. This formulation makes the semantic representation itself the decoding bottleneck and removes the need for word-level alignment.
Auditory and speech-based MEG decoding.
A complementary line of work maps MEG activity to representations closely tied to the acoustic structure of speech. Défossez et al. (2023) aligned MEG responses with Wav2Vec representations (Baevski et al., 2020), establishing a contrastive-learning framework and neural architecture that have influenced subsequent MEG decoding systems. More recent approaches align MEG with auditory representations (Yang et al., 2024c), adapt Whisper for neural speech decoding (Yang et al., 2024b), or incorporate neural signals into multimodal foundation-model architectures (Yang et al., 2024a). These approaches exploit neural information associated with the acoustic realization of speech, whereas our method deliberately targets higher-level semantic content.
Within this line of work, BrainECHO (Li et al., 2024) provides the closest comparison to our method. Like Brain2Semantics2Text, BrainECHO operates at the sentence level and does not require word-level alignment, but the two methods differ in the representation through which decoding proceeds. BrainECHO maps neural activity into a vector-quantized audio-spectrogram latent space before generating text with Whisper, whereas our method maps MEG directly into a sentence-level semantic embedding. BrainECHO therefore provides a particularly informative baseline for evaluating the use of semantic, rather than acoustic, representations as a bottleneck for sentence-level decoding.
Word-level classification.
Semantic representations have also been used for word-level MEG decoding. d’Ascoli et al. (2024) achieve strong closed-vocabulary word classification by aligning MEG responses with lexical semantic embeddings augmented by sentence context. Their approach demonstrates the utility of semantic representations for MEG decoding, but relies on exact word-level timing and formulates decoding as classification among candidate words. Our approach instead targets a single compositional representation of the complete sentence, removing the requirement for word-level alignment at the cost of a substantially less constrained reconstruction problem.
3 Method
The Brain2Semantics2Text method operates in two stages. In the first stage, MEG neural responses corresponding to continuously presented spoken sentences are mapped to vector representations in the pre-trained semantic embedding space. Training is guided by objectives that encourage alignment with the target embeddings while preserving their global statistical structure. In the second stage, the predicted semantic embedding is inverted into natural language using a pre-trained inversion model (Morris et al., 2023) that reconstructs text from semantic vectors.
3.1 Backbone
MEG-to-semantic mapping.
The MEG input signal , where denotes the number of sensors and the number of temporal samples, is first processed by a spatial attention module, followed by an initial convolution that projects the sensor dimension into a latent feature space. The resulting representation is then passed through a stack of dilated temporal convolutional blocks. Each block consists of dilated convolutions equipped with residual connections and gated linear units (GLUs), enabling the model to capture long-range temporal dependencies.
A subject-specific layer can optionally be inserted after the initial projection to model inter-subject variability. In the experiments reported here all data originate from a single subject.
Temporal aggregation.
To obtain a fixed-dimensional semantic representation from the time-resolved features, we use a 4-head self-attention Transformer over the temporal axis. We then pool over time with a masked mean
where is a temporal mask based on the true segment lengths, preventing length-related surface confounders from influencing the pooled embedding.
3.2 Semantic Embeddings
Many semantic embedding models are available, with different architectures, training objectives, and benchmark performance (Muennighoff et al., 2022). For Brain2Semantics2Text, however, standard evaluations on semantic similarity or retrieval tasks provide only partial guidance. Our method requires the embedding space to function as an invertible bottleneck between MEG responses and text, which introduces additional constraints beyond general semantic performance. In particular, the embedding must have an available inversion mechanism and must support reliable reconstruction from both exact text embeddings and imperfect neural predictions. We therefore evaluate candidate embedding spaces according to four criteria.
Expressivity.
The embedding space should preserve enough information about the original sentence to support reconstruction. We measure this using a round-trip reconstruction test, in which each sentence is embedded and then inverted back into text. Higher reconstruction quality indicates that more sentence-level information is retained by the embedding and its inversion procedure.
Soft reversibility.
At test time, the vectors being inverted are not exact text embeddings, but embeddings predicted from noisy MEG responses. The embedding space should therefore be robust to prediction error: vectors near the target should still invert to text with similar meaning. We assess this by perturbing target embeddings and calculating the relation between the introduced noise and reconstruction fidelity.
Length bias.
The embedding should encode sentence meaning without being dominated by surface-level properties such as sentence length. This is particularly important because stimulus duration and sentence length may be available to the neural decoder and could provide a shortcut that competes with semantic learning. We estimate length bias by computing the Spearman rank correlation () between sentence length and the leading principal components of each embedding space. Although this measure does not capture all forms of embedding sensitivity to sentence length, we find that it provides an effective empirical proxy for the extent to which sentence length is reflected in the global geometry of the embedding space. A visual illustration of this analysis is provided in Appendix F.
Intrinsic dimensionality.
The target space should be learnable from limited MEG data. We therefore prefer embedding spaces with lower intrinsic dimensionality, measured by effective rank, provided that they remain sufficiently expressive and reversible.
A further consideration is training-data transparency. In principle, a fully open embedding and inversion pipeline would be preferable, since it would allow us to verify that evaluation sentences were not present in the training data of either the embedding model or the inversion model. Among the embedding–inversion pairs we considered, however, we did not find a fully data-transparent option that also satisfied the practical requirements of expressivity and soft reversibility. We therefore control for possible corpus-leakage by comparing neural-based predictions against noise-control.
| Embedding | Expressivity | Soft reversibility | Length bias | Intrinsic dim. |
|---|---|---|---|---|
| SONAR | ||||
| T5 | ||||
| ADA |
3.3 Objectives for Manifold Learning
At the core of our training setup is a SigLIP-style contrastive loss (Zhai et al., 2023; d’Ascoli et al., 2024), which has been shown to be effective for aligning representations across modalities with different dimensionalities and statistical characteristics (Radford et al., 2021). However, in the low-data regime typical of non-invasive speech decoding, contrastive objectives alone are insufficient to learn the target manifold of semantic embeddings.
Previous brain-to-text and word-decoding approaches have largely relied on contrastive objectives to align neural signals with semantic embeddings (Défossez et al., 2023; d’Ascoli et al., 2024). However, we observe that in the low-data regimes typical of non-invasive speech decoding, contrastive loss primarily optimizes a retrieval objective. In this setting, the model learns a mapping that enables nearest-neighbor matching under cosine similarity (Minnema and Herbelot, 2019), but does not necessarily preserve the inter-vector distances or the scale of embedding magnitudes.
This limitation is problematic for our setting, where the goal is not merely to retrieve a correct target embedding, but to learn a mapping that faithfully captures the global geometry of the target semantic manifold. To address this, we draw inspiration from the manifold learning literature (Meilă and Zhang, 2023) and introduce several auxiliary losses in addition to the SigLIP that force the model to learn the global properties of the target manifold and prevent collapse. The resulting training objective consists of the following components (invariance, covariance and variance losses are adopted from VICReg (Bardes et al., 2021)):
SigLIP Loss: A contrastive alignment term that formulates predicted–target matching as independent pairwise classification rather than a batch-wise softmax objective. This is useful for sentence-level semantic decoding, where different non-matching sentences may still be semantically related and should not necessarily be treated as mutually exclusive classes. We also found this objective more stable in the low-data MEG setting, particularly with small batches.
| (1) |
Invariance Loss : mean squared distance between predicted and target embeddings.
| (2) |
Covariance Loss: a decorrelation term that penalizes off-diagonal covariances between embedding dimensions, reducing redundancy and preventing informational collapse.
| (3) |
| (4) |
Variance Loss: a hinge loss that enforces a minimum standard deviation across the batch for each embedding dimension, preventing collapse:
| (5) |
| (6) |
Global Cosine Alignment Loss: maximizes the average cosine similarity between predicted and target embedding vectors.
| (7) |
The final loss is:
| (8) |
To test whether each component contributes to the final decoding performance, we ablate individual terms from the training objective while keeping the rest of the pipeline fixed, as shown in Table 4.
3.4 Inverting Semantic Embeddings Back to Text
Embedding inversion (Morris et al., 2023) is formulated as an iterative conditional generation problem, where the objective is to recover a text sequence given only its embedding . The procedure initializes by sampling an initial hypothesis from a base generator,
At each iteration , the current hypothesis is re-embedded to obtain , and a learned correction model generates an improved hypothesis conditioned on the current text and the embedding discrepancy:
This iterative refinement progressively reduces the embedding distance , yielding increasingly faithful reconstructions without direct optimization in discrete token space.
3.5 Data
For training and evaluation, we use LibriBrain (Özdogan et al., 2025), the largest single-subject speech-decoding MEG dataset available at the time of writing. Specifically, we use the Sherlock Holmes subset, which provides over 62 hours of MEG recordings from a single participant listening to continuous spoken narrative. The validation and test sets are held-out recording sessions, allowing us to evaluate generalization across sessions rather than across randomly sampled sentences.
| Stimuli | Words | Unique | Sentences | Hours |
|---|---|---|---|---|
| Sherlock Holmes Books (Train) | 600,107 | 20,837 | 40,659 | 61.80 |
| Sherlock Holmes Books (Validation) | 3,427 | 1,155 | 198 | 0.36 |
| Sherlock Holmes Books (Test) | 3,577 | 1,210 | 172 | 0.38 |
| Sherlock Holmes Books (Total) | 607,111 | 20,971 | 41,029 | 62.54 |
Text and audio were manually corrected, normalized, and force-aligned, with sentence boundaries defined by corpus punctuation. In this work, we prioritize dataset scale and semantic variability as key factors for semantic decoding, while deferring subject variability and cross-subject generalization to future studies. We chose MEG as our recording modality because it occupies a middle ground between fMRI and EEG: it offers high temporal resolution while providing substantially better spatial specificity than EEG. If MEG spatial resolution proves sufficient for capturing distributed semantic representations, this would open the possibility of jointly exploiting slow, distributed semantic signals and high-frequency auditory features within a single non-invasive modality.
3.5.1 Preprocessing
The recordings were originally sampled at 1 kHz and downsampled to 250 Hz to preserve oscillations into the high-gamma range (70–125 Hz).
4 Experiments
We compare our results against prior Brain2Text approaches using standard text-generation metrics: WER, BLEU, ROUGE, and BERTScore. While WER, BLEU, and ROUGE primarily measure lexical overlap and word-level reconstruction accuracy, BERTScore provides a complementary estimate of sentence-level semantic similarity. This is particularly important for our setting, where successful decoding may preserve the meaning of a sentence even when its exact wording is not recovered. At the same time, text-generation metrics alone cannot determine whether a decoded sentence was driven by neural information or by linguistic and dataset-level priors in the generation model. Following Jo et al. (2024), we therefore include a noise-control analysis to estimate how much of the decoded output is attributable to the neural input rather than to textual priors alone.
Reversing semantic embeddings reliably requires learning a high-fidelity representation of the target semantic space; otherwise, inversion becomes infeasible. It is therefore critical to identify the minimum training-set size required for the method to become reliable. To this end, we report empirical scaling laws demonstrating that the fidelity of semantic-embedding mapping improves systematically with data scale, and we identify a minimum data regime beyond which the method becomes feasible.
| Method | WER | BLEU-1 | ROUGE-1 | BERTScore |
|---|---|---|---|---|
| Sentence-level decoding, no word-level alignment | ||||
| Ours | 1.925 0.100 | 0.100 0.008 | 0.132 0.003 | 0.830 0.001 |
| Ours (noise control) | 2.648 0.415 | 0.086 0.009 | 0.113 0.004 | 0.819 0.006 |
| BrainECHO | 1.018 0.030 | 0.061 0.008 | 0.091 0.007 | 0.828 0.002 |
| BrainECHO (noise control) | 1.012 0.005 | 0.056 0.006 | 0.085 0.008 | 0.825 0.002 |
| Word-level decoding with word-aligned supervision | ||||
| d’Ascoli et al. | 0.871 0.002 | 0.190 0.004 | 0.172 0.004 | 0.820 0.001 |
| d’Ascoli et al. (noise control) | 0.994 0.003 | 0.072 0.024 | 0.057 0.021 | 0.794 0.004 |
| Method | BERTScore |
|---|---|
| Full objective (SigLIP + VICReg + global cosine) | |
| w/o global cosine loss | |
| w/o VICReg invariance term | |
| w/o VICReg variance term | |
| w/o VICReg covariance term |
4.1 Results
Comparison of method performance.
Table 3 compares the proposed Brain2Semantics2Text approach with prior Brain2Text decoding methods. The word-level metrics reveal a clear performance gap between word-level decoding, represented by d’Ascoli (d’Ascoli et al., 2024), and sentence-level decoding. This gap is expected, since word-level decoding operates in a much more constrained setting, with a closed vocabulary and access to exact word-level alignment. By contrast, when compared with acoustic-based sentence-level decoding, represented by BrainECHO (Li et al., 2024), our method performs favorably, achieving better results on BLEU-1, ROUGE-1, and BERTScore. These improvements hold both in absolute terms and when measured as neural-signal uplift relative to the noise control (Figure 3). Our method tends to generate more words than appear in the ground-truth sentence, making WER less informative; recall-based metrics such as ROUGE-1, however, show a distinct signal-based improvement.
Signal uplift.
Figure 3 evaluates the contribution of the neural signal across methods, measured using both BERTScore and ADA cosine similarity. BERTScore serves as the standard metric for comparing sentence-level reconstructions, while ADA cosine similarity provides a complementary measure of semantic similarity that more directly reflects the explicit training target of Brain2Semantics2Text. Because all methods may exploit text priors and corpus-level regularities that are not driven by the brain signal, we measure performance relative to a noise-control baseline. This signal-dependent uplift estimates the contribution of the neural signal itself, rather than reconstruction quality attributable to the inversion model, language prior, or corpus-level bias.
On BERTScore, our method yields a 1.2-point uplift over its noise-control baseline, exceeding the uplift of the previous sentence-level method, BrainECHO, but remaining below the 3.2-point uplift of d’Ascoli et al. However, d’Ascoli et al. operate at the word level and require exact word-aligned supervision, whereas our method performs sentence-level decoding without word-level alignment. On ADA cosine similarity, our method shows the largest signal-dependent uplift, improving by 6.0 points and exceeding the corresponding uplift of the word-level method. This discrepancy between semantic metrics suggests that current evaluation measures capture different aspects of sentence-level decoding performance. More broadly, it highlights the need for better standardized metrics for evaluating semantic reconstruction in brain-decoding models.
Scaling behavior of the Brain2Semantics2Text method. The method shows promising scaling behavior with increasing amounts of training data. Performance generally improved as training data increased, with gains appearing to saturate around 55.6 hours. These results indicate that the method makes effective use of additional MEG recordings of spoken language, supporting the potential value of larger-scale data collection while suggesting that future improvements may also benefit from increased data diversity.
We report performance using a retrieval-based metric: “Discounted Cumulative Gain” (Järvelin and Kekäläinen, 2002), rather than the text-based metrics, as below a certain performance threshold the full semantic inversion pipeline does not operate reliably and sentences reconstructions are not available. The retrieval metric evaluates the model’s ability to identify the closest matching sentence in the embedding space and, as such, does not capture global structural properties of the semantic manifold. Nevertheless, when all other factors remain constant, the retrieval performance provides a meaningful proxy for the overall effectiveness of the proposed method.
5 Limitations and Future Work
The current work represents an initial attempt at speech decoding through a direct mapping of MEG signals into semantic representations. Learning this semantic manifold proved to be challenging. Beyond the usual constraints of non-invasive neural recordings, such as low SNR and limited training data, reliable semantic encoding is likely to require greater semantic variability in the training corpus. We observed that the model learned semantic structure that was strongly biased toward the specific corpus used in this study. Therefore, future work should involve data collection protocols that emphasize variability in topics and concepts, perhaps guided by the properties of the target semantic manifold (Pereira et al., 2018).
The inversion methods used in this work were treated as black-box components and may not be optimally suited for semantic decoding. Further work is needed to better understand the “soft reversibility” of predicted vectors, introduced in the soft-reversibility analysis (Section 3.2), and to optimize the preservation of semantic content, potentially by training inversion mechanisms specifically for this purpose.
The current work also does not address subject variability. Although this issue is beyond the present study, semantic representations may offer new routes for cross-subject generalization by modeling both shared semantic structure and participant-specific profiles in semantic representation space.
6 Conclusion
While fully non-invasive speech decoding remains a long-term goal, decoding contextual semantic content from brain activity represents an important step toward its realization. Neuroscientific evidence suggests that speech is processed across multiple levels of representation, from fast acoustic and lexical features to slower, higher-level semantic information. These levels may provide complementary targets for neural decoding. Recovering contextual semantics from MEG alongside lower-level speech information could therefore provide multiple sources of information for reconstruction, helping compensate for the limited signal quality of non-invasive neural recordings.
References
- Anumanchipalli et al. (2019) G. K. Anumanchipalli, J. Chartier, and E. F. Chang Speech synthesis from neural decoding of spoken sentences. Nature 568, pp. 493–498. External Links: Document Cited by: §1.
- Baevski et al. (2020) A. Baevski, H. Zhou, A. Mohamed, and M. Auli Wav2vec 2.0: a framework for self-supervised learning of speech representations. In Advances in Neural Information Processing Systems (NeurIPS 2020), Note: arXiv:2006.11477 External Links: Link Cited by: §2.
- Bardes et al. (2021) A. Bardes, J. Ponce, and Y. LeCun VICReg: variance-invariance-covariance regularization for self-supervised learning. In Advances in Neural Information Processing Systems (NeurIPS), External Links: Link, Document Cited by: §3.3.
- Card et al. (2024) N. S. Card, M. Wairagkar, C. Iacobacci, X. Hou, T. Singer-Clark, F. R. Willett, E. M. Kunz, C. Fan, M. Vahdati Nia, D. R. Deo, A. Srinivasan, E. Y. Choi, M. F. Glasser, L. R. Hochberg, K. Shahlaie, S. D. Stavisky, and D. M. Brandman An accurate and rapidly calibrating speech neuroprosthesis. New England Journal of Medicine 391 (7), pp. 609–618. External Links: Document, Link Cited by: §1.
- Défossez et al. (2023) A. Défossez, C. Caucheteux, J. Rapin, O. Kabeli, and J. King Decoding speech perception from non-invasive brain recordings. Nature Machine Intelligence 5, pp. 1097–1107. External Links: Document Cited by: §2, §3.3.
- Duquenne et al. (2023) P. Duquenne, H. Schwenk, and B. Sagot SONAR: sentence-level multimodal and language-agnostic representations. External Links: 2308.11466, Document, Link Cited by: 1st item.
- d’Ascoli et al. (2024) S. d’Ascoli, C. Bel, J. Rapin, H. Banville, Y. Benchetrit, C. Pallier, and J. King Decoding individual words from non-invasive brain recordings across 723 participants. arXiv preprint arXiv:2412.17829. External Links: Link Cited by: §2, §3.3, §3.3, §4.1, Table 3, Table 3.
- Goldstein et al. (2025) A. Goldstein, E. Ham, M. Schain, S. A. Nastase, B. Aubrey, Z. Zada, A. Grinstein-Dabush, H. Gazula, A. Feder, W. Doyle, S. Devore, P. Dugan, D. Friedman, M. Brenner, A. Hassidim, Y. Matias, O. Devinsky, N. Siegelman, A. Flinker, O. Levy, R. Reichart, and U. Hasson Temporal structure of natural language processing in the human brain corresponds to layered hierarchy of large language models. Nature Communications 16 (1), pp. 10529. External Links: Document, Link Cited by: §1.
- Gwilliams et al. (2025) L. Gwilliams, I. Bhaya-Grossman, Y. Zhang, T. Scott, S. Harper, and D. Levy Computational architecture of speech comprehension in the human brain. Annual Review of Linguistics 11, pp. 209–226. External Links: Document, Link Cited by: §1.
- Gwilliams et al. (2025) L. Gwilliams et al. Hierarchical dynamic coding coordinates speech comprehension in the human brain. bioRxiv Preprint. Note: Preprint. PMCID: PMC11042271; PMID: 38659750 External Links: Link, Document Cited by: §1, §1.
- Huth et al. (2016) A. G. Huth, W. A. de Heer, T. L. Griffiths, F. E. Theunissen, and J. L. Gallant Natural speech reveals the semantic maps that tile human cerebral cortex. Nature 532 (7600), pp. 453–458. External Links: Document, Link Cited by: §1.
- Järvelin and Kekäläinen (2002) K. Järvelin and J. Kekäläinen Cumulated gain-based evaluation of IR techniques. ACM Transactions on Information Systems 20 (4), pp. 422–446. Cited by: §4.1.
- Jha et al. (2025) R. Jha, C. Zhang, V. Shmatikov, and J. X. Morris Harnessing the universal geometry of embeddings. External Links: 2505.12540, Document, Link Cited by: 1st item.
- Jo et al. (2024) H. Jo, Y. Yang, J. Han, Y. Duan, H. Xiong, and W. H. Lee Are eeg-to-text models working?. arXiv preprint arXiv:2405.06459. External Links: Document, Link Cited by: §4.
- Li et al. (2024) J. Li, Z. Song, J. Wang, M. Zhang, and Z. Zhang BrainECHO: semantic brain signal decoding through vector-quantized spectrogram reconstruction for whisper-enhanced text generation. arXiv preprint arXiv:2410.14971. External Links: Document, Link Cited by: §2, §4.1.
- Meilă and Zhang (2023) M. Meilă and H. Zhang Manifold learning: what, how, and why. arXiv preprint arXiv:2311.03757. External Links: Link, Document Cited by: §3.3.
- Minnema and Herbelot (2019) G. Minnema and A. Herbelot From brain space to distributional space: the perilous journeys of fmri decoding. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: Student Research Workshop, F. Alva-Manchego, E. Choi, and D. Khashabi (Eds.), Florence, Italy, pp. 155–161. External Links: Document, Link Cited by: §3.3.
- Morris et al. (2023) J. X. Morris, W. Zhao, J. T. Chiu, V. Shmatikov, and A. M. Rush Language model inversion. arXiv abs/2311.13647. Note: Preprint External Links: Link Cited by: 1st item, §3.4, §3.
- Moses et al. (2021) D. A. Moses, S. L. Metzger, J. R. Liu, G. K. Anumanchipalli, J. G. Makin, P. F. Sun, J. Chartier, M. E. Dougherty, P. M. Liu, G. M. Abrams, A. Tu-Chan, K. Ganguly, and E. F. Chang Neuroprosthesis for decoding speech in a paralyzed person with anarthria. New England Journal of Medicine 385 (3), pp. 217–227. External Links: Document Cited by: §1.
- Muennighoff et al. (2022) N. Muennighoff, N. Tazi, L. Magne, and N. Reimers MTEB: massive text embedding benchmark. arXiv preprint arXiv:2210.07316. External Links: Link, Document Cited by: §3.2.
- Özdogan et al. (2025) M. Özdogan, G. Landau, G. Elvers, D. Jayalath, P. Somaiya, F. Mantegna, M. Woolrich, and O. Parker Jones LibriBrain: over 50 hours of within-subject MEG to improve speech decoding methods at scale. arXiv preprint arXiv:2506.02098. External Links: Link, Document Cited by: §3.5.
- Pereira et al. (2018) F. Pereira, B. Lou, B. Pritchett, S. Ritter, S. J. Gershman, N. Kanwisher, M. Botvinick, and E. Fedorenko Toward a universal decoder of linguistic meaning from brain activation. Nature Communications 9 (1), pp. 963. External Links: Document Cited by: §2, §5.
- Radford et al. (2021) A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning, pp. 8748–8763. External Links: Link Cited by: §3.3.
- Tang et al. (2023) J. Tang, A. LeBel, S. Jain, and A. G. Huth Semantic reconstruction of continuous language from non-invasive brain recordings. Nature Neuroscience 26 (5), pp. 858–866. External Links: Document Cited by: 3rd item, §2.
- Wang et al. (2023) B. Wang, X. Xu, L. Zhang, B. Xiao, X. Wu, and J. Chen Semantic reconstruction of continuous language from meg signals. arXiv abs/2309.07701. External Links: Link Cited by: §2.
- Willett et al. (2023) F. R. Willett, E. M. Kunz, C. Fan, D. T. Avansino, G. H. Wilson, E. Y. Choi, F. Kamdar, M. F. Glasser, L. R. Hochberg, S. Druckmann, K. V. Shenoy, and J. M. Henderson A high-performance speech neuroprosthesis. Nature 620 (7976), pp. 1031–1036. External Links: Document, Link Cited by: §1.
- Willett et al. (2024) F. R. Willett, J. Li, T. Le, C. Fan, M. Chen, E. Shlizerman, Y. Chen, X. Zheng, T. S. Okubo, T. Benster, H. D. Lee, M. Kounga, E. K. Buchanan, D. Zoltowski, S. W. Linderman, and J. M. Henderson Brain-to-text benchmark ’24: lessons learned. arXiv preprint arXiv:2412.17227. External Links: Link, Document Cited by: §1.
- Yang et al. (2024a) Y. Yang, Y. Duan, H. Jo, Q. Zhang, R. Xu, O. Parker Jones, X. Hu, C. Lin, and H. Xiong NeuGPT: unified multi-modal neural gpt. arXiv preprint arXiv:2410.20916. External Links: Link Cited by: §2.
- Yang et al. (2024b) Y. Yang, Y. Duan, Q. Zhang, H. Jo, J. Zhou, W. H. Lee, R. Xu, and H. Xiong NeuSpeech: decode neural signal as speech. arXiv preprint arXiv:2403.01748. External Links: Link Cited by: §2.
- Yang et al. (2024c) Y. Yang, H. Jo, Y. Duan, Q. Zhang, J. Zhou, W. H. Lee, R. Xu, and H. Xiong MAD: multi-alignment meg-to-text decoding. arXiv abs/2406.01512. External Links: Link Cited by: §2.
- Zhai et al. (2023) X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer Sigmoid loss for language image pre-training. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), External Links: Link Cited by: §3.3.
Appendix A Impact Statement
This paper presents work toward non-invasive speech decoding, with potential applications in brain-computer interfaces for individuals who have lost the ability to speak. Clinical deployment remains distant, as current performance is still below what communication aids require, and substantial further work is needed. We also note that neural decoding technologies raise clear privacy concerns, since they involve inferring mental content from brain activity. Our work uses only publicly available research datasets with their own ethics approvals and decodes perceived speech rather than covert thought. As decoding capabilities improve, the field will need norms around consent, data ownership, and the boundary between assistive and surveillant applications. These are questions we do not resolve here, but consider essential.
Appendix B Hyperparameters
| Parameter | Value |
|---|---|
| Dropout | 0.6 |
| Learning rate | |
| Optimizer | AdamW |
| Loss type | SigLIP |
| VICReg weight | 5.0 |
| Cosine loss weight | 6.0 |
| Contrastive loss weight | 2.0 |
| Aggregation | Attention, 4 heads |
Appendix C Compute Resources
All MEG-to-semantic embedding models were trained on a single NVIDIA GPU using 4 CPU cores and 64 GiB of system memory. A typical full-data run took approximately 16–18 GPU-hours, corresponding to about 6 minutes per epoch for 150–170 epochs. Final training across seven random seeds used approximately 110–130 GPU-hours, excluding exploratory runs, failed jobs, and downstream evaluation or decoding.
Appendix D Decoding Examples
Appendix E Neural Signal Ablations
Appendix F Semantic Embedding Analysis