Brain2Semantics2Text:用语义瓶颈实现非侵入式语音解码

HuggingFace Daily Papers(社区热门论文)·2026-09-09 08:00·1天前
AI 导读

Brain2Semantics2Text 通过中间语义嵌入空间重建文本,将句子级 MEG 脑磁响应映射到语义流形,再反演为自然语言,从而绕开词级对齐。该方法面向非侵入式语音解码中信噪比低、难以还原音素或单词的难题,利用皮层分布式、慢时间尺度的语义表征,在句子级结果上优于此前的非侵入式 Brain2Text 方法。

HuggingFace Daily Papers(社区热门论文)
33AI 编辑部评分,满分 100

Brain2Semantics2Text:用语义瓶颈实现非侵入式语音解码

2026-09-09 08:00· 1天前
AI 导读

Brain2Semantics2Text 通过中间语义嵌入空间重建文本,将句子级 MEG 脑磁响应映射到语义流形,再反演为自然语言,从而绕开词级对齐。该方法面向非侵入式语音解码中信噪比低、难以还原音素或单词的难题,利用皮层分布式、慢时间尺度的语义表征,在句子级结果上优于此前的非侵入式 Brain2Text 方法。

[Uncaptioned image] [Uncaptioned image] [Uncaptioned image]
Abstract

Non-invasive speech decoding remains constrained by the low signal-to-noise ratio of neural recordings, which makes fine-grained reconstruction of phonemes or individual words difficult. Motivated by neuroscientific evidence that high-level semantic representations are distributed across cortical regions and evolve over slower temporal scales, we hypothesize that semantic content may provide a more suitable target for non-invasive decoding than low-level acoustic or lexical features. We introduce Brain2Semantics2Text, a method that reconstructs text through an intermediate semantic embedding space. Our model maps sentence-level MEG responses into a semantic manifold and then inverts the predicted embeddings into natural language. This semantic bottleneck enables recovery of high-level meaning without word-level alignment. We describe the core principles of the approach, its implementation, and the strategies used to mitigate the challenges of learning a reliable neural-to-semantic mapping. Finally, we compare against prior non-invasive Brain2Text methods and show improved sentence-level results.

1 Introduction

Speech decoding BCIs have long been a sought-after goal in both neuroscience and healthcare. Recent progress in invasive Brain2Text systems has demonstrated the feasibility of translating neural activity into language (Anumanchipalli et al., 2019; Moses et al., 2021; Willett et al., 2023; Card et al., 2024). However, these invasive approaches require surgical implantation of intracranial electrodes, creating a strong incentive to develop non-invasive speech decoding solutions that are safer, more accessible, and easier to deploy.

Extending this paradigm to non-invasive systems remains a major challenge, largely due to their inherently lower signal-to-noise ratios. Lower signal fidelity makes it difficult to reliably decode fine-grained linguistic units such as phonemes or individual words. In contrast, higher-order contextual semantic representations are spatially distributed across the cortex (Huth et al., 2016), exhibit substantial redundancy, and evolve over slower temporal scales (Gwilliams and others, 2025). These properties may make semantic representations particularly amenable to decoding from non-invasive neural recordings, whose spatial and temporal characteristics are better matched to distributed, slowly evolving signals.

Refer to caption
Figure 1: (1) MEG responses are collected while the participant listens to spoken sentence stimuli. (2) The Brain Module extracts time-resolved neural features using spatial attention and dilated temporal convolutions. (3) A backbone with temporal masked attention pools the neural sequence into a fixed-dimensional vector. (4) The vector is trained to be aligned with a pre-trained sentence-embedding space representing semantic content. (5) The predicted semantic embedding is inverted back to text.

Recently, a growing body of work in both the speech decoding literature and the neuroscience of language has converged on the view that speech comprehension and production rely on a hierarchical organization of neural representations (Gwilliams and others, 2025). Lower levels of this hierarchy are dominated by auditory and articulatory representations closely tied to the acoustic structure of speech, while progressively higher levels abstract away from these surface properties, giving rise to representations that are less dependent on specific phonetic or lexical features and more closely associated with the meaning of speech (Gwilliams et al. (2025), Goldstein et al. (2025)).

Within this framework, the success of invasive approaches can be largely attributed to their ability to directly access high-fidelity neural signals from the auditory and articulatory components of speech processing, which occupy the lower levels of the cortical hierarchy (often supplemented by post-hoc language models that guide generation; (Willett et al., 2024)). The same hierarchical view also motivates directly targeting higher-level semantic representations as a distinct and parallel neural signal. If such representations can be reliably decoded, their defining properties—slow temporal dynamics, distributed cortical organization, and representational redundancy—are better matched to the spatial and temporal characteristics of non-invasive modalities such as fMRI, MEG, and EEG.

This work targets semantic representations in brain activity, focusing on higher-level stages of the speech-processing hierarchy. Brain2Semantics2Text maps sentence-length MEG responses during heard speech into a pretrained semantic embedding space, from which text is subsequently reconstructed. By constraining neural decoding to pass through this semantic bottleneck, the method aims to shift the objective toward higher-level speech representations that carry information about sentence-level meaning.

Several aspects distinguish our method from previous work on Brain2Text:

  • Semantic embedding inversion. We build on recent advances in semantic embedding inversion (Morris et al., 2023), which enables the reconstruction of text from semantic embeddings either as an intrinsic property of newer embedding models (Duquenne et al., 2023) or via general inversion techniques applicable to arbitrary pre-trained semantic spaces (Jha et al., 2025). This allows us to frame speech decoding as semantic reconstruction rather than word or phoneme prediction.

  • Sentence-level semantic decoding. Our method operates directly at the sentence level, targeting compositional semantic representations near the top of the speech-processing hierarchy. This formulation shifts the decoding problem away from exact word or phoneme recovery and toward reconstruction of the intended semantic content. As a result, it avoids dependence on precise word-level alignment and closed-vocabulary supervision, both of which are difficult to assume in realistic settings. Although full-sentence decoding from MEG is ambitious, sentence-level semantic decoding is well matched to the distributed and temporally extended nature of high-level language representations, making it a promising direction for future non-invasive communication systems.

  • Semantic decoding from MEG. Unlike most prior semantic decoding work, which relies on the high spatial resolution of fMRI (Tang et al., 2023), we leverage the largest single-subject heard-speech MEG dataset of its kind to date. Decoding semantic-level information from MEG enables the joint exploitation of slow, distributed semantic signals and local, high-frequency neural activity within the same recordings, opening new possibilities for improved speech decoding.

2 Related Work

Semantic decoding.

Semantic representations have previously been used as an intermediate target for non-invasive language decoding. Pereira et al. (2018) demonstrated that fMRI responses to sentences could be mapped into semantic embedding spaces, providing early evidence that distributed neural activity can be aligned with sentence-level meaning. More recently, Tang et al. (2023) reconstructed continuous perceived and imagined language from fMRI by mapping distributed cortical responses into semantic representations and using these representations to constrain language generation. These approaches exploit the high spatial resolution of fMRI to recover distributed semantic information, but sacrifice the temporal resolution available in electrophysiological recordings.

Wang et al. (2023) extended semantic reconstruction to MEG, demonstrating that semantic information can also be recovered from temporally resolved non-invasive recordings. Their approach, however, operates at the word level, reconstructing a temporally aligned sequence of contextual word embeddings that is subsequently used to generate continuous text. In contrast, our method treats the entire sentence as the unit of decoding, mapping sentence-length MEG responses directly into a single pre-trained sentence-level semantic embedding. This formulation makes the semantic representation itself the decoding bottleneck and removes the need for word-level alignment.

Auditory and speech-based MEG decoding.

A complementary line of work maps MEG activity to representations closely tied to the acoustic structure of speech. Défossez et al. (2023) aligned MEG responses with Wav2Vec representations (Baevski et al., 2020), establishing a contrastive-learning framework and neural architecture that have influenced subsequent MEG decoding systems. More recent approaches align MEG with auditory representations (Yang et al., 2024c), adapt Whisper for neural speech decoding (Yang et al., 2024b), or incorporate neural signals into multimodal foundation-model architectures (Yang et al., 2024a). These approaches exploit neural information associated with the acoustic realization of speech, whereas our method deliberately targets higher-level semantic content.

Within this line of work, BrainECHO (Li et al., 2024) provides the closest comparison to our method. Like Brain2Semantics2Text, BrainECHO operates at the sentence level and does not require word-level alignment, but the two methods differ in the representation through which decoding proceeds. BrainECHO maps neural activity into a vector-quantized audio-spectrogram latent space before generating text with Whisper, whereas our method maps MEG directly into a sentence-level semantic embedding. BrainECHO therefore provides a particularly informative baseline for evaluating the use of semantic, rather than acoustic, representations as a bottleneck for sentence-level decoding.

Word-level classification.

Semantic representations have also been used for word-level MEG decoding. d’Ascoli et al. (2024) achieve strong closed-vocabulary word classification by aligning MEG responses with lexical semantic embeddings augmented by sentence context. Their approach demonstrates the utility of semantic representations for MEG decoding, but relies on exact word-level timing and formulates decoding as classification among candidate words. Our approach instead targets a single compositional representation of the complete sentence, removing the requirement for word-level alignment at the cost of a substantially less constrained reconstruction problem.

3 Method

The Brain2Semantics2Text method operates in two stages. In the first stage, MEG neural responses corresponding to continuously presented spoken sentences are mapped to vector representations in the pre-trained semantic embedding space. Training is guided by objectives that encourage alignment with the target embeddings while preserving their global statistical structure. In the second stage, the predicted semantic embedding is inverted into natural language using a pre-trained inversion model (Morris et al., 2023) that reconstructs text from semantic vectors.

3.1 Backbone

MEG-to-semantic mapping.

The MEG input signal 𝐱C×T, where C denotes the number of sensors and T the number of temporal samples, is first processed by a spatial attention module, followed by an initial 1×1 convolution that projects the sensor dimension into a latent feature space. The resulting representation is then passed through a stack of dilated temporal convolutional blocks. Each block consists of dilated convolutions equipped with residual connections and gated linear units (GLUs), enabling the model to capture long-range temporal dependencies.

A subject-specific layer can optionally be inserted after the initial projection to model inter-subject variability. In the experiments reported here all data originate from a single subject.

Temporal aggregation.

To obtain a fixed-dimensional semantic representation from the time-resolved features, we use a 4-head self-attention Transformer over the temporal axis. We then pool over time with a masked mean

𝐲^=t=1TmtHtt=1Tmtd,

where mt is a temporal mask based on the true segment lengths, preventing length-related surface confounders from influencing the pooled embedding.

3.2 Semantic Embeddings

Many semantic embedding models are available, with different architectures, training objectives, and benchmark performance (Muennighoff et al., 2022). For Brain2Semantics2Text, however, standard evaluations on semantic similarity or retrieval tasks provide only partial guidance. Our method requires the embedding space to function as an invertible bottleneck between MEG responses and text, which introduces additional constraints beyond general semantic performance. In particular, the embedding must have an available inversion mechanism and must support reliable reconstruction from both exact text embeddings and imperfect neural predictions. We therefore evaluate candidate embedding spaces according to four criteria.

Expressivity.

The embedding space should preserve enough information about the original sentence to support reconstruction. We measure this using a round-trip reconstruction test, in which each sentence is embedded and then inverted back into text. Higher reconstruction quality indicates that more sentence-level information is retained by the embedding and its inversion procedure.

Soft reversibility.

At test time, the vectors being inverted are not exact text embeddings, but embeddings predicted from noisy MEG responses. The embedding space should therefore be robust to prediction error: vectors near the target should still invert to text with similar meaning. We assess this by perturbing target embeddings and calculating the relation between the introduced noise and reconstruction fidelity.

Length bias.

The embedding should encode sentence meaning without being dominated by surface-level properties such as sentence length. This is particularly important because stimulus duration and sentence length may be available to the neural decoder and could provide a shortcut that competes with semantic learning. We estimate length bias by computing the Spearman rank correlation (ρ) between sentence length and the leading principal components of each embedding space. Although this measure does not capture all forms of embedding sensitivity to sentence length, we find that it provides an effective empirical proxy for the extent to which sentence length is reflected in the global geometry of the embedding space. A visual illustration of this analysis is provided in Appendix F.

Intrinsic dimensionality.

The target space should be learnable from limited MEG data. We therefore prefer embedding spaces with lower intrinsic dimensionality, measured by effective rank, provided that they remain sufficiently expressive and reversible.

A further consideration is training-data transparency. In principle, a fully open embedding and inversion pipeline would be preferable, since it would allow us to verify that evaluation sentences were not present in the training data of either the embedding model or the inversion model. Among the embedding–inversion pairs we considered, however, we did not find a fully data-transparent option that also satisfied the practical requirements of expressivity and soft reversibility. We therefore control for possible corpus-leakage by comparing neural-based predictions against noise-control.

媒体内容 · 前往原文查看
Table 1: Comparison of candidate embedding spaces according to expressivity, soft reversibility, length bias, and intrinsic dimensionality. ADA was chosen for its combination of high soft reversibility, low length bias, and good expressivity.
Embedding Expressivity Soft reversibility Length bias Intrinsic dim.
SONAR 0.981±0.019 0.902 0.819 1024/197
T5 0.813±0.018 0.333 0.569 𝟕𝟔𝟖/𝟏𝟖𝟑
ADA 0.888±0.031 0.925 0.432 1536/195

3.3 Objectives for Manifold Learning

At the core of our training setup is a SigLIP-style contrastive loss (Zhai et al., 2023; d’Ascoli et al., 2024), which has been shown to be effective for aligning representations across modalities with different dimensionalities and statistical characteristics (Radford et al., 2021). However, in the low-data regime typical of non-invasive speech decoding, contrastive objectives alone are insufficient to learn the target manifold of semantic embeddings.

Previous brain-to-text and word-decoding approaches have largely relied on contrastive objectives to align neural signals with semantic embeddings (Défossez et al., 2023; d’Ascoli et al., 2024). However, we observe that in the low-data regimes typical of non-invasive speech decoding, contrastive loss primarily optimizes a retrieval objective. In this setting, the model learns a mapping that enables nearest-neighbor matching under cosine similarity (Minnema and Herbelot, 2019), but does not necessarily preserve the inter-vector distances or the scale of embedding magnitudes.

This limitation is problematic for our setting, where the goal is not merely to retrieve a correct target embedding, but to learn a mapping that faithfully captures the global geometry of the target semantic manifold. To address this, we draw inspiration from the manifold learning literature (Meilă and Zhang, 2023) and introduce several auxiliary losses in addition to the SigLIP that force the model to learn the global properties of the target manifold and prevent collapse. The resulting training objective consists of the following components (invariance, covariance and variance losses are adopted from VICReg (Bardes et al., 2021)):

SigLIP Loss: A contrastive alignment term that formulates predicted–target matching as independent pairwise classification rather than a batch-wise softmax objective. This is useful for sentence-level semantic decoding, where different non-matching sentences may still be semantically related and should not necessarily be treated as mutually exclusive classes. We also found this objective more stable in the low-data MEG setting, particularly with small batches.

ctr=1ni=1nj=1nlog(1+exp(tij(τ𝐲^i𝐲j+b))), (1)

Invariance Loss : mean squared distance between predicted and target embeddings.

inv=1ni=1n𝐱i𝐲i22. (2)

Covariance Loss: a decorrelation term that penalizes off-diagonal covariances between embedding dimensions, reducing redundancy and preventing informational collapse.

cov=1dijC(𝐱)i,j2. (3)
C(𝐱)=1n1k=1n(𝐱k𝐱¯)(𝐱k𝐱¯), (4)

Variance Loss: a hinge loss that enforces a minimum standard deviation across the batch for each embedding dimension, preventing collapse:

var=1dj=1dmax(0,γσj(𝐱)), (5)
σj(𝐱)=Var(xj)+ϵ. (6)

Global Cosine Alignment Loss: maximizes the average cosine similarity between predicted and target embedding vectors.

gcs=11ni=1n𝐲^i𝐲i𝐲^i2𝐲i2. (7)

The final loss is:

=αctr+βgcs+γinv+δvar+ηcov. (8)

To test whether each component contributes to the final decoding performance, we ablate individual terms from the training objective while keeping the rest of the pipeline fixed, as shown in Table 4.

3.4 Inverting Semantic Embeddings Back to Text

Embedding inversion (Morris et al., 2023) is formulated as an iterative conditional generation problem, where the objective is to recover a text sequence x given only its embedding e=f(x). The procedure initializes by sampling an initial hypothesis from a base generator,

x0pθ(xe).

At each iteration t, the current hypothesis xt is re-embedded to obtain et=f(xt), and a learned correction model generates an improved hypothesis conditioned on the current text and the embedding discrepancy:

xt+1pϕ(xxt,et,e).

This iterative refinement progressively reduces the embedding distance ete, yielding increasingly faithful reconstructions without direct optimization in discrete token space.

3.5 Data

For training and evaluation, we use LibriBrain (Özdogan et al., 2025), the largest single-subject speech-decoding MEG dataset available at the time of writing. Specifically, we use the Sherlock Holmes subset, which provides over 62 hours of MEG recordings from a single participant listening to continuous spoken narrative. The validation and test sets are held-out recording sessions, allowing us to evaluate generalization across sessions rather than across randomly sampled sentences.

媒体内容 · 前往原文查看
Table 2: LibriBrain’s Sherlock Holmes data splits used in our experiments. Validation and test sets are held-out recording sessions.
Stimuli Words Unique Sentences Hours
Sherlock Holmes Books (Train) 600,107 20,837 40,659 61.80
Sherlock Holmes Books (Validation) 3,427 1,155 198 0.36
Sherlock Holmes Books (Test) 3,577 1,210 172 0.38
Sherlock Holmes Books (Total) 607,111 20,971 41,029 62.54

Text and audio were manually corrected, normalized, and force-aligned, with sentence boundaries defined by corpus punctuation. In this work, we prioritize dataset scale and semantic variability as key factors for semantic decoding, while deferring subject variability and cross-subject generalization to future studies. We chose MEG as our recording modality because it occupies a middle ground between fMRI and EEG: it offers high temporal resolution while providing substantially better spatial specificity than EEG. If MEG spatial resolution proves sufficient for capturing distributed semantic representations, this would open the possibility of jointly exploiting slow, distributed semantic signals and high-frequency auditory features within a single non-invasive modality.

3.5.1 Preprocessing

The recordings were originally sampled at 1 kHz and downsampled to 250 Hz to preserve oscillations into the high-gamma range (70–125 Hz).

4 Experiments

We compare our results against prior Brain2Text approaches using standard text-generation metrics: WER, BLEU, ROUGE, and BERTScore. While WER, BLEU, and ROUGE primarily measure lexical overlap and word-level reconstruction accuracy, BERTScore provides a complementary estimate of sentence-level semantic similarity. This is particularly important for our setting, where successful decoding may preserve the meaning of a sentence even when its exact wording is not recovered. At the same time, text-generation metrics alone cannot determine whether a decoded sentence was driven by neural information or by linguistic and dataset-level priors in the generation model. Following Jo et al. (2024), we therefore include a noise-control analysis to estimate how much of the decoded output is attributable to the neural input rather than to textual priors alone.

Reversing semantic embeddings reliably requires learning a high-fidelity representation of the target semantic space; otherwise, inversion becomes infeasible. It is therefore critical to identify the minimum training-set size required for the method to become reliable. To this end, we report empirical scaling laws demonstrating that the fidelity of semantic-embedding mapping improves systematically with data scale, and we identify a minimum data regime beyond which the method becomes feasible.

媒体内容 · 前往原文查看
Table 3: Comparison of decoding methods (mean ± s.d.). Brain2Semantics2Text and BrainECHO operate at the sentence level and do not rely on word-level alignment, whereas d’Ascoli et al. (2024) use word-aligned supervision. Lexical-overlap metrics such as WER, BLEU-1, and ROUGE-1 therefore favor methods with access to word-level timing or lexical supervision, while BERTScore better reflects the sentence-level semantic reconstruction objective of our method.
Method WER BLEU-1 ROUGE-1 BERTScore
Sentence-level decoding, no word-level alignment
Ours 1.925 ± 0.100 0.100 ± 0.008 0.132 ± 0.003 0.830 ± 0.001
Ours (noise control) 2.648 ± 0.415 0.086 ± 0.009 0.113 ± 0.004 0.819 ± 0.006
BrainECHO 1.018 ± 0.030 0.061 ± 0.008 0.091 ± 0.007 0.828 ± 0.002
BrainECHO (noise control) 1.012 ± 0.005 0.056 ± 0.006 0.085 ± 0.008 0.825 ± 0.002
Word-level decoding with word-aligned supervision
d’Ascoli et al. 0.871 ± 0.002 0.190 ± 0.004 0.172 ± 0.004 0.820 ± 0.001
d’Ascoli et al. (noise control) 0.994 ± 0.003 0.072 ± 0.024 0.057 ± 0.021 0.794 ± 0.004
媒体内容 · 前往原文查看
Figure 2: Semantic embedding inversion. We evaluated semantic embedding inversion by adding controlled Gaussian noise with increasing standard deviation to the sentence embedding vectors and measuring the BERTScore F1 of the reconstructed text. Expressivity is defined as the reconstruction score obtained from the unperturbed embedding, corresponding to the ideal case of a perfectly predicted vector. Soft reversibility is measured by how smoothly reconstruction quality degrades as embedding noise increases, summarized by the R2 of the fitted relationship between noise level and BERTScore. Actual neural predictions are overlaid to show where model outputs fall along the resulting noise–performance curve (similar analyses for all candidate embeddings are provided in Appendix  F).
媒体内容 · 前往原文查看
Table 4: Loss ablation results measured as mean decoded-sentence BERTScore F1 on the test set. Values are mean ± standard deviation across runs. Higher values indicate better decoded-sentence semantic similarity.
Method BERTScore
Full objective (SigLIP + VICReg + global cosine) 0.8297±0.0008
   w/o global cosine loss 0.8103±0.0077
   w/o VICReg invariance term 0.8263±0.0025
   w/o VICReg variance term 0.8267±0.0015
   w/o VICReg covariance term 0.8270±0.0012

4.1 Results

媒体内容 · 前往原文查看
Figure 3: Signal uplift. Comparison of signal-dependent improvement over noise-control baselines across decoding methods.
媒体内容 · 前往原文查看
Figure 4: Scaling Laws
Comparison of method performance.

Table 3 compares the proposed Brain2Semantics2Text approach with prior Brain2Text decoding methods. The word-level metrics reveal a clear performance gap between word-level decoding, represented by d’Ascoli (d’Ascoli et al., 2024), and sentence-level decoding. This gap is expected, since word-level decoding operates in a much more constrained setting, with a closed vocabulary and access to exact word-level alignment. By contrast, when compared with acoustic-based sentence-level decoding, represented by BrainECHO (Li et al., 2024), our method performs favorably, achieving better results on BLEU-1, ROUGE-1, and BERTScore. These improvements hold both in absolute terms and when measured as neural-signal uplift relative to the noise control (Figure 3). Our method tends to generate more words than appear in the ground-truth sentence, making WER less informative; recall-based metrics such as ROUGE-1, however, show a distinct signal-based improvement.

Signal uplift.

Figure 3 evaluates the contribution of the neural signal across methods, measured using both BERTScore and ADA cosine similarity. BERTScore serves as the standard metric for comparing sentence-level reconstructions, while ADA cosine similarity provides a complementary measure of semantic similarity that more directly reflects the explicit training target of Brain2Semantics2Text. Because all methods may exploit text priors and corpus-level regularities that are not driven by the brain signal, we measure performance relative to a noise-control baseline. This signal-dependent uplift estimates the contribution of the neural signal itself, rather than reconstruction quality attributable to the inversion model, language prior, or corpus-level bias.

On BERTScore, our method yields a 1.2-point uplift over its noise-control baseline, exceeding the uplift of the previous sentence-level method, BrainECHO, but remaining below the 3.2-point uplift of d’Ascoli et al. However, d’Ascoli et al. operate at the word level and require exact word-aligned supervision, whereas our method performs sentence-level decoding without word-level alignment. On ADA cosine similarity, our method shows the largest signal-dependent uplift, improving by 6.0 points and exceeding the corresponding uplift of the word-level method. This discrepancy between semantic metrics suggests that current evaluation measures capture different aspects of sentence-level decoding performance. More broadly, it highlights the need for better standardized metrics for evaluating semantic reconstruction in brain-decoding models.

Scaling behavior of the Brain2Semantics2Text method. The method shows promising scaling behavior with increasing amounts of training data. Performance generally improved as training data increased, with gains appearing to saturate around 55.6 hours. These results indicate that the method makes effective use of additional MEG recordings of spoken language, supporting the potential value of larger-scale data collection while suggesting that future improvements may also benefit from increased data diversity.

We report performance using a retrieval-based metric: “Discounted Cumulative Gain” (Järvelin and Kekäläinen, 2002), rather than the text-based metrics, as below a certain performance threshold the full semantic inversion pipeline does not operate reliably and sentences reconstructions are not available. The retrieval metric evaluates the model’s ability to identify the closest matching sentence in the embedding space and, as such, does not capture global structural properties of the semantic manifold. Nevertheless, when all other factors remain constant, the retrieval performance provides a meaningful proxy for the overall effectiveness of the proposed method.

5 Limitations and Future Work

The current work represents an initial attempt at speech decoding through a direct mapping of MEG signals into semantic representations. Learning this semantic manifold proved to be challenging. Beyond the usual constraints of non-invasive neural recordings, such as low SNR and limited training data, reliable semantic encoding is likely to require greater semantic variability in the training corpus. We observed that the model learned semantic structure that was strongly biased toward the specific corpus used in this study. Therefore, future work should involve data collection protocols that emphasize variability in topics and concepts, perhaps guided by the properties of the target semantic manifold (Pereira et al., 2018).

The inversion methods used in this work were treated as black-box components and may not be optimally suited for semantic decoding. Further work is needed to better understand the “soft reversibility” of predicted vectors, introduced in the soft-reversibility analysis (Section 3.2), and to optimize the preservation of semantic content, potentially by training inversion mechanisms specifically for this purpose.

The current work also does not address subject variability. Although this issue is beyond the present study, semantic representations may offer new routes for cross-subject generalization by modeling both shared semantic structure and participant-specific profiles in semantic representation space.

6 Conclusion

While fully non-invasive speech decoding remains a long-term goal, decoding contextual semantic content from brain activity represents an important step toward its realization. Neuroscientific evidence suggests that speech is processed across multiple levels of representation, from fast acoustic and lexical features to slower, higher-level semantic information. These levels may provide complementary targets for neural decoding. Recovering contextual semantics from MEG alongside lower-level speech information could therefore provide multiple sources of information for reconstruction, helping compensate for the limited signal quality of non-invasive neural recordings.

References

  • Anumanchipalli et al. (2019) G. K. Anumanchipalli, J. Chartier, and E. F. Chang Speech synthesis from neural decoding of spoken sentences. Nature 568, pp. 493–498. External Links: Document Cited by: §1.
  • Baevski et al. (2020) A. Baevski, H. Zhou, A. Mohamed, and M. Auli Wav2vec 2.0: a framework for self-supervised learning of speech representations. In Advances in Neural Information Processing Systems (NeurIPS 2020), Note: arXiv:2006.11477 External Links: Link Cited by: §2.
  • Bardes et al. (2021) A. Bardes, J. Ponce, and Y. LeCun VICReg: variance-invariance-covariance regularization for self-supervised learning. In Advances in Neural Information Processing Systems (NeurIPS), External Links: Link, Document Cited by: §3.3.
  • Card et al. (2024) N. S. Card, M. Wairagkar, C. Iacobacci, X. Hou, T. Singer-Clark, F. R. Willett, E. M. Kunz, C. Fan, M. Vahdati Nia, D. R. Deo, A. Srinivasan, E. Y. Choi, M. F. Glasser, L. R. Hochberg, K. Shahlaie, S. D. Stavisky, and D. M. Brandman An accurate and rapidly calibrating speech neuroprosthesis. New England Journal of Medicine 391 (7), pp. 609–618. External Links: Document, Link Cited by: §1.
  • Défossez et al. (2023) A. Défossez, C. Caucheteux, J. Rapin, O. Kabeli, and J. King Decoding speech perception from non-invasive brain recordings. Nature Machine Intelligence 5, pp. 1097–1107. External Links: Document Cited by: §2, §3.3.
  • Duquenne et al. (2023) P. Duquenne, H. Schwenk, and B. Sagot SONAR: sentence-level multimodal and language-agnostic representations. External Links: 2308.11466, Document, Link Cited by: 1st item.
  • d’Ascoli et al. (2024) S. d’Ascoli, C. Bel, J. Rapin, H. Banville, Y. Benchetrit, C. Pallier, and J. King Decoding individual words from non-invasive brain recordings across 723 participants. arXiv preprint arXiv:2412.17829. External Links: Link Cited by: §2, §3.3, §3.3, §4.1, Table 3, Table 3.
  • Goldstein et al. (2025) A. Goldstein, E. Ham, M. Schain, S. A. Nastase, B. Aubrey, Z. Zada, A. Grinstein-Dabush, H. Gazula, A. Feder, W. Doyle, S. Devore, P. Dugan, D. Friedman, M. Brenner, A. Hassidim, Y. Matias, O. Devinsky, N. Siegelman, A. Flinker, O. Levy, R. Reichart, and U. Hasson Temporal structure of natural language processing in the human brain corresponds to layered hierarchy of large language models. Nature Communications 16 (1), pp. 10529. External Links: Document, Link Cited by: §1.
  • Gwilliams et al. (2025) L. Gwilliams, I. Bhaya-Grossman, Y. Zhang, T. Scott, S. Harper, and D. Levy Computational architecture of speech comprehension in the human brain. Annual Review of Linguistics 11, pp. 209–226. External Links: Document, Link Cited by: §1.
  • Gwilliams et al. (2025) L. Gwilliams et al. Hierarchical dynamic coding coordinates speech comprehension in the human brain. bioRxiv Preprint. Note: Preprint. PMCID: PMC11042271; PMID: 38659750 External Links: Link, Document Cited by: §1, §1.
  • Huth et al. (2016) A. G. Huth, W. A. de Heer, T. L. Griffiths, F. E. Theunissen, and J. L. Gallant Natural speech reveals the semantic maps that tile human cerebral cortex. Nature 532 (7600), pp. 453–458. External Links: Document, Link Cited by: §1.
  • Järvelin and Kekäläinen (2002) K. Järvelin and J. Kekäläinen Cumulated gain-based evaluation of IR techniques. ACM Transactions on Information Systems 20 (4), pp. 422–446. Cited by: §4.1.
  • Jha et al. (2025) R. Jha, C. Zhang, V. Shmatikov, and J. X. Morris Harnessing the universal geometry of embeddings. External Links: 2505.12540, Document, Link Cited by: 1st item.
  • Jo et al. (2024) H. Jo, Y. Yang, J. Han, Y. Duan, H. Xiong, and W. H. Lee Are eeg-to-text models working?. arXiv preprint arXiv:2405.06459. External Links: Document, Link Cited by: §4.
  • Li et al. (2024) J. Li, Z. Song, J. Wang, M. Zhang, and Z. Zhang BrainECHO: semantic brain signal decoding through vector-quantized spectrogram reconstruction for whisper-enhanced text generation. arXiv preprint arXiv:2410.14971. External Links: Document, Link Cited by: §2, §4.1.
  • Meilă and Zhang (2023) M. Meilă and H. Zhang Manifold learning: what, how, and why. arXiv preprint arXiv:2311.03757. External Links: Link, Document Cited by: §3.3.
  • Minnema and Herbelot (2019) G. Minnema and A. Herbelot From brain space to distributional space: the perilous journeys of fmri decoding. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: Student Research Workshop, F. Alva-Manchego, E. Choi, and D. Khashabi (Eds.), Florence, Italy, pp. 155–161. External Links: Document, Link Cited by: §3.3.
  • Morris et al. (2023) J. X. Morris, W. Zhao, J. T. Chiu, V. Shmatikov, and A. M. Rush Language model inversion. arXiv abs/2311.13647. Note: Preprint External Links: Link Cited by: 1st item, §3.4, §3.
  • Moses et al. (2021) D. A. Moses, S. L. Metzger, J. R. Liu, G. K. Anumanchipalli, J. G. Makin, P. F. Sun, J. Chartier, M. E. Dougherty, P. M. Liu, G. M. Abrams, A. Tu-Chan, K. Ganguly, and E. F. Chang Neuroprosthesis for decoding speech in a paralyzed person with anarthria. New England Journal of Medicine 385 (3), pp. 217–227. External Links: Document Cited by: §1.
  • Muennighoff et al. (2022) N. Muennighoff, N. Tazi, L. Magne, and N. Reimers MTEB: massive text embedding benchmark. arXiv preprint arXiv:2210.07316. External Links: Link, Document Cited by: §3.2.
  • Özdogan et al. (2025) M. Özdogan, G. Landau, G. Elvers, D. Jayalath, P. Somaiya, F. Mantegna, M. Woolrich, and O. Parker Jones LibriBrain: over 50 hours of within-subject MEG to improve speech decoding methods at scale. arXiv preprint arXiv:2506.02098. External Links: Link, Document Cited by: §3.5.
  • Pereira et al. (2018) F. Pereira, B. Lou, B. Pritchett, S. Ritter, S. J. Gershman, N. Kanwisher, M. Botvinick, and E. Fedorenko Toward a universal decoder of linguistic meaning from brain activation. Nature Communications 9 (1), pp. 963. External Links: Document Cited by: §2, §5.
  • Radford et al. (2021) A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, G. Krueger, and I. Sutskever Learning transferable visual models from natural language supervision. In Proceedings of the 38th International Conference on Machine Learning, pp. 8748–8763. External Links: Link Cited by: §3.3.
  • Tang et al. (2023) J. Tang, A. LeBel, S. Jain, and A. G. Huth Semantic reconstruction of continuous language from non-invasive brain recordings. Nature Neuroscience 26 (5), pp. 858–866. External Links: Document Cited by: 3rd item, §2.
  • Wang et al. (2023) B. Wang, X. Xu, L. Zhang, B. Xiao, X. Wu, and J. Chen Semantic reconstruction of continuous language from meg signals. arXiv abs/2309.07701. External Links: Link Cited by: §2.
  • Willett et al. (2023) F. R. Willett, E. M. Kunz, C. Fan, D. T. Avansino, G. H. Wilson, E. Y. Choi, F. Kamdar, M. F. Glasser, L. R. Hochberg, S. Druckmann, K. V. Shenoy, and J. M. Henderson A high-performance speech neuroprosthesis. Nature 620 (7976), pp. 1031–1036. External Links: Document, Link Cited by: §1.
  • Willett et al. (2024) F. R. Willett, J. Li, T. Le, C. Fan, M. Chen, E. Shlizerman, Y. Chen, X. Zheng, T. S. Okubo, T. Benster, H. D. Lee, M. Kounga, E. K. Buchanan, D. Zoltowski, S. W. Linderman, and J. M. Henderson Brain-to-text benchmark ’24: lessons learned. arXiv preprint arXiv:2412.17227. External Links: Link, Document Cited by: §1.
  • Yang et al. (2024a) Y. Yang, Y. Duan, H. Jo, Q. Zhang, R. Xu, O. Parker Jones, X. Hu, C. Lin, and H. Xiong NeuGPT: unified multi-modal neural gpt. arXiv preprint arXiv:2410.20916. External Links: Link Cited by: §2.
  • Yang et al. (2024b) Y. Yang, Y. Duan, Q. Zhang, H. Jo, J. Zhou, W. H. Lee, R. Xu, and H. Xiong NeuSpeech: decode neural signal as speech. arXiv preprint arXiv:2403.01748. External Links: Link Cited by: §2.
  • Yang et al. (2024c) Y. Yang, H. Jo, Y. Duan, Q. Zhang, J. Zhou, W. H. Lee, R. Xu, and H. Xiong MAD: multi-alignment meg-to-text decoding. arXiv abs/2406.01512. External Links: Link Cited by: §2.
  • Zhai et al. (2023) X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer Sigmoid loss for language image pre-training. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), External Links: Link Cited by: §3.3.

Appendix A Impact Statement

This paper presents work toward non-invasive speech decoding, with potential applications in brain-computer interfaces for individuals who have lost the ability to speak. Clinical deployment remains distant, as current performance is still below what communication aids require, and substantial further work is needed. We also note that neural decoding technologies raise clear privacy concerns, since they involve inferring mental content from brain activity. Our work uses only publicly available research datasets with their own ethics approvals and decodes perceived speech rather than covert thought. As decoding capabilities improve, the field will need norms around consent, data ownership, and the boundary between assistive and surveillant applications. These are questions we do not resolve here, but consider essential.

Appendix B Hyperparameters

媒体内容 · 前往原文查看
Table 5: Hyperparameters used for MEG-to-semantic embedding training.
Parameter Value
Dropout 0.6
Learning rate 3×105
Optimizer AdamW
Loss type SigLIP
VICReg weight 5.0
Cosine loss weight 6.0
Contrastive loss weight 2.0
Aggregation Attention, 4 heads

Appendix C Compute Resources

All MEG-to-semantic embedding models were trained on a single NVIDIA GPU using 4 CPU cores and 64 GiB of system memory. A typical full-data run took approximately 16–18 GPU-hours, corresponding to about 6 minutes per epoch for 150–170 epochs. Final training across seven random seeds used approximately 110–130 GPU-hours, excluding exploratory runs, failed jobs, and downstream evaluation or decoding.

Appendix D Decoding Examples

Refer to caption
Figure 5: Semantic geometry of reconstructed sentence embeddings. PCA projections of sentence-level embeddings for ground-truth sentences and reconstructed predictions. Highlighted examples illustrate that relative positions and global structure are largely preserved under reconstruction. Examples are shown for geometric comparison only, and textual reconstructions may differ from the ground truth.
Refer to caption
Figure 6: Qualitative decoding examples. Example decoded sentences shown for qualitative illustration of semantic similarity between the target sentence and the reconstructed output.

Appendix E Neural Signal Ablations

Refer to caption
Figure 7: Neural signal properties supporting semantic decoding. We examine how different properties of the MEG signal affect semantic mapping performance, providing an interpretable view of which aspects of the neural response contribute most to decoding.

Appendix F Semantic Embedding Analysis

Refer to caption
Figure 8: PCA comparison of ADA and SONAR semantic embeddings, colored by sentence length. While SONAR embeddings segregate sentences by length, forming length-dependent regions in the embedding space, ADA embeddings remain more invariant to sentence length. This suggests that ADA is less vulnerable to the sentence-length confound.
媒体内容 · 前往原文查看
Figure 9: Comparison of candidate embedding spaces across embedding-space diagnostics. Candidate semantic embedding spaces are compared according to expressivity, soft reversibility, length bias, and intrinsic dimensionality. A full explanation of how these measures are computed is provided in the appendix.

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org