Source-linked AI summary
Brain-to-Text Decoding: A Non-invasive Approach via Typing
Jarod Lévy, Mingfang Zhang, Svetlana Pinet, Jérémy Rapin, Hubert Banville, Stéphane d'Ascoli, Jean-Rémi King
TL;DR
Invasive communication BCIs require neurosurgery and are difficult to scale, while existing non-invasive EEG systems remain moderate in performance and demanding for users. Brain2Qwerty is a three-stage deep neural network trained on brain recordings from 35 participants typing briefly memorized sentences, using either EEG or MEG. Brain2Qwerty achieved 32±0.6% character-error-rate (CER) with MEG, with the best-performing participants reaching a CER as low as 19%.
Problem
Invasive communication BCIs require neurosurgery and are difficult to scale, while existing non-invasive EEG systems remain moderate in performance and demanding for users.
Method
Brain2Qwerty is a three-stage deep neural network trained on brain recordings from 35 participants typing briefly memorized sentences, using either EEG or MEG.
Results
Brain2Qwerty achieved 32±0.6% character-error-rate (CER) with MEG, with the best-performing participants reaching a CER as low as 19%.
Takeaways & Limitations
The results provide a stepping stone toward safer and more accessible non-invasive BCIs for people who have partially or completely lost the ability to communicate.
Takeaways & Limitations
The method is not real time and was tested only with healthy participants using supervised training that requires the timing and identity of each character.
Abstract
from arXiv · showhide
Modern neuroprostheses can now restore communication in patients who have lost the ability to speak or move. However, these invasive devices entail risks inherent to neurosurgery. Here, we introduce a non-invasive method to decode the production of sentences from brain activity and demonstrate its efficacy in a cohort of 35 healthy volunteers. For this, we present Brain2Qwerty, a new deep learning architecture trained to decode sentences from either electro- (EEG) or magneto-encephalography (MEG), while participants typed briefly memorized sentences on a QWERTY keyboard. With MEG, Brain2Qwerty reaches, on average, a character-error-rate (CER) of 32% and substantially outperforms EEG (CER: 67%). For the best participants, the model achieves a CER of 19%, and can perfectly decode a variety of sentences outside of the training set. While error analyses suggest that decoding depends on motor processes, the analysis of typographical errors suggests that it also involves higher-level cognitive factors. Overall, these results narrow the gap between invasive and non-invasive methods and thus open the path for developing safe brain-computer interfaces for non-communicating patients.
1 Introduction
Invasive BCIs can restore sentence-level communication but require risky, difficult-to-maintain implants. Brain2Qwerty combines higher-SNR MEG, natural typing, and deep learning to decode text production non-invasively.
- Invasive neuroprostheses can produce full sentences but require neurosurgery, limiting their scalability for many non- or poorly responsive patients.Reported risks include brain hemorrhage and infection, while maintaining cortical implants over time remains challenging.
- Current non-invasive BCIs generally use EEG with demanding attention or motor-imagery tasks, yet decoding performance remains moderate.These limitations leave existing non-invasive methods short of a fast and reliable BCI.
- MEG offers higher signal-to-noise ratio than EEG, while recent deep-learning models have improved natural-language reconstruction from brain signals.Together, these advances motivate decoding language production from non-invasive recordings.
- Brain2Qwerty decodes briefly memorized sentences typed by 35 participants using EEG or MEG recordings.The study evaluates EEG from 20 participants and MEG from 20 participants, while focusing on decoding rather than the neural mechanisms of language production.
2 Results
Linear decoding confirms that the typing protocol elicits expected motor-related brain responses. MEG provides stronger hand-related decoding than EEG in the reported analysis.
- 2.1 Linear Decoding: 74±1.3% MEG accuracy is reached when classifying left- versus right-handed key presses at the peak time of 40 ms after pressing.A subject-specific linear ridge classifier was evaluated at each time sample relative to key presses.
- 2.1 Linear Decoding: The evoked-response topographies are typical of cortical motor activity, supporting the expected neural response to left- and right-handed typing.The protocol therefore provides a motor-related signal that can be classified around key presses.
- 2.1 Linear Decoding: Figure 2 compares sensor responses, time-resolved hand and character classifiers, baseline models, and Brain2Qwerty ablations using HER and CER.Stars mark significant decoding scores, and points represent participant-level averages in the model comparisons.
2.2 Brain2Qwerty Performance
Brain2Qwerty decodes individual characters from non-invasive M/EEG signals, with substantially lower error using MEG than EEG. Performance varies across participants, reaching 19% CER for the best MEG participant.
- 2.2 Brain2Qwerty Performance: Brain2Qwerty is trained to decode individual characters from M/EEG signals and evaluated using both hand-error-rate and character-error-rate.
- 2.2 Brain2Qwerty Performance: 32±0.6% CER with MEG versus 67±1.5% CER with EEG demonstrates a substantial recording-device difference.The difference is statistically significant at p<10^-8.
- 2.2 Brain2Qwerty Performance: 19±1.1% CER is achieved by the best MEG participant, while the worst MEG participant reaches 45±1.2% CER.For EEG, the best and worst participants reach 61±2.0% and 71±2.3% CER, respectively.
2.3 Comparing Brain2Qwerty to Baseline Models
Brain2Qwerty outperforms classic linear and EEGNet baselines in character decoding, although EEGNet itself is stronger than the linear model in several comparisons.
- 2.3 Comparing Brain2Qwerty to Baseline Models: Brain2Qwerty achieves a 1.14-fold improvement in EEG CER over EEGNet, with p<10^-5.EEGNet remains less effective than Brain2Qwerty in the reported comparison.
- 2.3 Comparing Brain2Qwerty to Baseline Models: EEGNet outperforms the linear model for MEG HER and CER, but for EEG this advantage is reported only for HER.The corresponding significance values are p=0.008 and p<10^-4 for MEG, and p=0.03 for EEG HER.
2.4 Brain2Qwerty Ablations
Ablation tests show that sentence-level contextualization and language-model regularities progressively improve Brain2Qwerty’s character decoding.
- 2.4 Brain2Qwerty Ablations: The convolutional module alone outperforms EEGNet on both EEG and MEG for hands- and character-error rates.
- 2.4 Brain2Qwerty Ablations: The transformer and language-model components each improve character decoding beyond the convolutional module alone.The transformer improves CER for both EEG and MEG, while the language model provides an additional CER improvement.
- 2.4 Brain2Qwerty Ablations: The ablation analysis attributes gains to sentence-level contextualization and natural-language statistical regularities.
- 2.4 Brain2Qwerty Ablations: Figure 3 summarizes sentence-level performance for best, median, and worst MEG subjects using per-sentence CER and example predictions.
2.5 Analyses of Decoded Sentences
MEG decoding can produce fully correct sentences outside the training set, including cases where the model corrects participants’ typographical errors.
- 2.5 Analyses of Decoded Sentences: Several MEG sentences are perfectly decoded, including examples outside the training set.In one example, the model correctly decoded “el beneficio supera los riesgos” despite substantial typing errors by the participant.
2.6 Impact of Word Type and Frequency
Decoding varies with linguistic and character frequency: frequent items are easier, rare or unseen words are harder, and more training data improves CER.
- 2.6 Impact of Word Type and Frequency: 17±1.9% CER is achieved for determiners, while out-of-vocabulary words remain decodable but reach 68±2.1% CER.Frequent words are decoded better than rare words, and all tested part-of-speech categories perform above chance.
- 2.6 Impact of Word Type and Frequency: R=0.85, p < 10−8 links character frequency with decoding accuracy, while rare Spanish characters are not decoded above chance.The rare characters “z,” “k,” and “w” each account for less than 0.1% of the sentence characters.
- 2.6 Impact of Word Type and Frequency: The supplied sentence examples compare decoded outputs across EEG and MEG or model ablations using color-coded correct characters, mistakes, and typing errors.
- 2.6 Impact of Word Type and Frequency: R=0.93, p<10−7 shows that CER decreases as the amount of training data increases.
- 2.6 Impact of Word Type and Frequency: Figure 4 evaluates CER across part-of-speech categories, word and character frequency, out-of-vocabulary words, and training-set recording time.
2.7 Impact of Keyboard Layout
Keyboard distance predicts character confusions, supporting a relationship between decoding errors and the QWERTY motor layout.
- 2.7 Impact of Keyboard Layout: R=0.73, p=0.02 indicates a strong correlation between physical keyboard distance and character confusion rate.The analysis tests whether decoding errors relate to the QWERTY layout expected from motor-cortex activity.
2.8 Impact of Typing Errors
Typing errors are associated with longer inter-key intervals and substantially worse character decoding, even when sentence context is minimized.
- 114±12 ms inter-key intervals for incorrect characters versus 50±7 ms for correct characters suggest hesitation or monitoring during typing errors.The difference is statistically significant at p=10−7.
- 65% CER for incorrectly typed characters versus 38% for correctly typed characters indicates poorer decoding when motor execution is inaccurate.The same pattern remains without transformer-based contextualization: 71% versus 52% CER in the Convolutional Module.
- The error analysis suggests that decoding performance diminishes when motor processes are inaccurately executed.
3 Discussion
Brain2Qwerty demonstrates sentence decoding from non-invasive brain recordings, with MEG substantially outperforming EEG and deep-learning contextualization improving character decoding. The method narrows but does not eliminate the gap with invasive BCIs, while real-time operation and clinical generalization remain unresolved.
- 3 Discussion: Transformer-based sentence context and a pretrained language model substantially improve character decoding beyond the convolutional architecture alone.Adding the transformer improves CER, while the language model provides an additional CER improvement for both EEG and MEG.
- 3 Discussion: Brain2Qwerty uses a comparatively easier typing task than traditional P300-speller, SSVEP, and fMRI-localizer protocols, which relied on handcrafted processing and shallow classifiers.
- 3 Discussion: The results narrow the gap between non-invasive and invasive BCIs, but invasive systems still report lower error rates and faster communication.Reported invasive benchmarks include 15.2% CER at 79 words per minute and below 6% CER at 90 characters per minute.
- 3 Discussion: Clinical adaptation is limited because the current system is not real time, requires MEG segments aligned to keystrokes, and was tested only with healthy participants using supervised typing.Locked-in individuals cannot perform the keyboard typing task, motivating imagination-based tasks or stronger cross-participant generalization.
- 3 Discussion: Current MEG systems are not wearable, although optically pumped magnetometers may address this constraint.
- 3 Discussion: The study positions non-invasive decoding as a step toward safer and more accessible BCIs for people who have lost communication abilities.
4 Methods
The study records brain activity while participants type briefly memorized sentences and decodes each keystroke with Brain2Qwerty, a three-stage neural architecture. The protocol combines non-invasive EEG or MEG recordings, controlled sentence splits, and character-level evaluation.
- 4.1 Experimental Protocol: The experiment recruited 35 healthy, right-handed, skilled typists who typed memorized sentences without visual feedback.Participants were native Spanish speakers and had demonstrated at least 80% typing accuracy.
- 4.1 Experimental Protocol: EEG and MEG were recorded at 1 kHz, then filtered, resampled to 50 Hz, and segmented into windows around each key press.Each window spans from 0.2 s before to 0.3 s after the key press, followed by baseline correction and robust scaling.
- 4.1 Experimental Protocol: Sentences were divided into maximally diverse 80% train, 10% validation, and 10% test splits to limit sentence memorization.A clustering-based splitter grouped similar sentences before assigning them to partitions.
- 4.2 Decoder: Brain2Qwerty predicts each keystroke from 0.5 s M/EEG windows across 29 character classes.The classes cover the Latin alphabet plus spaces, numbers, and other special characters.
- 4.2.1 Architecture: The architecture successively applies a convolutional module, a sentence-level transformer, and a pretrained character-level language model.The convolutional module encodes sensor signals, the transformer exploits sentence context, and the language model regularizes predictions using Spanish character statistics.
- 4.2.1 Architecture: The convolutional and transformer modules are trained jointly end-to-end with cross-entropy loss across subjects, using the same hyperparameters for EEG and MEG.The model contains approximately 400 million parameters and uses AdamW training with early stopping.