Paper deep dive
Phoneme- vs. Character-Level Targets and Selective State-Space Models for Intracortical Brain-to-Text
Lucas Zamora Vera, Jose A. Gonzalez-Lopez
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/4/2026, 11:05:32 AM
Summary
This paper evaluates intracortical brain-to-text decoding systems by comparing recurrent neural networks (GRU) against selective state-space models (Mamba variants) and analyzing phonetic versus character-level output targets. Using the Brain-to-Text '25 benchmark, the study finds that the GRU baseline with phonetic targets achieves the lowest Word Error Rate (21.19%), outperforming Mamba-based hybrids and character-level decoders. Error analysis reveals that phonetic targets expose articulatory confusions, while character targets expose lexical and word-boundary errors.
Entities (10)
Relation Signals (6)
GRU → achievesbestperformanceon → Phonetic Target
confidence 95% · The recurrent baseline remains strongest: the best phonetic GRU reaches 12.62% PER and 21.19% WER
Brain-to-Text '25 → usedby → GRU
confidence 95% · On the public Brain-to-Text '25 benchmark, we study... GRU
Phonetic Target → exposeserrormode → Articulatory Confusions
confidence 92% · articulatory-like phoneme confusions vs. lexical and word-boundary errors
Character Target → exposeserrormode → Lexical Errors
confidence 92% · lexical and word-boundary errors
Mamba → iscompetitivebutnotsuperiorto → GRU
confidence 90% · The Mamba hybrid is competitive but does not surpass it.
ConvMambaGRU → combines → ConvMamba
confidence 88% · adds a final unidirectional GRU after the Mamba blocks
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:State-of-the-art intracortical brain-to-text systems pair a neural-sequence phone decoder with an external language model. Two design axes remain underexplored: whether selective state-space models (Mamba) improve on recurrent decoders, and how the output target (phonetic vs.\ character) interacts with that choice. On the public Brain-to-Text '25 benchmark, we study a controlled 2x2 grid (GRU vs.\ hybrid Mamba decoder; phonetic vs.\ character targets) trained with a CTC objective under one reproducible protocol. The recurrent baseline remains strongest: the best phonetic GRU reaches 12.62\% PER and 21.19\% WER, while the best textual GRU after LM rescoring reaches 13.39\% CER and 26.28\% WER. The Mamba hybrid is competitive but does not surpass it. Ablations isolate architectural contributions, and error analysis shows representation-dependent failures: articulatory-like phoneme confusions vs.\ lexical and word-boundary errors.
Tags
Links
- Source: https://arxiv.org/abs/2607.26751v1
- Canonical: https://arxiv.org/abs/2607.26751v1
Trouble viewing inline? Open PDF directly →
Full Text
35,045 characters extracted from source content.
Expand or collapse full text
Phoneme- vs. Character-Level Targets and Selective State-Space Models for Intracortical Brain-to-Text Lucas Zamora Vera 1 , Jose A. Gonzalez-Lopez ID 2,3,1 1 Universitat Oberta de Catalunya (UOC), Spain 2 Dpt. of Signal Theory, Telematics and Communications, University of Granada, Spain 3 Research Centre for Information and Communication Technologies (CITIC-UGR), University of Granada, Spain lzamorav@uoc.edu, joseangl@ugr.es Abstract State-of-the-art intracortical brain-to-text systems pair a neural- sequence phone decoder with an external language model. Two design axes remain underexplored: whether selective state- space models (Mamba) improve on recurrent decoders, and how the output target (phonetic vs. character) interacts with that choice. On the public Brain-to-Text ’25 benchmark, we study a controlled 2x2 grid (GRU vs. hybrid Mamba decoder; phonetic vs. character targets) trained with a CTC objective un- der one reproducible protocol. The recurrent baseline remains strongest: the best phonetic GRU reaches 12.62% PER and 21.19% WER, while the best textual GRU after LM rescoring reaches 13.39% CER and 26.28% WER. The Mamba hybrid is competitive but does not surpass it. Ablations isolate archi- tectural contributions, and error analysis shows representation- dependent failures: articulatory-like phoneme confusions vs. lexical and word-boundary errors. Index Terms: brain-computer interface, speech neuroprosthe- sis, intracortical speech decoding, state-space models, brain-to- text 1. Introduction People with amyotrophic lateral sclerosis (ALS), brainstem stroke or locked-in syndrome can lose intelligible speech while retaining cognition, and the recovery of a reliable communica- tion channel substantially improves autonomy and social par- ticipation [1, 2, 3]. Brain-computer interfaces (BCIs) aim to re- store this channel by translating neural activity into communica- tive output [2], a paradigm first established for motor and cursor control in people with paralysis [4, 5, 6]. Speech decoding has been pursued across recording modalities: non-invasive record- ings support coarse perception decoding [7, 8], electrocorticog- raphy (ECoG) enables phrase- and sentence-level decoding and synthesis [9, 10, 11, 12, 13], and intracortical microelectrode arrays offer the highest spatiotemporal resolution, driving the most accurate recent speech decoders [1, 14, 15]. The neuro- functional basis for this is that ventral sensorimotor cortex re- tains rich articulatory information even when motor output is lost [1, 16, 17]. A brain-to-text system casts speech decoding as a sequence-to-sequence problem with no explicit alignment be- tween the multichannel neural signal and the target symbol sequence, which is naturally handled by the Connectionist Temporal Classification (CTC) objective [18]. The dominant paradigm is two-stage: a recurrent decoder maps neural activity to sub-lexical units (typically phonemes), and a language model (LM) reconstructs the most plausible sentence [14, 15, 1, 19]. Landmark intracortical results, including high-rate handwriting decoding [20], real-time decoding in anarthria [21], a large- vocabulary speech neuroprosthesis [14], a high-performance avatar-and-speech system [22], generalizable spelling [23], in- stantaneous voice synthesis [24], and a rapidly calibrating sys- tem reaching sub-5% word error rates over extended use [15], established the clinical relevance of this line of work, approach- ing the accuracy of modern automatic speech recognition (ASR) systems [25, 26]. The recent Brain-to-Text ’24 benchmark re- ported that the largest gains in decoding performance came from better training, ensembling, and LM rescoring rather than from replacing the recurrent decoding backbone, with Trans- formers showing no clear advantage [27, 28]. Two design choices, however, remain comparatively under- studied. First, selective state-space models (SSMs), specifi- cally Mamba [29], offer a linear-time recurrent formulation with input-dependent selectivity that is attractive for long, noisy neu- ral sequences. SSMs have shown promise in ASR, both as stan- dalone speech encoders [30, 31] and in convolution-augmented forms [32, 33] that echo the success of attention-based se- quence models [34, 35], the Conformer [36] and its efficient variants [37, 38]. Whether they help for intracortical decod- ing is an open question. Second, most systems commit to a phonetic target, yet a direct character-level target removes the dependency on a pronunciation lexicon and brings the decoder closer to the final output. A recent end-to-end character-level Conformer study on the same dataset showed that meaningful direct character decoding is possible without an external lan- guage model, but also found that dominant errors arise from incorrect word-boundary segmentation [39]. However, a con- trolled comparison of the target representation crossed with the decoder architecture remains missing. This paper asks: do selective state-space models improve on recurrent decoders for intracortical brain-to-text, and how does the output target (phoneme vs. character) interact with that choice? We make three contributions: (i) a controlled 2× 2 study (GRU vs. a hybrid ConvMambaGRU; phonetic vs. tex- tual targets) under one reproducible CTC pipeline on the pub- lic Brain-to-Text ’25 benchmark; (i) to our knowledge, the first adaptation of a ConvMamba-style backbone to intracorti- cal brain-to-text, with ablations isolating its convolutional front- end and recurrent refinement stage; and (i) a representation- aware error analysis contrasting phoneme-level and character- level failure modes. Our finding is that the recurrent baseline remains the strongest decoder and the phonetic two-stage vari- ant gives the lowest WER; the SSM hybrid is competitive but does not surpass it, consistent with the data-limited regime of single-participant intracortical decoding [27]. arXiv:2607.26751v1 [cs.CL] 29 Jul 2026 Figure 1: Overview of the experimental pipeline. Intracortical neural activity is preprocessed and passed to a neural decoder. We study two factors: the sequential backbone (GRU, Mamba variants and ConvMambaGRU) and the target representation (phonetic vs. textual). The resulting sequences are decoded with CTC and optionally refined through beam search and language- model rescoring to obtain the final sentence-level transcription. 2. Methods 2.1. General procedure Fig. 1 summarizes the pipeline shared by all configurations. A trial of intracortical neural features is first passed through a per-session adaptation layer that normalizes day-to-day distri- bution shifts, then through temporal patching that shortens the effective sequence. A sequential core (the recurrent baseline or one of the state-space variants) produces a per-step distribution over an output vocabulary, trained end-to-end with a CTC ob- jective [18]. The two design axes we study sit at the two ends of this core: the target representation (phonemes vs. characters), which sets the output vocabulary, and the sequential backbone. A final linguistic-decoding stage (greedy or beam search) op- tionally rescored by a language model turns the logits into the output sentence; in the phonetic branch this stage also performs the phoneme-to-text conversion. Keeping every stage fixed ex- cept the backbone and the target lets us attribute differences to those two factors. 2.2. Dataset and signal processing We use the public Brain-to-Text ’25 benchmark, released with the speech neuroprosthesis of Card et al. [15] and used by the Brain-to-Text benchmark line of work [27]. The data come from a single participant with ALS, severe dysarthria and tetra- paresis but intact cognition, implanted with four 64-channel microelectrode arrays (256 intracortical electrodes) in the left ventral precentral gyrus [15]. Each trial is a prompted speech attempt with an aligned transcript, encoded as a variable-length matrix X ∈R T×512 , where T is the number of 20 ms time bins and the 512 features are two measures per electrode (binned threshold-crossing counts and spike-band power). We use the official partitions: 8,072 training and 1,426 validation trials; the 1,450 test trials are excluded as their transcripts are not public, so all comparisons are on validation. Signal processing follows the baseline pipeline: Gaussian temporal smoothing and per-session normalization, the latter reducing inter-session distribution shift from electrode drift and impedance changes, a well-documented source of non- stationarity in chronic recordings [40, 41] addressed by prior work through recalibration [42] and manifold alignment [43]. During training only, we apply lightweight SpecAugment-style augmentations [44] as regularization: per-channel gain, white noise, constant offsets, random-walk noise, and random start crops. 2.3. Problem formulation Given the trial representation X ∈R T×512 , the decoder emits a sequence of per-frame logits over an output vocabulary of ei- ther phonemes or characters. Because there is no explicit align- ment between X and the target y, we train with CTC [18]: the probability of y given X is the sum over all alignments π that collapse to y, P (y | X) = P π∈B −1 (y) P (π | X), and the loss isL CTC = − logP (y | X). At inference, CTC collapsing removes blanks and consecutive repeats. Variable-length trials are batched with zero-padding while preserving true lengths, so padding does not affect the loss. 2.4. Decoder architectures All backbones share the pipeline of Section 2.1 (per-session adaptation, temporal patching, sequential core, linear projec- tion), so performance differences are attributable to the core and the target representation. GRU baseline. A reproducible recurrent decoder aligned with recent brain-to-text systems: a per-session adaptation layer, temporal patching to shorten the effective sequence, a stack of unidirectional GRU layers [45], and a final projection. State-space models. As an alternative core we consider Mamba [29], a selective state-space model (SSM). A continu- ous SSM maps an input signal x(t) to an output y(t) through a latent state h(t): h ′ (t) = Ah(t) + B x(t),y(t) = C h(t),(1) where A governs the internal state dynamics, B how the input drives the state, and C how the state is read out. Discretizing in time yields a linear recurrence, h t = ̄ A t h t−1 + ̄ B t x t ,y t = C t h t ,(2) which scales linearly with sequence length, unlike the quadratic cost of self-attention, making it attractive for the long, noisy neural sequences here. The key idea of Mamba is selectiv- ity: the parameters ( ̄ B t ,C t ) and the discretization step be- come functions of the current input, so the model can choose which information to keep, update or discard at each step [29]. This input-dependent gating is well suited to neural recordings, where not every bin is equally informative. On top of this core we evaluate three increasingly struc- tured variants. (i) Mamba applies residual Mamba blocks di- rectly after patching, isolating the contribution of the SSM core alone [29, 30]. (i) ConvMamba prepends a residual Conv1D front-end before the Mamba blocks. The motivation, follow- ing convolution-augmented SSMs that replace Conformer self- attention with Mamba layers [32, 33], is that local depth- wise convolutions efficiently capture short-range articulatory- like patterns that complement the longer-range dependencies Table 1: Evaluated backbones, instantiated as phonetic (P) and textual (T) variants. Output classes include the CTC blank. BackboneParams (P) Params (T) Temporal core GRU (baseline)44.3M44.3M5× GRU Mamba36.2M36.2M5× Mamba ConvMamba50.6M50.6MConv1D + 5× Mamba ConvMambaGRU50.7M54.2MConv1D + 5× Mamba + GRU modeled by the SSM, which is particularly useful when train- ing data are scarce [31]. (i) ConvMambaGRU, our proposed variant, adds a final unidirectional GRU after the Mamba blocks as a temporal-refinement stage before projection, combining lo- cal convolutional context, linear-time selective state-space mod- eling, and recurrent refinement. The temporal patching itself follows patch-based tokenization from vision transformers [46], reducing sequence length while preserving local context. Each backbone is instantiated as a phonetic variant (ARPA- bet targets, later converted to text by the LM stage) and a textual variant (character targets, decoded directly). Table 1 summa- rizes the configurations. 2.5. Linguistic decoding and rescoring From the T × P logit matrix we obtain hypotheses by greedy CTC decoding and by beam search. On the beam-search hy- potheses we apply LM rescoring: the final score combines the neural (CTC) score with the LM plausibility, Score(y) = Score CTC (y) + α Score LM (y), where α weights the linguistic component. In the phonetic variant the LM is essential, as it performs the phoneme-to-text conversion; in the textual variant it refines an already character-level hypothesis. We compare a generalist Transformer LM (GPT-2 variants [47]) and a domain- trained n-gram LM (KenLM [48]). For textual decoding, we use CTC beam search with beam width 100 and retain the top 100 hypotheses for LM rescoring. The LM weight α is selected on the validation set for each decoder–LM pair according to the lowest WER. 3. Experimental setup The dataset, splits and signal processing are described in Sec- tion 2.2; all results are reported on the 1,426-trial validation set. Metrics. Phoneme Error Rate (PER) for the phonetic de- coder; Character Error Rate (CER) and Word Error Rate (WER) for the final text are used to evaluate the proposed systems. Implementation. Models are implemented in PyTorch and trained on a single NVIDIA 3060 GPU with 12 GB of VRAM. For reproducibility, configuration files, checkpoints, validation predictions and metric logs are retained for every run. Random seeds are fixed for each experiment. Code and trained-model artifacts will be released with the final version of the paper. Statistical reporting. For the phonetic backbone compari- son, each model is trained with five random seeds, and we report mean and standard deviation of PER across runs. The textual configurations are reported as single-run validation results. For the utterance-level paired analysis, we focus on the GRU base- line and ConvMambaGRU, the best-performing Mamba-based variant in the multi-seed phonetic comparison. Paired differ- ences in PER and WER are computed on the same 1,426 valida- tion trials and assessed with the Wilcoxon signed-rank test [49], following recommended practice for comparing learning sys- tems [50]. We report both two-sided tests and directional tests in the observed direction. Table 2: Initial backbone exploration on the phonetic validation task. Values are PER (%) over five independent runs; the last column reports mean± standard deviation across seeds. BackboneR1 R2 R3 R4 R5 Mean± std GRU12.62 13.15 12.79 12.55 13.68 12.96± 0.47 Mamba15.37 14.93 14.99 15.03 14.57 14.98± 0.29 ConvMamba17.47 17.91 16.89 17.50 17.21 17.40± 0.38 MambaGRU14.05 13.95 14.06 13.78 22.41 15.65± 3.78 ConvMambaGRU 13.13 13.16 13.67 13.68 13.58 13.44± 0.28 Table 3: Main validation results. Phonetic models use the best runs from Table 2 with the official 1-gram OpenWebText base- line LM. Textual models use GPT-2 XL rescoring with the best validation α per decoder. PER/CER are intermediate metrics; WER is the final text metric. TargetModelPER/CER (%) WER (%) PhoneticGRU12.62 (PER)21.19 PhoneticConvMambaGRU13.13 (PER)24.11 TextualGRU13.39 (CER)26.28 TextualConvMambaGRU15.29 (CER)29.86 4. Results 4.1. Backbone performance on the phonetic task We first explore how the sequential backbones behave on the phonetic task, using validation PER over five independent runs (Table 2) to identify the strongest SSM variant to carry for- ward. The GRU baseline obtains the lowest mean PER and is the most stable model. Among the SSM backbones, perfor- mance improves as convolutional preprocessing and recurrent refinement are added to the bare Mamba core: Mamba-only and ConvMamba lag clearly behind, MambaGRU is competitive on average but unstable (one diverging run inflates its variance), and ConvMambaGRU is the strongest and most consistent SSM configuration, narrowing the gap to the recurrent baseline. This indicates that selective SSM blocks alone do not improve intra- cortical phoneme decoding in this data-limited regime, becom- ing competitive only when combined with both local convolu- tional context and recurrent refinement, consistent with bench- mark observations that recurrent baselines remain hard to beat with limited data [27, 28]. We therefore adopt ConvMamba- GRU as the representative SSM model. 4.2. Main comparison For the main comparison, final WER for the phonetic systems is computed with the official baseline language model from the Brain-to-Text benchmark pipeline: a 1-gram OpenWebText lan- guage model with silence-token handling. For the textual sys- tems, final WER is computed after character-level beam search and GPT-2 XL rescoring, using the best validation weight for each decoder (α = 1.5 for GRU and α = 2.0 for ConvMamba- GRU; Table 4). We compare ConvMambaGRU, the strongest SSM back- bone, against the GRU baseline across both target represen- tations (Table 3); phonetic values correspond to the best run from Table 2, textual models are single-run results. The GRU is the strongest decoder in both representations, and the two-stage phonetic models outperform their textual counterparts (21.19% vs. 26.28% WER). ConvMambaGRU is competitive on PER Table 4: Language-model rescoring for the textual target on the validation set. For each neural decoder and language model, α is selected according to the lowest validation WER. ModelLMCER (%) WER (%)α GRUGPT-2 XL13.3926.281.5 GRUGPT-213.5126.442.0 GRUKenLM 5-gram14.1927.805.0 GRUKenLM 3-gram14.2027.884.5 ConvMambaGRU GPT-2 XL15.2929.862.0 ConvMambaGRU KenLM 3-gram16.1330.947.5 ConvMambaGRU KenLM 5-gram16.1330.957.5 ConvMambaGRU GPT-215.8531.100.8 but does not surpass the GRU on final WER. We speculate that the phonetic advantage arises because the neural signal, originating in speech-motor cortex, encodes articulatory struc- ture more directly than orthographic characters, placing a pho- netic target closer to the information present in the recordings. The textual variant nonetheless retains practical appeal: de- coding characters directly removes the need for a pronuncia- tion lexicon and phoneme-to-text conversion, although in this study its best performance still requires GPT-2 XL rescoring. Finally, small intermediate differences propagate non-linearly: because the LM resolves hypotheses at the word level, a sub- 1-point PER change can shift lexical and word-boundary deci- sions into a ∼3-point WER gap, confirming that final perfor- mance is the joint product of the neural decoder and the linguis- tic stage [27, 28]. 4.3. Effect of the language model We next examine the linguistic decoding stage, sweeping the rescoring weight α described in Section 2.5 for both decoders and several language models on the textual target. Table 4 re- ports, for each neural decoder and LM, the configuration with the lowest validation WER. The domain n-gram (KenLM) and the generalist GPT-2 variants trade off differently; rescoring consistently improves over greedy decoding, but cannot fully compensate for a weaker neural decoder. 4.4. Error analysis Errors grow with sentence length in both variants, and there is appreciable inter-session variability, consistent with the ex- ploratory analysis of signal stability. Crucially, the two repre- sentations fail differently, and the target determines how errors become visible. In the phonetic variant, substitutions are not uniform across the inventory. Table 5 lists the most frequent phoneme sub- stitutions, in both ARPAbet and IPA, aggregated over the five GRU phonetic validation runs. The dominant confusions are reciprocal and phonetically plausible: the most frequent pair is D /d/→T /t/, followed by its reverse, and the same voicing contrast appears for Z /z/→S /s/ and S /s/→Z /z/. Vowel confusions between neighbouring qualities are also prominent (e.g. AH /@/↔IH /I/, IH /I/→EH /E/, IY /i/→IH /I/). These patterns indicate systematic difficulty discriminating articulato- rily and acoustically close phones rather than random sequence- level errors. In the textual variant, the most frequent errors concentrate on short, high-frequency words (Table 6), mainly function- word and near-homophone substitutions such as the→they, Table 5: Most frequent phoneme substitutions for the phonetic variant on the validation set. Counts are aggregated over the five GRU phonetic validation runs. Reference Hypothesis Count D /d/T /t/631 Z /z/S /s/354 AH /@/IH /I/349 T /t/D /d/339 IH /I/AH /@/272 S /s/Z /z/267 IH /I/EH /E/239 IY /i/IH /I/233 Table 6: Most frequent word-level confusions for the textual variant on the validation set, using beam search with language- model rescoring. Reference Hypothesis Count thethey9 doto8 theirthere7 tooto6 hasas5 init5 do→to, too→to, and their→there. Even after beam search with LM rescoring, the decoder can still confuse brief lexi- cal items whose acoustic, orthographic or contextual evidence is weak. This contrasts with the phonetic analysis, where er- rors surface as local phone substitutions, and is complemen- tary to a recent character-level Conformer study on the same data, which reports dominant character-level errors involving the space token [39]. Together, the two analyses show that the target representation governs how decoding errors become vis- ible: phonetic targets expose articulatory and acoustic confu- sions, whereas textual targets expose lexical and segmentation- sensitive errors. 5. Conclusion We presented a controlled study of intracortical brain-to-text de- coding crossing the target representation (phonetic vs. textual) with the decoder architecture (recurrent vs. a selective state- space hybrid) under one reproducible CTC pipeline. The re- current baseline remains the strongest decoder, and the pho- netic two-stage variant gives the lowest WER. In particular, the GRU neural decoder obtains lower error than ConvMam- baGRU in PER (12.62% vs. 13.13%, a difference of -0.51 per- centage points; Wilcoxon two-sided p = 0.023; directional p = 0.0115) and in WER (21.19% vs. 24.11%, a difference of -2.92 percentage points; Wilcoxon two-sided p = 3.7× 10 −10 ; directional p = 1.85× 10 −10 ). Final performance is therefore the joint product of the neural decoder and the linguistic stage, and the dominant error modes are representation-specific. Fu- ture work will explore three directions aligned with recent re- search on low-resource intracortical decoding [27, 28]: richer target representations, stronger training and rescoring strategies, and more task-specific adaptation of SSM architectures to neu- ral signals, rather than simply substituting the recurrent back- bone. 6. Acknowledgments This work was supported by grants PID2022-141378OB- C22andAIA2025-163317-C32fundedbyMI- CIU/AEI/10.13039/501100011033 and ERDF/EU. 7. Generative AI Use Disclosure During the preparation of this work, the authors used genera- tive AI tools for language editing and to assist with code devel- opment. These tools were not used to generate scientific con- tent, design the methodology, or interpret the results. After us- ing these tools, the authors reviewed and edited the content as needed and take full responsibility for the content of the publi- cation. 8. References [1] A. B. Silva, K. T. Littlejohn, J. R. Liu, D. A. Moses, and E. F. Chang, “The speech neuroprosthesis,” Nature Reviews Neuro- science, vol. 25, no. 7, p. 473–492, 2024. [2] J. R. Wolpaw, N. Birbaumer, D. J. McFarland, G. Pfurtscheller, and T. M. Vaughan, “Brain-computer interfaces for communica- tion and control,” Clinical Neurophysiology, vol. 113, no. 6, p. 767–791, 2002. [3] E. F. Chang, “Brain-computer interfaces for restoring communi- cation,” New England Journal of Medicine, vol. 391, no. 7, p. 654–657, 2024. [4] L. R. Hochberg, D. Bacher, B. Jarosiewicz, N. Y. Masse, J. D. Simeral, J. Vogel, S. Haddadin, J. Liu, S. S. Cash, P. van der Smagt, and J. P. Donoghue, “Reach and grasp by people with tetraplegia using a neurally controlled robotic arm,” Nature, vol. 485, no. 7398, p. 372–375, 2012. [5] C. Pandarinath, P. Nuyujukian, C. H. Blabe, B. L. Sorice, J. Saab, F. R. Willett, L. R. Hochberg, K. V. Shenoy, and J. M. Hender- son, “High performance communication by people with paralysis using an intracortical brain-computer interface,” eLife, vol. 6, p. e18554, 2017. [6] M. J. Vansteensel, E. G. M. Pels, M. G. Bleichner, M. P. Branco, T. Denison, Z. V. Freudenburg, P. Gosselaar, S. Leinders, T. H. Ottens, M. A. van den Boom, P. C. van Rijen, E. J. Aarnoutse, and N. F. Ramsey, “Fully implanted brain-computer interface in a locked-in patient with ALS,” New England Journal of Medicine, vol. 375, no. 21, p. 2060–2066, 2016. [7] J. Tang, A. LeBel, S. Jain, and A. G. Huth, “Semantic reconstruc- tion of continuous language from non-invasive brain recordings,” Nature Neuroscience, vol. 26, no. 5, p. 858–866, 2023. [8] A. D ́ efossez, C. Caucheteux, J. Rapin, O. Kabeli, and J.-R. King, “Decoding speech perception from non-invasive brain record- ings,” Nature Machine Intelligence, vol. 5, no. 10, p. 1097–1107, 2023. [9] C. Herff, D. Heger, A. de Pesters, D. Telaar, P. Brunner, G. Schalk, and T. Schultz, “Brain-to-text: Decoding spoken phrases from phone representations in the brain,” Frontiers in Neuroscience, vol. 9, p. 217, 2015. [10] J. G. Makin, D. A. Moses, and E. F. Chang, “Machine translation of cortical activity to text with an encoder–decoder framework,” Nature Neuroscience, vol. 23, no. 4, p. 575–582, 2020. [11] G. K. Anumanchipalli, J. Chartier, and E. F. Chang, “Speech syn- thesis from neural decoding of spoken sentences,” Nature, vol. 568, no. 7753, p. 493–498, 2019. [12] M. Angrick, C. Herff, E. Mugler, M. C. Tate, M. W. Slutzky, D. J. Krusienski, and T. Schultz, “Speech synthesis from ECoG using densely connected 3D convolutional neural networks,” Journal of Neural Engineering, vol. 16, no. 3, p. 036019, 2019. [13] S. Martin, P. Brunner, C. Holdgraf, H.-J. Heinze, N. E. Crone, J. Rieger, G. Schalk, R. T. Knight, and B. N. Pasley, “Decoding spectrotemporal features of overt and covert speech from the hu- man cortex,” Frontiers in Neuroengineering, vol. 7, p. 14, 2014. [14] F. R. Willett, E. M. Kunz, C. Fan, D. T. Avansino, G. H. Wil- son, E. Y. Choi, F. Kamdar, L. R. Hochberg, J. M. Henderson, and K. V. Shenoy, “A high-performance speech neuroprosthesis,” Nature, vol. 620, no. 7976, p. 1031–1036, 2023. [15] N. S. Card, M. Wairagkar, C. Iacobacci, P. Bhatt, T. Singer-Clark, F. R. Willett, K. C. Ames, J. Liu, P. Rezaii, L. R. Hochberg, J. M. Henderson, K. V. Shenoy, and D. M. Brandman, “An accurate and rapidly calibrating speech neuroprosthesis,” New England Journal of Medicine, vol. 391, no. 7, p. 609–618, 2024. [16] G. Hickok and D. Poeppel, “The cortical organization of speech processing,” Nature Reviews Neuroscience, vol. 8, no. 5, p. 393– 402, 2007. [17] F. Pulverm ̈ uller, “Neural reuse of action perception circuits for language, concepts and communication,” Progress in Neurobiol- ogy, vol. 160, p. 1–44, 2018. [18] A. Graves, S. Fern ́ andez, F. Gomez, and J. Schmidhuber, “Con- nectionist temporal classification: Labelling unsegmented se- quence data with recurrent neural networks,” in Proc. Interna- tional Conference on Machine Learning (ICML), 2006, p. 369– 376. [19] J. S. Brumberg, A. Nieto-Castanon, P. R. Kennedy, and F. H. Guenther, “Brain-computer interfaces for speech communica- tion,” Speech Communication, vol. 52, no. 4, p. 367–379, 2010. [20] F. R. Willett, D. T. Avansino, L. R. Hochberg, J. M. Henderson, and K. V. Shenoy, “High-performance brain-to-text communica- tion via handwriting,” Nature, vol. 593, no. 7858, p. 249–254, 2021. [21] D. A. Moses, S. L. Metzger, J. R. Liu, G. K. Anumanchipalli, J. G. Makin, P. F. Sun, J. Chartier, M. E. Dougherty, P. M. Liu, G. M. Abrams, A. Tu-Chan, K. Ganguly, and E. F. Chang, “Neuropros- thesis for decoding speech in a paralyzed person with anarthria,” New England Journal of Medicine, vol. 385, no. 3, p. 217–227, 2021. [22] S. L. Metzger, K. T. Littlejohn, A. B. Silva, D. A. Moses, M. P. Seaton, R. Wang, M. E. Dougherty, J. R. Liu, P. Wu, M. A. Berger, I. Zhuravleva, A. Tu-Chan, K. Ganguly, G. K. Anumanchipalli, and E. F. Chang, “A high-performance neuroprosthesis for speech decoding and avatar control,” Nature, vol. 620, no. 7976, p. 1037–1046, 2023. [23] S. L. Metzger, J. R. Liu, D. A. Moses, M. E. Dougherty, M. P. Seaton, K. T. Littlejohn, J. Chartier, G. K. Anumanchipalli, A. Tu- Chan, K. Ganguly, and E. F. Chang, “Generalizable spelling using a speech neuroprosthesis in an individual with severe limb and vocal paralysis,” Nature Communications, vol. 13, p. 6510, 2022. [24] M. Wairagkar, N. S. Card, T. Singer-Clark, X. Hou, C. Iacobacci, L. R. Hochberg, D. M. Brandman, and S. D. Stavisky, “An instan- taneous voice-synthesis neuroprosthesis,” Nature, vol. 644, p. 145–152, 2025. [25] A. Baevski, Y. Zhou, A. Mohamed, and M. Auli, “wav2vec 2.0: A framework for self-supervised learning of speech representa- tions,” in Advances in Neural Information Processing Systems, vol. 33, 2020, p. 12 449–12 460. [26] A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, and I. Sutskever, “Robust speech recognition via large-scale weak su- pervision,” Proc. International Conference on Machine Learning (ICML), p. 28 492–28 518, 2023. [27] F. R. Willett, J. Li, T. Le, C. Fan, M. Chen, and E. Shlizerman, “Brain-to-text benchmark ’24: Lessons learned,” arXiv preprint arXiv:2412.17227, 2024. [28] J. Li, T. Le, C. Fan, M. Chen, and E. Shlizerman, “Brain-to-text decoding with context-aware neural representations and large lan- guage models,” arXiv preprint arXiv:2411.10657, 2024. [29] A. Gu and T. Dao, “Mamba: Linear-time sequence modeling with selective state spaces,” arXiv preprint arXiv:2312.00752, 2023. [30] K. Miyazaki, Y. Masuyama, and M. Murata, “Exploring the ca- pability of Mamba in speech applications,” in Proc. Interspeech, 2024, p. 176–180. [31] X. Jiang, Y. A. Li, A. N. Florea, C. Han, and N. Mesgarani, “Speech Slytherin: Examining the performance and efficiency of Mamba for speech separation, recognition, and synthesis,” arXiv preprint arXiv:2407.09732, 2024. [32] H. Hou et al., “ConMamba: A convolution-augmented Mamba encoder model for efficient end-to-end ASR systems,” in Proc. IEEE International Conference on Signal, Information and Data Processing (ICSIDP), 2024. [33] R. Zevallos, M. Cortada-Garcia, S. Solito, C. Mena, A. Peiro- Lilja, and J. Hernando, “Assessing the performance and efficiency of Mamba ASR in low-resource scenarios,” in Proc. Interspeech, 2025, p. 5198–5202. [34] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Advances in Neural Information Processing Systems, vol. 30, 2017. [35] W. Chan, N. Jaitly, Q. Le, and O. Vinyals, “Listen, attend and spell: A neural network for large vocabulary conversational speech recognition,” in Proc. IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2016, p. 4960–4964. [36] A. Gulati, J. Qin, C.-C. Chiu, N. Parmar, Y. Zhang, J. Yu, W. Han, S. Wang, Z. Zhang, Y. Wu, and R. Pang, “Conformer: Convolution-augmented transformer for speech recognition,” in Proc. Interspeech, 2020, p. 5036–5040. [37] M. Burchi and V. Vielzeuf, “Efficient Conformer: Progressive downsampling and grouped attention for automatic speech recog- nition,” in Proc. IEEE Automatic Speech Recognition and Under- standing Workshop (ASRU), 2021, p. 8–15. [38] Y. Peng, S. Dalmia, I. Lane, and S. Watanabe, “Branchformer: Parallel MLP-attention architectures to capture local and global context for speech recognition and understanding,” in Proc. In- ternational Conference on Machine Learning (ICML), 2022, p. 17 627–17 643. [39] O. M. Khanday, J. A. Gonzalez-Lopez, M. Ouellet, A. Galdon, and G. Olivares Granados, “End-to-end intracortical speech de- coding from neural activity,” arXiv preprint arXiv:2605.24313, 2026. [40] J. D. Simeral, S.-P. Kim, M. J. Black, J. P. Donoghue, and L. R. Hochberg, “Neural control of cursor trajectory and click by a human with tetraplegia 1000 days after implant of an intracorti- cal microelectrode array,” Journal of Neural Engineering, vol. 8, no. 2, p. 025027, 2011. [41] J. A. Gallego, M. G. Perich, L. E. Miller, and S. A. Solla, “Neural manifolds for the control of movement,” Neuron, vol. 94, no. 5, p. 978–984, 2017. [42] B. Jarosiewicz, A. A. Sarma, D. Bacher, N. Y. Masse, J. D. Simeral, B. Sorice, E. M. Oakley, C. Blabe, C. Pandarinath, V. Gilja, S. S. Cash, E. N. Eskandar, G. Friehs, J. M. Hender- son, K. V. Shenoy, J. P. Donoghue, and L. R. Hochberg, “Virtual typing by people with tetraplegia using a self-calibrating intracor- tical brain-computer interface,” Science Translational Medicine, vol. 7, no. 313, p. 313ra179, 2015. [43] A. D. Degenhart, W. E. Bishop, E. R. Oby, E. C. Tyler-Kabara, S. M. Chase, A. P. Batista, and B. M. Yu, “Stabilization of a brain- computer interface via the alignment of low-dimensional spaces of neural activity,” Nature Biomedical Engineering, vol. 4, no. 7, p. 672–685, 2020. [44] D. S. Park, W. Chan, Y. Zhang, C.-C. Chiu, B. Zoph, E. D. Cubuk, and Q. V. Le, “SpecAugment: A simple data augmen- tation method for automatic speech recognition,” in Proc. Inter- speech, 2019, p. 2613–2617. [45] K. Cho, B. van Merri ̈ enboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase rep- resentations using RNN encoder–decoder for statistical machine translation,” in Proc. Conference on Empirical Methods in Natu- ral Language Processing (EMNLP), 2014, p. 1724–1734. [46] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An image is worth 16x16 words: Transformers for image recognition at scale,” in Proc. International Conference on Learning Representations (ICLR), 2021. [47] A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language models are unsupervised multitask learn- ers,” OpenAI, Tech. Rep., 2019. [48] K. Heafield, “KenLM: Faster and smaller language model queries,” in Proc. Sixth Workshop on Statistical Machine Trans- lation (WMT), 2011, p. 187–197. [49] F. Wilcoxon, “Individual comparisons by ranking methods,” Bio- metrics Bulletin, vol. 1, no. 6, p. 80–83, 1945. [50] G. Santaf ́ e, I. Inza, and J. A. Lozano, “Dealing with the evaluation of supervised classification algorithms,” Artificial Intelligence Re- view, vol. 44, no. 4, p. 467–508, 2015.