Paper deep dive
MEBM-Phoneme: Multi-scale Enhanced BrainMagic for End-to-End MEG Phoneme Classification
Liang Jinghua, Zhang Zifeng, Li Songyi, Zheng Linze
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 7/20/2026, 7:07:49 AM
Summary
The paper introduces MEBM-Phoneme, a multi-scale enhanced neural decoder for end-to-end phoneme classification from non-invasive magnetoencephalography (MEG) signals. Built upon the BrainMagic backbone, the model integrates a short-term multi-scale convolutional module and a convolutional attention layer to capture fine-grained temporal dependencies. It employs a stacking-based local validation strategy and weighted cross-entropy loss to address class imbalance and distributional shifts, achieving competitive performance in the NeurIPS 2025 LibriBrain Competition.
Entities (8)
Relation Signals (7)
MEBM-Phoneme → processes → MEG
confidence 99% · We propose MEBM-Phoneme, a multi-scale enhanced neural decoder for phoneme classification from non-invasive magnetoencephalography (MEG) signals.
MEBM-Phoneme → builtupon → BrainMagic
confidence 95% · Built upon the BrainMagic backbone, MEBM-Phoneme integrates a short-term multi-scale convolutional module...
MEBM-Phoneme → usedin → LibriBrain Competition 2025
confidence 92% · Comprehensive evaluations on LibriBrain Competition 2025 Track2 demonstrate robust generalization...
LibriBrain Competition 2025 → partof → NeurIPS 2025
confidence 90% · Our approach is developed for Track 2 of the NeurIPS 2025 LibriBrain Competition
MEBM-Phoneme → usesdataset → Sherlock1
confidence 88% · The offline validation set was constructed using the official validation and test sessions (Sherlock1, sessions 11–12)
MEBM-Phoneme → implementedin → PyTorch
confidence 85% · The proposed MEBM-Phoneme model was implemented in PyTorch
MEBM-Phoneme → trainedonhardware → NVIDIA A800 GPU
confidence 85% · trained on a single NVIDIA A800 GPU (80 GB)
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We propose MEBM-Phoneme, a multi-scale enhanced neural decoder for phoneme classification from non-invasive magnetoencephalography (MEG) signals. Built upon the BrainMagic backbone, MEBM-Phoneme integrates a short-term multi-scale convolutional module to augment the native mid-term encoder, with fused representations via depthwise separable convolution for efficient cross-scale integration. A convolutional attention layer dynamically weights temporal dependencies to refine feature aggregation. To address class imbalance and session-specific distributional shifts, we introduce a stacking-based local validation set alongside weighted cross-entropy loss and random temporal augmentation. Comprehensive evaluations on LibriBrain Competition 2025 Track2 demonstrate robust generalization, achieving competitive phoneme decoding accuracy on the validation and official test leaderboard. These results underscore the value of hierarchical temporal modeling and training stabilization for advancing MEG-based speech perception analysis.
Tags
Links
- Source: https://arxiv.org/abs/2603.02254v1
- Canonical: https://arxiv.org/abs/2603.02254v1
Trouble viewing inline? Open PDF directly →
Full Text
14,829 characters extracted from source content.
Expand or collapse full text
MEBM-Phoneme: Multi-scale Enhanced BrainMagic for End-to-End MEG Phoneme Classification Liang Jinghua1 &Zhang Zifeng1,2 &Li Songyi1 &Zheng Linze1 1Speech and Hearing Research Center, School of Intelligence Science and Technology 2Center for BioMed-X Research, Academy for Advanced Interdisciplinary Studies Peking University, Beijing, China 100871 zifengzhang25@stu.pku.edu.cn Abstract We propose MEBM-Phoneme, a multi-scale enhanced neural decoder for phoneme classification from non-invasive magnetoencephalography (MEG) signals. Built upon the BrainMagic backbone, MEBM-Phoneme integrates a short-term multi-scale convolutional module to augment the native mid-term encoder, with fused representations via depthwise separable convolution for efficient cross-scale integration. A convolutional attention layer dynamically weights temporal dependencies to refine feature aggregation. To address class imbalance and session-specific distributional shifts, we introduce a stacking-based local validation set alongside weighted cross-entropy loss and random temporal augmentation. Comprehensive evaluations on LibriBrain Competition 2025 Track 2 demonstrate robust generalization, achieving competitive phoneme decoding accuracy on the validation and official test leaderboard. These results underscore the value of hierarchical temporal modeling and training stabilization for advancing MEG-based speech perception analysis. 1 Introduction Phoneme decoding from brain signals has long been a central goal of neural speech decoding research. Recent advances in invasive neuroprosthetic technologies achieve remarkable accuracy by directly mapping neural activity to phoneme categories and then decoding text through language models metzger2023 ; willett2023 ; card2024 . However, replicating such performance using non-invasive neuroimaging techniques, such as MEG, remains highly challenging due to the lower signal-to-noise ratio. To address this problem, we propose MEBM-Phoneme, an enhanced end-to-end framework for MEG-based phoneme classification. Our method is developed for Track 2 of the NeurIPS 2025 LibriBrain Competition landau2025 ; ozdogan2025 , which focuses on decoding phonemic representations from non-invasive MEG recordings. Our approach centers on three key contributions: 1. Model Architecture: We augment the BrainMagic defossez2023 architecture with a short-term multi-scale convolutional module, capturing fine-grained temporal dependencies. The resulting features are fused with mid-term representations through a depthwise separable convolution, followed by a convolutional attention layer that aggregates temporal information. 2. Validation Strategy: To address severe class imbalance and better approximate the holdout distribution, we construct a session-aware local validation set using a stacking-based sampling method, ensuring statistical alignment with the competition’s evaluation protocol. 3. Training Protocol: To enhance robustness and address class imbalance, we adopt a stochastic sample construction strategy that randomly selects a phoneme class per iteration and dynamically averages a variable number of instances. Together with random temporal offsets and an adaptive weighted cross-entropy loss, this approach promotes balanced learning and stable convergence. 2 Methods Our approach for MEG-based phoneme classification is designed to enhance temporal feature modeling, handle severe class imbalance, and improve training robustness. Below, we present the detailed description of our method. 2.1 Model Architecture Figure 1: Overall architecture of the proposed MEBM-Phoneme model. (a) The complete processing pipeline. (b) The spatial attention module enhances sensor-level representations by learning spatial relevance weights across MEG channels. (c) The BM encoder extracts mid-term contextual features from spatially weighted signals. (d) The short-term multi-scale convolutional module captures fine-grained temporal dependencies using multiple receptive fields. (e) The depthwise separable convolutional layer further refines temporal representations with lightweight channel-wise and pointwise filtering. As illustrated in Figure 1, our proposed MEBM-Phoneme model builds upon the original BrainMagic architecture by introducing a dedicated short-term feature extraction pathway and an enhanced fusion mechanism. Given the MEG input ∈ℝCin×TX ^C_in× T, where CinC_in denotes the number of MEG sensor channels and T represents the total number of temporal samples, the model first applies a spatial attention module that dynamically re-weights sensor-wise activations, producing a spatially enhanced representation s∈ℝD×TH_s ^D× T, with D being the dimensionality of the projected feature space. This representation is then processed in parallel by two temporal streams: 12 multi-scale convolutional blocks comprising a stack of dilated convolutional blocks designed to capture local temporal dependencies across multiple receptive fields, and 3 BrainMagic (BM) encoders responsible for extracting mid-term contextual features. The outputs from both branches are concatenated along the channel dimension and passed through a depthwise separable convolution for efficient fusion. This operation not only reduces computational overhead but also enforces feature disentanglement across temporal scales. Subsequently, a convolutional attention layer aggregates temporal information through a channel-compressive operation. Specifically, a 11D convolution reduces the feature dimension to 11, yielding an attention map ∈ℝ1×TA ^1× T. A softmax normalization is then applied along the temporal axis to obtain attention weights tW_t, which are multiplied back with the fused representation to reweight each time step by its learned importance: att=t⊙fused.H_att=W_t _fused. Finally, a sum pooling operation collapses the temporal dimension, and a linear layer with softmax activation produces the phoneme-level class probabilities. 2.2 Validation and Training Sampling Strategy To ensure robust and distributionally aligned evaluation, we design a unified data construction rule for both validation and training samples, differing only in the degree of stochasticity. For each phoneme class, we estimate the average number of samples per session n, and determine the number of averaged single samples n′n according to: nval′=100,n>100,n,50≤n<100,1.5n,n<50,n _val= cases100,&n>100,\\ n,&50≤ n<100,\\ 1.5n,&n<50, cases (1) ntrain′=100,n>100,rand[n−5,min(n+5,100)],50≤n<100,2n,n<50.n _train= cases100,&n>100,\\ rand[n-5, (n+5,100)],&50≤ n<100,\\ 2n,&n<50. cases (2) Here, nval′n _val defines a deterministic sampling rule for validation, while ntrain′n _train introduces controlled randomness during training to enhance generalization and reduce overfitting. At each training iteration, a single phoneme class is randomly selected, and its samples are averaged following Eq. 2. To further improve temporal robustness, we apply a random temporal jittering scheme: the starting point of each segment is uniformly sampled from the interval [onset−3,onset+3][onset-3,onset+3], and a fixed 0.50.5 s window is subsequently extracted. This perturbation increases invariance to onset timing variability inherent in MEG signals. Finally, training is guided by an adaptive weighted cross-entropy loss, designed not only for class balancing but also to reduce confusion among acoustically or articulatorily similar phonemes. 3 Experiments 3.1 Experimental Setup The offline validation set was constructed using the official validation and test sessions (Sherlock1, sessions 11–12) to approximate the holdout distribution defined by the LibriBrain challenge. For reproducibility, we fixed random seeds and performed eight independent sampling iterations per phoneme class, discarding classes with insufficient samples to meet the required n′n values from Eq. 1. This procedure ensured that the resulting validation data statistically aligned with the holdout distribution while mitigating class imbalance and session-specific bias. Before training, the continuous MEG signals of each session were normalized along the temporal dimension independently. After sample extraction and averaging, the resulting averaged samples were normalized again along the temporal axis. The proposed MEBM-Phoneme model was implemented in PyTorch and trained on a single NVIDIA A800 GPU (80 GB) for approximately three hours. The network contained 4.7 M trainable parameters. Each input MEG sequence consisted of Cin=306C_in=306 channels and T=125T=125 time points, producing Cout=39C_out=39 phoneme probabilities. The intermediate feature dimension was set to D=128D=128 with a dropout rate of 0.02. Training was conducted for 80 epochs using the AdamW optimizer with a learning rate of 1×10−31× 10^-3, batch size of 256, and 40,000 samples per epoch. All convolutional layers adopted padding=’same’ to preserve temporal resolution. Model selection and hyperparameter tuning were performed using the offline validation set constructed from the Sherlock1 Session 11–12 data. 3.2 Results and Ablation We report the performance of the proposed MEBM-Phoneme model and its ablated variants on the offline validation set. All results are averaged over six random seeds 0,1,2,3,4,5\0,1,2,3,4,5\ for reproducibility. Evaluation metrics include F1macro(%), Top-3 Accmacro (%), and Top-5 Accmacro (%). Table 1 summarizes the performance of our proposed MEBM-Phoneme model and its ablated variants on the validation set. The full model achieves an average F1macro of 60.95%, Top-3 Accmacro of 89.54%, and Top-5 Accmacro of 95.08% across six random seeds. Removing any individual component leads to a consistent degradation in performance, demonstrating the effectiveness of each module in the proposed architecture. Moreover, all model variants maintain relatively high Top-3 and Top-5 Accmacro scores, indicating that even when the top prediction is incorrect, the correct phoneme often lies among the top few candidates. This suggests that the model already possesses a strong discriminative capacity for phoneme categorization, and could further benefit from integration with a language model to leverage contextual linguistic information. On the online test set, our model achieved a peak decoding accuracy of up to 72% on the first half, but exhibited degraded performance on the second half. We conjecture that this discrepancy may be partly attributed to our submission strategy. Nevertheless, results on the local evaluation set indicate that our approach remains robust and demonstrates strong generalization capability. Table 1: Results and ablation analysis on the local validation set under six random seeds (0–5). Metrics include F1macro, Top-3 Accmacro, and Top-5 Accmacro (mean ± std). Model Variant F1macro (%) Top-3 Accmacro (%) Top-5 Accmacro (%) Full Model 60.95±0.90 89.54±0.48 95.08±0.61 w/o Weighted Loss 59.97±0.90 88.87±1.14 94.75±0.63 w/o Multi-scale Conv 59.75±0.68 88.98±1.12 94.67±1.03 w/o BM Encoder 54.43±2.07 84.96±1.69 92.19±1.28 w/o Conv. Attention 59.60±0.82 88.47±1.46 94.17±1.13 4 Conclusion This work presents MEBM-Phoneme, our enhanced framework for MEG-based phoneme classification in the NeurIPS 2025 LibriBrain Competition. By augmenting the BrainMagic architecture with a short-term multi-scale convolutional module and an attention-based temporal aggregation mechanism, the model effectively captures both fine-grained and contextual temporal dependencies from non-invasive MEG signals. Additionally, our session-aware validation strategy and stochastic training protocol improve robustness against class imbalance and distributional variation. Experimental results under multiple random seeds demonstrate that each component of MEBM-Phoneme contributes to stable performance improvements, achieving competitive results on the official evaluation set. It is important to note, however, that our study relies on averaged MEG signals to boost the signal-to-noise ratio. A significant remaining challenge—and the focus of our future work—is to perform accurate phoneme classification on single-trial, continuous MEG data, which is essential for developing practical, real-time neural speech decoding systems. References (1) Metzger, S.L., Littlejohn, K.T., Silva, A.B., Moses, D.A., Seaton, M.P., Wang, R., Dougherty, M.E., Liu, J.R., Wu, P., Berger, M.A., Zhuravleva, I., Tu-Chan, A., Ganguly, K., Anumanchipalli, G.K. & Chang, E.F. (2023) A high-performance neuroprosthesis for speech decoding and avatar control. Nature, 620(7976), 1037–1046. (2) Willett, F.R., Kunz, E.M., Fan, C., Avansino, D.T., Wilson, G.H., Choi, E.Y., Kamdar, F., Glasser, M.F., Hochberg, L.R., Druckmann, S., Shenoy, K.V. & Henderson, J.M. (2023) A high-performance speech neuroprosthesis. Nature, 620(7976), 1031–1036. (3) Card, N.S., Wairagkar, M., Iacobacci, C., Hou, X., Singer-Clark, T., Willett, F.R., Kunz, E.M., Fan, C., Vahdati Nia, M., Deo, D.R., Srinivasan, A., Choi, E.Y., Glasser, M.F., Hochberg, L.R., Henderson, J.M., Shahlaie, K., Stavisky, S.D. & Brandman, D.M. (2024) An accurate and rapidly calibrating speech neuroprosthesis. New England Journal of Medicine, 391(7):609–618. (4) Landau, G., Özdogan, M., Elvers, G., Mantegna, F., Somaiya, P., Jayalath, D., Kurth, L., Kwon, T., Shillingford, B., Farquhar, G., Jiang, M., Jerbi, K., Abdelhedi, H., Ramos, Y.M., Gulcehre, C., Woolrich, M., Voets, N. & Jones, O.P. (2025) The 2025 PNPL Competition: Speech Detection and Phoneme Classification in the LibriBrain Dataset. arXiv, 2506.10165. doi:10.48550/arXiv.2506.10165. (5) Özdogan, M., Landau, G., Elvers, G., Jayalath, D., Somaiya, P., Mantegna, F., Woolrich, M. & Jones, O.P. (2025) LibriBrain: Over 50 hours of within-subject MEG to improve speech decoding methods at scale. arXiv, 2506.02098. doi:10.48550/arXiv.2506.02098. (6) Défossez, A., Caucheteux, C., Rapin, J., Kabeli, O. & King, J.-R. (2023) Decoding speech perception from non-invasive brain recordings. Nature Machine Intelligence, 5(10):1097–1107. Appendix A. Adaptive Loss Weights for Phoneme Classes Table 2 lists the adaptive loss weights used for each phoneme class. The weights were empirically tuned to balance class frequency and confusion. The remaining phoneme weights were all set to 1.0. Table 2: Adaptive loss weights for each phoneme class. Phoneme /ey/ /ay/ /uh/ /uw/ /s/ /sh/ /m/ /ae/ /jh/ /ah/ Weight 0.05 3.00 10.00 3.00 0.80 3.00 3.00 3.00 1.50 2.00