Paper deep dive
Copilot-Assisted Second-Thought Framework for Brain-to-Robot Hand Motion Decoding
Yizhe Li, Shixiao Wang, Jian K. Liu
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 3/31/2026, 2:06:04 AM
Summary
This paper introduces a CNN-attention hybrid model for decoding hand kinematics (3D trajectories) from EEG and EEG-EMG multimodal signals during grasp-and-lift tasks. To improve trajectory fidelity, the authors propose a 'copilot' framework that uses a motion-state-aware critic and a finite-state machine to filter low-confidence decoded points, demonstrating improved performance in MuJoCo robotic arm simulations.
Entities (5)
Relation Signals (3)
CNN-attention hybrid model â decodes â hand kinematics
confidence 95% ¡ we propose a CNNâattention hybrid model for decoding hand kinematics from EEG
CNN-attention hybrid model â testedon â WAY-EEG-GAL
confidence 95% ¡ The WAY-EEG-GAL dataset contains scalp EEG recordings... The model architecture is illustrated in Fig. 1.
Copilot framework â refines â decoded trajectories
confidence 92% ¡ To enhance trajectory fidelity, we introduce a copilot framework that filters low-confidence decoded points
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Motor kinematics prediction (MKP) from electroencephalography (EEG) is an important research area for developing movement-related brain-computer interfaces (BCIs). While traditional methods often rely on convolutional neural networks (CNNs) or recurrent neural networks (RNNs), Transformer-based models have shown strong ability in modeling long sequential EEG data. In this study, we propose a CNN-attention hybrid model for decoding hand kinematics from EEG during grasp-and-lift tasks, achieving strong performance in within-subject experiments. We further extend this approach to EEG-EMG multimodal decoding, which yields substantially improved results. Within-subject tests achieve PCC values of 0.9854, 0.9946, and 0.9065 for the X, Y, and Z axes, respectively, computed on the midpoint trajectory between the thumb and index finger, while cross-subject tests result in 0.9643, 0.9795, and 0.5852. The decoded trajectories from both modalities are then used to control a Franka Panda robotic arm in a MuJoCo simulation. To enhance trajectory fidelity, we introduce a copilot framework that filters low-confidence decoded points using a motion-state-aware critic within a finite-state machine. This post-processing step improves the overall within-subject PCC of EEG-only decoding to 0.93 while excluding fewer than 20% of the data points.
Tags
Links
- Source: https://arxiv.org/abs/2603.27492v1
- Canonical: https://arxiv.org/abs/2603.27492v1
Trouble viewing inline? Open PDF directly â
Full Text
30,476 characters extracted from source content.
Expand or collapse full text
Copilot-Assisted Second-Thought Framework for Brain-to-Robot Hand Motion Decoding Yizhe Li â University of Birmingham yxl1897@alumni.bham.ac.uk Shixiao Wang â University of Birmingham sxw1238@student.bham.ac.uk Jian K. Liu University of Birmingham j.liu.22@bham.ac.uk AbstractâMotor kinematics prediction (MKP) from electroen- cephalography (EEG) is a key research area for developing movement-related brainâcomputer interfaces (BCIs). While tra- ditional methods often rely on CNNs or RNNs, Transformers have demonstrated superior capabilities for modeling long sequential EEG data. In this study, we propose a CNNâattention hybrid model for decoding hand kinematics from EEG during grasp- and-lift tasks, achieving superb performance in within-subject experiments. We further extend this approach to EEGâEMG multimodal decoding, yielding significantly improved perfor- mance. Within-subject tests achieve PCC values of 0.9854, 0.9946, and 0.9065 (X, Y, Z), respectively, computed on the midpoint tra- jectory between the thumb and index finger, while cross-subject tests result in 0.9643, 0.9795, and 0.5852. The decoded trajectories from both modalities were used to control a Franka Panda robotic arm in a MuJoCo simulation. To enhance trajectory fidelity, we introduce a copilot framework that filters low-confidence decoded points using a motion-state-aware critic within a finite- state machine. This post-processing step improves the overall within-subject PCC of EEG-only decoding to 0.93 while excluding fewer than 20% of the data points. The code is available at https://github.com/ssshiFang/EEGDecoding Robotarm. Index TermsâBCI, EEG, EMG, copilot, hybrid modality, motor kinematics prediction, CNN, transformer, robot arm I. INTRODUCTION Electroencephalography (EEG) is a fundamental tool in neuroscience, offering the potential for controlling external devices directly through brain activity, bypassing the neu- romuscular pathway [1]. This concept is realized through brainâcomputer interfaces (BCIs). The central nervous system (CNS) typically generates movement in response to external stimuli, acting on effectors such as muscles or glands [2], [3]. BCIs introduce an alternative pathway for this interaction by leveraging digital technologies. They can utilize signals from task-related brain regions to enhance, monitor, or even partially replace traditional CNSâeffector communication [4]. While BCIs can theoretically use any brain-derived signal as input, practical limitations exist, such as random mental activity or external noise. Designing an effective BCI paradigm is therefore critical. By using various types of stimuli and tasks, a mapping can be established between specific inputs and corresponding brain activity patterns [5]. Current BCI systems are primarily built on paradigms like motor imagery, event-related potentials, and steady-state visually evoked po- â These authors contributed equally to this work tentials, all of which heavily rely on non-invasive EEG for data acquisition [6]. Decoding algorithms are another crucial component, trans- lating measured brain signals into executable control com- mands. However, this step is particularly challenging due to the inherent properties of EEG. Scalp EEG represents a summa- tion of currents from multiple local field potentials, reflecting broad brain activity [7], [8]. Furthermore, interference between current sources complicates spatial localization, and signal transmission through various tissue layers results in a very low signal-to-noise ratio. The advent of deep learning models, such as Shallow ConvNet and EEGNet, has demonstrated superiority over traditional handcrafted feature methods [9]â [11]. Transformers have further shown powerful capabilities in handling long sequences, overcoming limitations of CNNs and RNNs [12]â[14]. Consequently, hybrid models combining these architectures are increasingly applied to classification tasks [15], [16]. Despite these advances, most kinematic trajec- tory prediction tasks using movement-related EEG, beyond the core paradigm classifications, still rely on filter-based, linear, or CNN/RNN approaches [17] and continue to struggle with decoding accuracy [18]. This work proposes a CNNâattention hybrid model for EEG-based hand kinematics reconstruction. We further inves- tigate EEGâEMG multimodal decoding, reconstruct motion trajectories in a MuJoCo robotic arm simulation, and design a copilot algorithm to refine decoding points through secondary evaluation. I. RELATED WORK EEG-based kinematics decoding aims to reconstruct con- tinuous limb movement trajectories (e.g., position, velocity) from non-invasive EEG signals. Prior studies have targeted various body parts, including ankle plantar flexion [19], hand kinematics [20], finger movements [21], and even saccadic eye movements [22]. These tasks are closely linked to the temporalâspectral dynamics of EEG and have been explored in both motor imagery and motor execution paradigms [23], [24]. Among these, hand-movement decoding is the most extensively studied. Early work often assumed EEG lacked sufficient informa- tion for complex hand movement representation, a notion later challenged by [25]. [26] further elucidated the role of contralateral motor cortical regions in upper-limb movement arXiv:2603.27492v1 [cs.RO] 29 Mar 2026 control. These studies laid the foundation for hand-kinematics decoding and advanced understanding beyond simple single- region activation models. Deep learning models have demonstrated superior per- formance in capturing nonlinear relationships between EEG features and motor parameters, especially when combined with timeâfrequency or spatialâspectral preprocessing. A CNNâLSTM model decoded hand kinematics from cortical sources with a mean correlation value (CV) of 0.62Âą0.08 [20]. For other body parts, biceps curl trajectory estimation reached a PCC of 0.7 [27], and a multi-directional CNN-BiLSTM net- work for 3D arm tasks achieved a grand-averaged correlation coefficient (C) of 0.47 [28]. A comparison in [29] showed deep learning models (rEEGNet, rDeepConvNet, rShallow- ConvNet) significantly outperformed mLR, with mean PCCs reaching 0.78â0.79 for x/y axes and 0.62â0.64 for the z-axis. Despite progress, emerging technologies offer new so- lutions. Hybrid CNNâtransformer architectures have shown strong capabilities in EEG decoding tasks, such as EEGformer [30], EASM [31], EEG-TCNTransformer [32], and EEG- ConvTransformerNetwork [33]. Given that most kinematics decoding models remain CNN- or RNN-based, the success of the AI copilot framework in [34] motivated our exploration of hybrid modeling and trajectory optimization for non-invasive EEG-based hand decoding. The contributions of this study are as follows: (1) We propose a CNNâattention hybrid model to predict 3D coor- dinates of the index finger and thumb from EEG. The model achieves an overall PCC of 0.8728 in within-subject exper- iments (0.8376, 0.9229, 0.7208 for X, Y, Z axes), reaching up to 0.8916 with a larger input window. (2) We investigate EEGâEMG multimodal fusion for kinematics decoding. The best within-subject PCC is 0.9707 (0.9854, 0.9946, 0.9065 for X, Y, Z), and the best cross-subject total PCC reaches 0.9063 (0.9643, 0.9795, 0.5852 for X, Y, Z). (3) We reconstruct hand movement trajectories in a MuJoCo robotic arm simulator and design a copilot framework based on a state machine and knowledge graph to filter unreliable decoding points. I. MATERIALS AND METHODS A. Dataset Description [35] The WAY-EEG-GAL (Wearable interfaces for hand function recovery EEG grasp and lift) dataset contains scalp EEG recordings from a grasp-and-lift task designed to decode sensory, intentional, and motor-related signals. Twelve partic- ipants each performed 328 trials, totaling 3,936 trials. Each trial required participants to: (1) reach for a small object upon cue, (2) grasp it with the thumb and index finger, (3) lift and hold it briefly, (4) place it back, and (5) return the hand to the start position. The dataset includes 32-channel EEG, electromyography (EMG) from five muscles, and contact force/torque and 3D position data for the hand (index finger, thumb, wrist) and the object. Fig. 1: Structure of the decoding model. B. Data Preprocessing EEG preprocessing was performed using MNE-Python. Raw data were band-pass filtered between 0.1 and 40 Hz with an infinite impulse response (IIR) filter to remove low- frequency drifts and high-frequency noise. Common average referencing (CAR) was then applied across all electrodes to reduce spatially distributed artifacts and enhance localized neural activity. EMG preprocessing involved two steps. First, a 4th-order Butterworth band-pass filter (20â450 Hz), implemented in second-order sections (SOS) for stability, removed high- frequency noise and baseline drift. Second, EMG data were downsampled from 4000 Hz to 500 Hz to match the EEG sampling rate. Finally, kinematic data were scaled via minâmax normal- ization, mapping the processed data to the range [0, 1] as K[t] = k[t]âk min k max âk min . C. Algorithm Structure 1) Model Overview: The model architecture is illustrated in Fig. 1. It consists of five main components: a multi-convolution block, a Squeeze-and-Excitation (SE) block, an embedding block, a self-attention block, and a fully connected block. The input is a preprocessed EEG segment in the form of a 2D matrix (channels Ă time points). 2) Multi-Convolution Block: The multi-convolution block (Fig. 2) is partially inspired by EEGNet. First, a 1D convolu- tion with a large temporal kernel captures long-range depen- dencies [36]. Subsequent layers employ multiple smaller ker- nels of varying sizes to extract multi-scale temporal features. The 1D convolution is implemented as y[t] = P Kâ1 k=0 w[k]¡ x[t+k], where y[t] is the output for the t-th element. A pooling layer then simplifies features and reduces computational cost for the subsequent attention mechanism. The average pooling operation is defined as y i = 1 k P kâ1 j=0 x i¡s+j . 3) Squeeze-and-Excitation (SE) Block: The SE block high- lights informative input channels. Following [37], for an input Fig. 2: Structure of the multi-convolution block (top). Structure of the self-attention block (bottom). feature map XâR BĂCĂFĂT , the block performs: Squeeze: z = 1 F ¡ T F X i=1 T X j=1 X(i,j),(1) Excitation: s = Ď (W 2 ¡ δ (W 1 ¡ z)),(2) Scale: Ě X = sâ X,(3) where Ď and δ are activation functions. 4) Embedding and Self-Attention Blocks: The embedding block (Fig. 2) abstracts the channelâfeature dimension (C Ă F) into tokens. The attention block then establishes long-range dependencies between these tokens. The attention mechanism is defined as [38]: Attention(Q,K,V ) = softmax QK ⤠â d k V.(4) A single attention block is used, but a multi-head strategy captures diverse channel relationships: MultiHead(Q,K,V ) = Concat(head 1 ,..., head h )W O , (5) head i = Attention(QW Q i ,KW K i ,V W V i ). (6) 5) Fully Connected Layer: This module processes the attention output ZâR TĂD . Global average pooling is applied over the T dimension, followed by layer normalization and two fully connected layers to produce the final output. D. Modality Fusion As shown in Fig. 1, EMG signals undergo the same multi- convolution and SE block processing as EEG. Fusion of EEG and EMG signals occurs within the embedding module. Fig. 3: Structure of the copilot and decoding point filtering workflow. The grasp-and-lift movement is segmented into five states (SEARCHING, LIFTING, HOLDING, PUTTING, RETURNING), each with different confidence thresholds. UNRELY is a virtual state representing points misclassified as belonging to the current state, which are assigned a higher decoding threshold for filtering. E. Copilot Framework As depicted in Fig. 3, the copilot module integrates three components: (1) a decoding model that predicts spatial coor- dinates, (2) a label prediction model (identical structure) that classifies motion states, and (3) a critic model that estimates confidence scores for each decoded point. A knowledge graph interacts with external sensor information to drive state tran- sitions within a finite state machine. IV. EXPERIMENT DETAILS A. Experiment Design EEG and EMG signals from each grasp-and-lift trial served as decoding inputs. Considering MRCP latency and the trans- formerâs sequence modeling capability, various window sizes (50â1000 samples) and delays (100â700 ms) were evaluated. Model input was a 3D tensor (batchĂ channelĂ time points), with outputs being minâmax normalized 3D coordinates for the index finger and thumb (6 values total). The dataset was split following [29]: For each participant, we randomly selected 30 trials as the validation set and another 30 non- overlapping trials as the test set, the rest for training, using MSELoss. For evaluating window sizes and delays, data from partic- ipant 4 was used. All delays and slicing were applied post- preprocessing to avoid artifact spread from filtering bound- aries; segments with mismatched EEG and kinematics data after delay were discarded. For within-subject model comparisons, the proposed model was tested against EEGNet, DeepConvNet, TCN, and a transformer baseline (EEG-only). The modified rEEGNet and rDeepConvNet from [29] were used. The multimodal version of our model used combined EEGâEMG input. Tests used a 250-sample window and 200 ms delay. Within-subject tests involved participants 3, 4, 5, 7, 9; cross-subject tests decoded Fig. 4: (Left) Within-subject decoding performance for differ- ent models (EEGNet, DeepConvNet, Transformer, EEG-TCN, EEG-only, and EMGâEEG fusion). (Right) Within-subject and cross-subject decoding performance for EEG-only and EMGâEEG fusion models. participant 1âs data using models trained on these five. Abla- tion studies validated model components. Decoded trajectories were simulated on a 7-DOF Franka Panda robotic arm in MuJoCo. For each trial, the midpoint of decoded index finger/thumb trajectories was mapped to the arm workspace, converted to joint angles via inverse kinematics, and interpolated for continuous motion. B. Copilot Implementation EEG preprocessing for the copilot was identical to previous experiments. The dataset was split differently: We reused the same trial-level train/validation/test split as previously for training the critic model, motion stage classifier, and state transition model. C. Performance Metrics 1) Pearson Correlation Coefficient (PCC): PCC mea- sures linear correlation between decoded and actual 3D coordinate trajectories (thumb and index finger). The for- mula for two sequences C x and C y is: PCC xy = P n i=1 (C x i â Ě C x )(C y i â Ě C y ) â P n i=1 (C x i â Ě C x ) 2 â P n i=1 (C y i â Ě C y ) 2 . The overall PCC is com- puted by concatenating all x, y, z outputs into a single vector. 2) Root Mean Square Error (RMSE): RMSE quanti- fies the deviation between predicted and actual trajectories: RMSE xy = q 1 N P N i=1 (C x i â C y i ) 2 . V. RESULTS AND DISCUSSION A. Model Comparison 1) Within-Subject Decoding: As shown in Fig. 4 (Left), except for the EEGâEMG multimodal approach, all EEG- only models showed notable inter-subject variability. The multimodal model remained stable across participants 3, 4, 5, 7, and 9, with PCC values of 0.94â0.97 and RMSE of 0.09â0.15. Participant 7 achieved the highest accuracy (PCC = 0.9707, RMSE = 0.0989), while participant 5 had the lowest (PCC = 0.9369, RMSE = 0.1447). In contrast, EEG-only models performed worse. Our proposed CNNâattention model Fig. 5: (Left) Model performance with varying input window sizes (50â1000 samples) for participant 4 with a 200 ms delay. Step size for the sliding window was one-fifth of the window length. (Right) Model performance with varying kinematic data delays (50â350 samples) for participant 4 with a fixed input length of 250 samples. performed best among them, achieving a PCC of 0.8728 and RMSE of 0.1908 for participant 4âonly 0.08 PCC lower than the multimodal result. The transformer baseline followed (PCC = 0.8353, RMSE = 0.2208), while EEGNet, DeepConvNet, and TCN showed comparable results (PCC = 0.76â0.78, RMSE = 0.24â0.26). Participant 3 yielded the weakest EEG- only results, with PCC around 0.7 for both our model and the transformer, and RMSE near 0.3. TCN performed worst overall (PCC = 0.6036, RMSE = 0.3263). These results indicate that our model provides a significant advantage over prior EEG-only models (EEGNet, DeepCon- vNet) for hand kinematics decoding, with an average PCC im- provement of approximately 0.05, despite remaining sensitive to inter-subject variability. 2) Cross-Subject Decoding: As shown in Fig. 4 (Right), cross-subject decoding resulted in lower PCC and higher RMSE than within-subject decoding. For the EEGâEMG model, cross-subject decoding showed an average PCC decrease of 0.09 and RMSE increase of 0.08. The model trained on participant 9 generalized best (PCC = 0.9063, RMSE = 0.1781), showing the smallest drop from its within-subject performance. The model trained on participant 4 performed worst (PCC = 0.8388, RMSE = 0.2278). The impact was more severe for the EEG-only model, with an average PCC drop of 0.23 and RMSE increase of 0.1. The best cross-subject performance came from the model trained on participant 4 (PCC = 0.6646, RMSE = 0.3139), still worse than the lowest within-subject result. The poorest performance was from the model trained on participant 7 (PCC = 0.4985, RMSE = 0.3894). These findings suggest EEGâEMG multimodal signals are more robust to inter-subject variability. Both models showed a parallel trend between within-subject and cross-subject performance, implying that datasets yielding strong within- subject decoding also tend to support better cross-subject generalization under our deep learning architecture. 3) Effect of Window Size and Delay: We first study two practical factors that directly shape decoding quality: the EEG input window length and the temporal delay between EEG and kinematics. All results in this ablation are obtained in the Fig. 6: MuJoCo robotic arm simulation (top). Trials with relatively clear trajectories (high PCC) from participants 3, 4, 5, 7, and 9, reconstructing the grasp-and-lift movement. For participant 3, the trajectory was rotated 180° around the z-axis to resolve visual overlap (bottom). EEG-only setting using participant 4, where PCC/RMSE are computed on the 3D midpoint trajectory between the thumb and index finger. Window size. Fig 5 (Left) shows a clear trade-off between temporal context and noise accumulation. As the window length increases, PCC improves steadily and reaches a turning point at 200 samples, indicating that short windows provide insufficient context for stable trajectory estimation. Beyond 200 samples, performance gradually declines without sharp oscillations, suggesting that overly long windows may dilute informative cues and introduce non-stationary interference. The lowest accuracy occurs at 50 samples (PCC = 0.8004), while the best performance is achieved at 750 samples (PCC = 0.8916). The axis-wise trends follow the same pattern, with the Y-axis consistently tracking the overall improvement most closely. Delay. Fig 5 (Right) evaluates the influence of kinematic delay. Performance peaks at a 200 ms delay and then degrades as the delay increases, consistent with the intuition that exces- sive lag weakens the time alignment between neural activity and motor execution. The best overall accuracy and error are obtained at 200 ms (PCC = 0.8728, RMSE = 0.1908), whereas the worst decoding is observed at 700 ms (PCC = 0.7993, RMSE = 0.2351). Across delays, axis-specific performance consistently follows Y > X > Z. The Y-axis achieves its highest PCC at 100 ms (0.9241), the X-axis peaks at 200 ms (0.8376), while the Z-axis is less stable and reaches its maximum at 300 ms (0.7554). B. Hand Kinematics Reconstruction on a Robotic Arm As shown in Fig. 6, EEG-only decoding occasion- ally produced trajectory reversals during all movement phasesâparticularly in grasping/return (participants 3, 7) and lift-hold stages (participants 3, 4, 9). Fluctuations during hold- Fig. 7: Performance of the test dataset for participant 4 after copilot filtering, showing the proportion of retained points (bottom-left graph). Peripheral plots show trajectory changes before/after filtering for a randomly selected trial, compared to the ground truth. The red box indicates the overall PCC for a single filtered trial. ing indicate residual precision limitations. Incorporating EMG significantly reduced the magnitude of reversals, especially for participants 3 and 4. While the trajectories in Fig. 6 appear relatively complete, this applies only to trials with very high PCC (>0.9). For most trajectories (PCC 0.83â0.88), reversals are more common and pronounced. To address this, we introduced the copilot filtering module, whose effect is shown in Fig. 7. The graph shows that PCC continuously increased as the retention ratio decreased, up to a threshold. When the retained points dropped to 27.88% (672/2410), performance metrics (including PCC) changed significantly and irregularly. This indicates that applying an appropriate confidence threshold in the copilot can enhance decoding quality, but excessive filter- ing is detrimental. The peripheral trajectory plots demonstrate that obvious reversals are filtered out, resulting in a clearer grasp-and-lift movement profile. In summary, copilot-filtered, EEG-based decoding trajec- tories can roughly reconstruct the grasp-and-lift motion and generate a stable robotic arm trajectory. However, for low- accuracy trajectories, the copilot currently cannot correct pointsâit can only filter unreliable ones. Higher-precision trajectory reconstruction from non-invasive EEG still requires more robust decoding models and potentially improved hard- ware. VI. CONCLUSION AND FUTURE WORK This paper proposed a CNNâattention hybrid model for predicting hand kinematic trajectories from movement-related EEG or combined EEGâEMG signals. We also designed a copilot framework to filter decoding points, improving the quality of EEG-only trajectories. Future work will integrate additional sensors to provide assistance tailored to different motion patterns and explore more advanced methods for decoding EEG and controlling arms in a more robust fash- ion [39]. REFERENCES [1] J. Wolpaw and E. W. Wolpaw, BrainâComputer Interfaces: Principles and Practice. Oxford University Press, 01 2012. [2] P. Brodal, The central nervous system : structure and function / Per Brodal., 4th ed., 2010. [3] Z. Yin, J. K. Liu, and K. Kornysheva, âDistributed neural dynamics underlie the shift from movement preparation to execution,â bioRxiv, Dec. 2025. [4] J. R. Wolpaw, J. D. R. Mill Ě an, and N. F. Ramsey, âBrain-computer interfaces: Definitions and principles,â Handbook of clinical neurology, vol. 168, p. 15â23, 2020. [5] Z. Yu, J. K. Liu, S. Jia, Y. Zhang, Y. Zheng, Y. Tian, and T. Huang, âToward the Next Generation of Retinal Neuroprosthesis: Visual Com- putation with Spikes,â Engineering, vol. 6, no. 4, p. 449â461, Apr. 2020. [6] M.-H. Lee, O.-Y. Kwon, Y.-J. Kim, H.-K. Kim, Y.-E. Lee, J. Williamson, S. Fazli, and S.-W. Lee, âEeg dataset and openbmi toolbox for three bci paradigms: an investigation into bci illiteracy,â Gigascience, vol. 8, no. 5, 2019. [7] B. Z. Allison, S. Dunne, R. Leeb, J. D. R. Milln, and A. Nijholt, Towards Practical Brain-Computer Interfaces: Bridging the Gap from Research to Real-World Applications. Springer, 2012. [8] R. Portillo-Lara, B. Tahirbegi, C. Chapman, J. Goding, and R. Green, âMind the gap: State-of-the-art technologies and applications for eeg- based brainâcomputer interfaces,â APL Bioengineering, vol. 5, p. 031507, 09 2021. [9] I. H. de Oliveira and A. C. Rodrigues, âEmpirical comparison of deep learning methods for eeg decoding,â Frontiers in neuroscience, vol. 16, p. 1003984, 2023. [10] R. T. Schirrmeister, J. T. Springenberg, L. D. J. Fiederer, M. Glasstetter, K. Eggensperger, M. Tangermann, F. Hutter, W. Burgard, and T. Ball, âDeep learning with convolutional neural networks for eeg decoding and visualization,â Human brain mapping, vol. 38, no. 11, p. 5391â5420, 2017. [11] V. J. Lawhern, A. J. Solon, N. R. Waytowich, S. M. Gordon, C. P. Hung, and B. J. Lance, âEegnet: a compact convolutional neural network for eeg-based brain-computer interfaces,â Journal of neural engineering, vol. 15, no. 5, p. 56013, 2018. [12] T. M. Ingolfsson, M. Hersche, X. Wang, N. Kobayashi, L. Cavigelli, and L. Benini, âEeg-tcnet: An accurate temporal convolutional network for embedded motor-imagery brain-machine interfaces,â in Conference proceedings - IEEE International Conference on Systems, Man, and Cybernetics, vol. 2020-. IEEE, 2020, p. 2958â2965. [13] Y. Song, X. Jia, L. Yang, and L. Xie, âTransformer-based spatial- temporal feature learning for eeg decoding,â 2021. [14] G. Cisotto, A. Zanga, J. Chlebus, I. Zoppis, S. Manzoni, and U. Markowska-Kaczmar, âComparison of attention-based deep learning models for eeg classification,â 2020. [15] Y. Song, Q. Zheng, B. Liu, and X. Gao, âEeg conformer: Convolutional transformer for eeg decoding and visualization,â IEEE Transactions on Neural Systems and Rehabilitation Engineering, vol. P, p. 1â1, 12 2022. [16] D. Wang and Q. Wei, âSmanet: A model combining sincnet, multi- branch spatial-temporal cnn, and attention mechanism for motor imagery bci,â IEEE transactions on neural systems and rehabilitation engineer- ing, vol. 33, p. 1497â1508, 2025. [17] R. J. Kobler, A. I. Sburlea, V. Mondini, M. Hirata, and G. R. M Ě uller- Putz, âDistance- and speed-informed kinematics decoding improves m/eeg based upper-limb movement decoder accuracy,â Journal of neural engineering, vol. 17, no. 5, p. 56027, 2020. [18] T. Trammel, N. Khodayari, S. J. Luck, M. J. Traxler, and T. Y. Swaab, âDecoding semantic relatedness and prediction from eeg: A classification method comparison,â NeuroImage, vol. 277, p. 120268, 2023. [19] M. Jochumsen, I. K. Niazi, N. Mrachacz-Kersting, D. Farina, and K. Dremstrup, âDetection and classification of movement-related cortical potentials associated with task force and speed,â Journal of neural engineering, vol. 10, no. 5, p. 56015, 2013. [20] A. Jain and L. Kumar, âEeg cortical source feature based hand kine- matics decoding using residual cnn-lstm neural network,â in 2023 45th Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC). IEEE, Jul. 2023, p. 1â4. [21] A. Y. Paek, H. A. Agashe, and J. L. Contreras-Vidal, âDecoding repetitive finger movements with brain activity acquired via non-invasive electroencephalography,â Frontiers in neuroengineering, vol. 7, p. 3, 2014. [22] Y. Jia and C. W. Tyler, âMeasurement of saccadic eye movements by electrooculography for simultaneous eeg recording,â Behavior research methods, vol. 51, no. 5, p. 2139â2151, 2019. [23] A. Chaddad, Y. Wu, R. Kateb, and A. Bouridane, âElectroencephalogra- phy signal processing: A comprehensive review and analysis of methods and techniques,â Sensors, vol. 23, no. 14, p. 6434, 2023. [24] K. Erat, E. B. S ̧ ahin, F. Do Ě gan, N. Merdano Ě glu, A. Akcakaya, and P. O. Durdu, âEmotion recognition with eeg-based brain-computer interfaces: a systematic literature review,â Multimedia tools and applications, vol. 83, no. 33, p. 79 647â79 694, 2024. [25] T. J. Bradberry, R. J. Gentili, and J. L. Contreras-Vidal, âReconstructing three-dimensional hand movements from noninvasive electroencephalo- graphic signals,â The Journal of neuroscience, vol. 30, no. 9, p. 3432â 3437, 2010. [26] P. Ofner, A. Schwarz, J. Pereira, and G. R. M Ě uller-Putz, âUpper limb movements can be decoded from the time-domain of low-frequency eeg,â PloS one, vol. 12, no. 8, p. e0182578, 2017. [27] M. Saini, A. Jain, S. P. Muthukrishnan, S. Bhasin, S. Roy, and L. Kumar, âBicurnet: Premovement eeg-based neural decoder for biceps curl trajectory estimation,â IEEE transactions on instrumentation and measurement, vol. 73, p. 1â11, 2024. [28] J.-H. Jeong, K.-H. Shim, D.-J. Kim, and S.-W. Lee, âBrain-controlled robotic arm system based on multi-directional cnn-bilstm network using eeg signals,â IEEE transactions on neural systems and rehabilitation engineering, vol. 28, no. 5, p. 1226â1238, 2020. [29] A. Jain and L. Kumar, âEsi-gal: Eeg source imaging-based kinematics parameter estimation for grasp and lift task,â arXiv, 2024. [30] Z. Wan, M. Li, S. Liu, J. Huang, H. Tan, and W. Duan, âEegformer: A transformerâbased brain activity classification method using eeg signal,â Frontiers in neuroscience, vol. 17, p. 1148855, 2023. [31] M. Singh, S. Chauhan, A. K. Rajput, I. Verma, and A. K. Tiwari, âEasm: An efficient attnsleep model for sleep apnea detection from eeg signals,â Multimedia tools and applications, vol. 84, no. 4, p. 1985â2003, 2025. [32] A. H. P. Nguyen, O. Oyefisayo, M. A. Pfeffer, and S. H. Ling, âEeg- tcntransformer: A temporal convolutional transformer for motor imagery brainâcomputer interfaces,â Signals, vol. 5, no. 3, p. 605â632, 2024. [33] S. Bagchi and D. R. Bathula, âEeg-convtransformer for single-trial eeg based visual stimuli classification,â 2021. [34] J. Y. Lee, S. Lee, A. Mishra, X. Yan, B. Mcmahan, B. Gaisford, C. Kobashigawa, M. Qu, C. Xie, and J. C. Kao, âBrainâcomputer interface control with artificial intelligence copilots,â Nature Machine Intelligence. [35] M. David Luciw, E. Jarocka, and B. Edin, âWay-eeg-gal: Multi-channel eeg recordings during 3,936 grasp and lift trials with varying weight and friction,â Nov 2014. [36] S. Bai, J. Z. Kolter, and V. Koltun, âAn empirical evaluation of generic convolutional and recurrent networks for sequence modeling,â 2018. [37] J. Hu, L. Shen, S. Albanie, G. Sun, and E. Wu, âSqueeze-and-excitation networks,â IEEE transactions on pattern analysis and machine intelli- gence, vol. 42, no. 8, p. 2011â2023, 2020. [38] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, âAttention is all you need,â arXiv, vol. abs/1706.03762, 2017. [39] Z. Yang, S. Guo, Y. Fang, Z. Yu, and J. K. Liu, âSpiking Variational Policy Gradient for Brain Inspired Reinforcement Learning,â IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 47, no. 3, p. 1975â1990, Mar. 2025.