Paper deep dive
EEG-JEPA: Structured Latent Prediction for EEG Foundation Models
Jinhao Li, Zhiyuan Ma, Xueqiao Han, Zhongye Xia, Xinche Zhang, Shanghong Xie, Yixuan Liu, Yongjian Li, Runmin Gan, Tianlin Huo, Sen Song
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 8/4/2026, 3:53:26 AM
Summary
The paper introduces EEG-JEPA, a structured latent-prediction framework for EEG foundation models. Unlike traditional masked waveform reconstruction, EEG-JEPA predicts contextual latent states using an exponential-moving-average (EMA) target encoder. It employs Neurotopology-Aware Multi-scale Electrode-Temporal Masking (N-MET) for structured target support and hierarchical supervision across multiple encoder depths. The model demonstrates improved performance in frozen multitask transfer and full fine-tuning on various EEG benchmarks compared to baseline methods like CBraMod.
Entities (8)
Relation Signals (7)
EEG-JEPA ā uses ā N-MET
confidence 95% Ā· EEG-JEPA organizes target design along three complementary dimensions... target support specifies where prediction occurs... through Neurotopology-Aware Multi-scale Electrode-Temporal Masking (N-MET)
EEG-JEPA ā uses ā EMA Encoder
confidence 92% Ā· contextual latent states produced by an exponential-moving-average target encoder
EEG-JEPA ā outperforms ā CBraMod
confidence 90% Ā· EEG-JEPA improves the 14-task frozen macro balanced accuracy from 40.49% to 50.42% over CBraMod-style masked waveform reconstruction.
EEG-JEPA ā trainedon ā TUEG
confidence 88% Ā· trains all objective variants on TUEG for 100 epochs
EEG-JEPA ā evaluatedon ā EEG-FM-Bench
confidence 85% Ā· evaluated on EEG-FM-Bench
EEG-JEPA ā trainedon ā TDBRAIN
confidence 85% Ā· Stage 2 continues pretraining... on a mixture of TUEG, TDBRAIN, and HBN
EEG-JEPA ā trainedon ā HBN
confidence 85% Ā· Stage 2 continues pretraining... on a mixture of TUEG, TDBRAIN, and HBN
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Electroencephalography (EEG) foundation models aim to learn reusable representations from large-scale unlabeled recordings. A common pretraining strategy is masked waveform reconstruction, but applying supervision directly to noisy EEG may encourage models to recover predictable background activity, acquisition effects, and artifacts rather than neural structure that transfers across tasks. This raises a central question: what should an EEG foundation model predict to learn transferable representations? We introduce EEG-JEPA a structured latent-prediction framework for EEG foundation modeling. Rather than reconstructing masked voltage samples, a masked context encoder and predictor infer contextual latent states produced by an exponential-moving-average target encoder that observes the complete input. EEG-JEPA organizes target design along three complementary dimensions: target content specifies what representation is predicted, target support specifies where prediction occurs over structured electrode--time regions through Neurotopology-Aware Multi-scale Electrode-Temporal Masking (N-MET), and target depth specifies at which encoder layers supervision is applied. Together, these designs shift EEG pretraining from recovering missing measurements to inferring latent states from structured electrode--time context. We evaluate EEG-JEPA through controlled objective comparisons, frozen multitask transfer, and full fine-tuning. Under the same backbone, pretraining corpus, and training duration, EEG-JEPA improves the 14-task frozen macro balanced accuracy from 40.49% to 50.42% over CBraMod-style masked waveform reconstruction. Multi-source continuation further raises this result to 52.94%, the highest average among the EEG foundation models evaluated on EEG-FM-Bench. Under protocol-matched full fine-tuning, EEG-JEPA also improves the nine-task average balanced accuracy from 68.98% to 70.65%.
Tags
Links
- Source: https://arxiv.org/abs/2608.00114v1
- Canonical: https://arxiv.org/abs/2608.00114v1
Trouble viewing inline? Open PDF directly ā
Full Text
43,059 characters extracted from source content.
Expand or collapse full text
EEG-JEPA: Structured Latent Prediction for EEG Foundation Models Jinhao Li 1,2ā , Zhiyuan Ma 1,3ā , Xueqiao Han 1ā , Zhongye Xia 1,3 , Xinche Zhang 1,3 , Shanghong Xie 1,4 , Yixuan Liu 1,3 , Yongjian Li 1,3 , Runmin Gan 1,3 , Tianlin Huo 5ā , Sen Song 1,3ā 1 Tsinghua Laboratory of Brain and Intelligence, Tsinghua University 2 School of Basic Medical Sciences, Tsinghua Medicine, Tsinghua University 3 School of Biomedical Engineering, Tsinghua Medicine, Tsinghua University 4 Academy for Advanced Interdisciplinary Studies, Peking University 5 Department of Computer Science and Technology, Tsinghua University huotianlin@gmail.com, songsen@tsinghua.edu.cn Abstract Electroencephalography (EEG) foundation models aim to learn reusable representations from large-scale unlabeled recordings. A common pretraining strategy is masked wave- form reconstruction, but applying supervision directly to noisy EEG may encourage models to recover predictable back- ground activity, acquisition effects, and artifacts rather than neural structure that transfers across tasks. This raises a cen- tral question: what should an EEG foundation model predict to learn transferable representations? We introduce EEG-JEPA, a structured latent-prediction framework for EEG foundation modeling. Rather than reconstructing masked voltage sam- ples, a masked context encoder and predictor infer contextual latent states produced by an exponential-moving-average tar- get encoder that observes the complete input. EEG-JEPA or- ganizes target design along three complementary dimensions: target content specifies what representation is predicted, tar- get support specifies where prediction occurs over structured electrodeātime regions through Neurotopology-Aware Multi- scale Electrode-Temporal Masking (N-MET), and target depth specifies at which encoder layers supervision is applied. To- gether, these designs shift EEG pretraining from recovering missing measurements to inferring latent states from struc- tured electrodeātime context. We evaluate EEG-JEPA through controlled objective comparisons, frozen multitask transfer, and full fine-tuning. Under the same backbone, pretraining corpus, and training duration, EEG-JEPA improves the 14- task frozen macro balanced accuracy from 40.49% to 50.42% over CBraMod-style masked waveform reconstruction. Multi- source continuation further raises this result to 52.94%, the highest average among the EEG foundation models evalu- ated on EEG-FM-Bench. Under protocol-matched full fine- tuning, EEG-JEPA also improves the nine-task average bal- anced accuracy from 68.98% to 70.65%. Code is available at https://github.com/SWF-hao/EEG-JEPA-official. Introduction Electroencephalography (EEG) provides a non-invasive win- dow into brain activity and underpins applications in clinical monitoring, sleep analysis, brainācomputer interfaces, and ā These authors contributed equally. ā Corresponding author. cognitive neuroscience (Obeid and Picone 2016; Khalighi et al. 2016; Schalk et al. 2004; Grootswagers et al. 2022). Large-scale self-supervised pretraining offers a promising route to learning representations that transfer across these diverse settings (Kostas, Aroca-Ouellette, and Rudzicz 2021; Jiang, Zhao, and Lu 2024; Wang et al. 2025). Yet EEG presents distinctive challenges for foundation modeling. Scalp recordings have a low signal-to-noise ratio and non- stationary statistics, vary substantially across subjects, sites, and acquisition systems, and combine task-relevant neural dynamics with ongoing background activity, device effects, and physiological or environmental artifacts (Lotte et al. 2018; Melnik et al. 2017; Jiang, Bian, and Tian 2019). More- over, downstream tasks depend on structure at different tem- poral and spatial scales: transient waveforms may be critical for event detection, rhythmic dynamics may characterize mo- tor, cognitive, or sleep states, and informative activity may be distributed across multiple channels (Farwell and Donchin 1988; Pfurtscheller and Da Silva 1999; Khalighi et al. 2016; Ma et al. 2026). An effective EEG foundation model must therefore capture reusable neural structure across these scales rather than regularities specific to individual datasets or tasks. This requirement raises a central objective-design question: what should an EEG foundation model be trained to predict? Current EEG foundation models mainly rely on con- trastive learning, masked signal reconstruction, discrete to- ken prediction, representation alignment, or combinations of these objectives (Kostas, Aroca-Ouellette, and Rudzicz 2021; Yang, Westover, and Sun 2023; Jiang, Zhao, and Lu 2024; Wang et al. 2024, 2025; Jiang et al. 2025). Although these approaches have enabled large-scale EEG pretraining, they leave a fundamental question about what information the model is encouraged to retain. Reconstruction-based meth- ods reward the recovery of all predictable signal compo- nents, while token-based methods inherit the information preserved by their tokenizers. For low-SNR EEG, both may emphasize background activity, subject- or device-specific patterns, and artifacts that are predictable but do not trans- fer reliably across tasks. Contrastive and alignment-based methods avoid exact signal recovery, but instead depend on the definitions of positive pairs, negative samples, and signal arXiv:2608.00114v1 [eess.SP] 31 Jul 2026 transformations (Kostas, Aroca-Ouellette, and Rudzicz 2021; Wang et al. 2024). These assumptions are difficult to make task-independent: temporal shifts, channel perturbations, or frequency transformations may preserve the information re- quired by one EEG task while altering that required by an- other. Hybrid objectives combine reconstruction and align- ment but do not resolve which neural information should dominate the learned representation. Latent prediction offers an alternative by predicting contextual representations rather than raw observations, but existing EEG studies have not yet demonstrated consistently stronger transfer across diverse downstream tasks (Mohammadi Foumani et al. 2024). This motivates a more fundamental question: how should latent targets be constructed to emphasize reusable neural structure across channels and temporal scales? To answer this question, we introduce EEG-JEPA, a struc- tured latent-prediction framework for EEG foundation mod- eling. Building on contextual latent prediction (Baevski et al. 2022; Assran et al. 2023), EEG-JEPA organizes latent-target design along three dimensions: target content specifies what representation is predicted, target support specifies where prediction occurs on the electrodeātime lattice, and target depth specifies at which encoder layers supervision is ap- plied. For target content, an exponential-moving-average en- coder processes the complete EEG crop to produce contex- tual latent states, which a masked context encoder and pre- dictor must infer from the visible electrodeātime context. For target support, Neurotopology-Aware Multi-scale Electrode- Temporal Masking (N-MET) selects structured temporal, fo- cal, regional, bilateral, and sensor-level regions rather than independent patches. For target depth, hierarchical predic- tion applies distinct targets at multiple encoder layers. To- gether, these designs shift EEG pretraining from recovering locally predictable voltage samples toward inferring struc- tured latent states across the electrodeātime lattice and rep- resentation hierarchy. Under the same backbone, Stage-1 cor- pus, and training duration, EEG-JEPA improves the 14-task frozen macro balanced accuracy from 40.49% to 50.42% over masked waveform reconstruction. Multi-source contin- uation further raises this result to 52.94%, while protocol- matched full fine-tuning improves the nine-task average BA from 68.98% to 70.65%. Our contributions are threefold: ⢠We formulate EEG foundation-model pretraining as a structured latent-prediction problem that jointly spec- ifies contextual target content, neurotopology-aware electrodeātime support, and representation depth, rather than reconstructing masked waveform samples. ⢠We instantiate this formulation with exponential-moving- average contextual targets, Neurotopology-Aware Multi- scale Electrode-Temporal Masking (N-MET), and depth- specific hierarchical prediction, requiring the model to infer latent states across temporal, focal, regional, bilat- eral, and sensor-level supports while directly supervising multiple encoder depths. ⢠We validate the proposed objective through strictly matched comparisons and progressive ablations, and demonstrate broad transferability through 14-task frozen transfer and nine-task full fine-tuning across heteroge- neous EEG applications. Related Work EEG Foundation Models. EEG foundation models seek transferable representations across tasks, datasets, and acqui- sition settings. BENDR uses contrastive predictive pretrain- ing, while LaBraM, CBraMod, and EEGPT explore discrete neural-code prediction, masked waveform reconstruction, and masked modeling with momentum-target representation alignment, respectively (Kostas, Aroca-Ouellette, and Rudz- icz 2021; Jiang, Zhao, and Lu 2024; Wang et al. 2025, 2024). Other work addresses heterogeneous inputs: BIOT supports cross-data biosignal learning, NeuroLM connects EEG with language for multitask learning, and REVE enables transfer across electrode montages (Yang, Westover, and Sun 2023; Jiang et al. 2025; El Ouahidi et al. 2025). Together, these studies establish the feasibility of EEG pretraining. However, because they differ in pretraining data, tokenization, channel handling, architecture, and learning objective, the effect of target design remains difficult to isolate. Existing work has not explicitly formulated EEG pretraining as joint target de- sign over contextual content, electrodeātime support, and representation depth. Invariance and Reconstruction Objectives. Self- supervised objectives have mainly defined supervision through either agreement between related views or recon- struction of masked inputs. Contrastive and non-contrastive methods align representations of related views, using negative samples or architectural and statistical constraints to prevent representation collapse (Oord, Li, and Vinyals 2018; Grill et al. 2020; Bardes, Ponce, and LeCun 2022). In time-series and EEG learning, these views are constructed through contextual consistency or temporal, spectral, and channel transformations (Eldele et al. 2021; Yue et al. 2022; Banville et al. 2021; Mohsenvand, Izadi, and Maes 2020; Kostas, Aroca-Ouellette, and Rudzicz 2021). However, transformations that preserve information for one EEG task may remove information required by another, while negative samples may not represent physiologically distinct states. Masked reconstruction instead predicts hidden samples, patches, or tokens from visible context, providing dense supervision at masked positions (Vincent et al. 2008; Devlin et al. 2019; He et al. 2022; Jiang, Zhao, and Lu 2024; Wang et al. 2025). For EEG, however, such targets contain background rhythms, device effects, and artifacts alongside transferable neural structure. Temporal continuity and cross-channel redundancy may also allow masked content to be recovered from nearby observations without learning higher-level representations. Hybrid objectives combine representation alignment and reconstruction, but do not determine which variations should be ignored and which information should be retained. This limitation motivates predicting targets in a learned representation space. Latent Predictive Learning. Latent predictive learning replaces agreement between predefined views and input re- construction with the prediction of learned target represen- tations from visible context. Data2vec uses an exponential- x-encodery-encoder D(sā,sįµ§) sā sįµ§ (a) Invariance-based Representation Learning decoder x-encoder D(Å·,y) Å· (b)Generative Reconstruction-based Learning predictor x-encodery-encoder D(Åįµ§,sįµ§) Åįµ§ sįµ§ (d) Predictive Latent Representation Learning x-encoder y-encoder D(sā,sįµ§) sā sįµ§ decoder Å· D(Å·,y) (c) Hybrid Contrastive-Generative Learning Raw EEG Wave D(sā,sįµ§) Supervision On Raw EEG D(sā,sįµ§) Supervision On Latent sā Latent Representation Condition x y x y x y x y x z z z z Figure 1: Comparison of major EEG self-supervised learning paradigms by supervision space. (a) Invariance-based methods align latent representations of two EEG views. (b) Generative methods reconstruct raw EEG signals from a corrupted input. (c) Hybrid methods combine waveform reconstruction with latent alignment. (d) Predictive latent methods infer the representation of a target EEG view from contextual representations and conditioning information, avoiding direct waveform reconstruction. Figure 2: Neurotopology-Aware Multi-scale Electrode- Temporal Masking (N-MET) selects structured target loca- tions on the electrodeātime lattice. moving-average teacher that observes the complete input to generate contextual targets, whereas I-JEPA predicts repre- sentations of masked target regions from surrounding context and positional information (Baevski et al. 2022; Assran et al. 2023). EEG2Rep extends this approach to EEG by predicting masked latent representations rather than reconstructing raw signals (Mohammadi Foumani et al. 2024). Together, these studies establish latent prediction as a viable alternative to direct signal reconstruction. However, they do not jointly de- fine the target representation, the electrodeātime positions to be predicted, and the encoder depths at which prediction tar- gets are imposed. EEG-JEPA addresses this gap by explicitly specifying what is predicted, where prediction is performed, and at which encoder depths supervision is applied. EEG-JEPA: Structured Latent Predictive Learning Framework Overview EEG-JEPA learns by predicting full-context latent states rather than masked waveform values. Given a complete EEG crop x, N-MET first selects a set of target positionsT on the electrodeātime lattice. The target encoder observes the com- plete crop and produces reference representations at these positions. In parallel, the context encoder processes the same crop with the target waveform patches zeroed. The predictor receives only context representations from visible positions, together with a positional query identifying each target loca- tion, and estimates the corresponding full-context represen- tation. Gradients update the context encoder and predictor, while the target encoder follows the context encoder through an exponential moving average. This prediction problem is governed by three coupled de- sign decisions: target content specifies which latent rep- resentation is predicted, target support specifies which electrodeātime positions must be inferred, and target depth specifies where along the encoder hierarchy prediction is en- forced. The subsections instantiate these decisions for EEG. Target Content: EMA Contextual States For a target electrodeātime position t and encoder depth ā, EEG-JEPA defines the prediction target as y (ā) t = sg LN f (ā) Ģ Īø (x) t ,(1) where f Ģ Īø is a target encoder that receives the complete crop, LN controls target scale, and sg stops gradients through the target branch. Unlike a waveform patch, y (ā) t is a contextual representation: it describes position t after information from the complete crop has been integrated. The target param- eters are updated from the context encoder parameters by exponential moving average. The context encoder f Īø instead receives Ģx = ZeroMask(x,T ), h (ā) Īø = f (ā) Īø ( Ģx),(2) where ZeroMask replaces every target waveform patch with a fixed all-zero vector before convolutional and spectral patch embedding. Although the context encoder retains the com- plete electrodeātime lattice, only representations at visible positionsC are gathered into z C and passed to the predictor. Representations at target positions are therefore excluded from the prediction input. Each target location is specified by m t = e q + p 2D t ,(3) Context Encoder Multi-levels fusion MLP PĻ Multi-level Predictor stop-grad Context loss Ī»dĀ· Smooth L1 stop-grad Prediction loss Smooth L1 ā= 1 ķæ % ā"# $ ā %&% ā +ķ '%( ā '%( ā +ķ )*+ ā )*+ +ķ ',) ā ',) Patchtify Patchtify EĪø # EĪø Target Encoder EMA Variance Covariance N-MET replace Covariance VCReg Variance Context Token Target Token PaddingTarget WaveForm TargetQueryToken (Learned Target Query + 2D Position) Hierarchical Latent Hierarchical Latent EEGSignal Figure 3: Overview of the EEG-JEPA pretraining framework. N-MET selects structured target locations on the electrodeātime lattice. The masked crop is processed by the context encoder, while an EMA-updated target encoder receives the complete crop and produces stop-gradient contextual targets at multiple depths. A shared predictor combines hierarchical context features with learned position-conditioned target queries to predict the corresponding L3, L6, L9, and final-layer target representations. Training jointly minimizes latent prediction and context-consistency losses, together with variance and covariance regularization. where e q is a learned target-query token and p 2D t is a fixed electrodeātime sineācosine positional embedding. The pre- dictor estimates Ėy (ā) t = g (ā) Ļ (z C ,m t ).(4) The query reveals where a target is located but not its sig- nal content. Consequently, the prediction must be formed from visible EEG context, while the full-input EMA encoder supplies the reference state. Target Support: Neurotopology-Aware Multi-scale Electrode-Temporal Masking (N-MET) Having defined what is predicted, we next specify where prediction is required. EEG patches form an electrodeātime lattice whose axes have different semantics. A target may represent a missing event interval, a focal electrodeātime pattern, a regional scalp field, a bilateral relation, or an ab- sent sensor. Independently sampled patches do not encode these structures and, because of EEGās temporal continu- ity and cross-channel redundancy, may be predictable from immediate neighbors alone. N-MET constructs T by combining six electrodeā temporal primitives (Figure 2 and Table 1. For each crop, primitives are sampled until 40ā55% of valid tokens are se- lected; padded positions are excluded from both attention and loss. The mixture emphasizes temporal gaps while retain- ing focal, regional, bilateral, and sensor-dropout supports, thereby creating prediction problems across multiple spatial and temporal scales. Topology-paired primitives probe bilateral correspon- dence and asymmetry, with full-pair masking requiring broader context than time-limited stripes. Full-channel mask- ing simulates sensor loss. These masking patterns allow the model to learn temporal continuity, spatial correlation, and sensor robustness. PrimitiveMass Design motivation Temporal stripe 35.0% Missing event interval; infer morphology and rhythm con- tinuity. Channel stripe 17.5% Regional sensor loss; infer neighboring scalp fields. Local block17.5% Focal activity localized in electrode and time. Topo-pair stripe 10.0% Bilateral correspondence and hemispheric asymmetry. Full channel15.0% Bad or absent electrode. Full topo-pair5.0% Bilateral dropout requiring global context. Table 1: N-MET primitives and conditional structure. Target Depth: Hierarchical Supervision Target content and support define prediction at a particular encoder layer, but supervision applied only to the final layer constrains only the endpoint of the representation hierarchy. EEG-JEPA therefore predicts target states at L3, L6, L9, and the final layer, spanning relatively local signal structure and deeper contextual organization. Concretely, the visible context representation z C used above is constructed by concatenating the 200-dimensional online states from these four depths and fusing them with an 800ā 400ā 200 MLP. The fused states and target queries are processed by a shared 100-dimensional predictor trunk. Four independent 100ā 200 heads then estimate the cor- responding L3, L6, L9, and final-layer targets. The shared trunk couples information across prediction depths, while the separate heads preserve depth-specific target spaces. Besides supervising the complete encoder, this design makes intermediate prefixes usable: frozen L3, L6, and L9 evaluations terminate the encoder at the corresponding block, providing a progressive accuracyācapacity trade-off. Training objective. Let D = 3, 6, 9, 12 denote the su- pervised depths andP =T āŖC the valid target and context positions. We define a unified latent prediction loss L latent = 1 |D| X āāD mean iāP w i Ļ Ėy (ā) i ā y (ā) i , (5) where Ļ is the Smooth L1 loss and w i balances supervision over masked targets and visible context positions. Invalid or padded positions are excluded. We additionally apply varianceācovariance regularization (VCReg) to the valid predictor outputs: L VCReg = Ī» var L var + Ī» cov L cov .(6) The variance component prevents collapsed feature dimen- sions, while the covariance component reduces redundancy. The complete objective is L =L latent +L VCReg .(7) Experiments Pretraining Setup Stage 1 uses CBraModās fixed 19-channel tokenization and spatio-temporal encoder and trains all objective variants on TUEG for 100 epochs (Wang et al. 2025). All variants share the same channel layout, patch definition, encoder ca- pacity, pretraining corpus, and training duration, enabling controlled comparisons of raw and latent targets, indepen- dently optimized and EMA target encoders, structured tar- get support, hierarchical prediction, context consistency, and VCReg. Starting from the Stage-1 EEG-JEPA checkpoint, Stage 2 continues pretraining for 25 epochs on a mixture of TUEG, TDBRAIN, and HBN (Table 2). Each training exam- ple is a valid 30-s crop. Stage 1 samples 1,115,752 crops per epoch, whereas Stage 2 samples 1,200,000 using weights of 75%, 15%, and 10% for TUEG, TDBRAIN, and HBN. SplitSourcesRatio Records 30s crops Stage 1TUEG100% 39,758 6,535,247 Stage 2 totalTUEG/TDBRAIN/HBN 100% 44,486 6,648,828 TUEG partTUEG75% 39,758 6,535,247 TDBRAIN part TDBRAIN BVA15% 2,668 48,542 HBN partHBN R810% 2,060 65,039 Table 2: Two-stage trainnig corpus; ratios are stage-specific sampling weights. Data Preprocessing Pretraining follows the CBraMod preprocessing protocol (Wang et al. 2025). Recordings are resampled to 200 Hz, mapped to a canonical 19-channel montage with a fixed chan- nel order, segmented into 30-s crops, and divided into non- overlapping 1-s patches. For downstream full fine-tuning, each benchmark retains its prescribed channel set and or- dering rather than being mapped to the pretraining montage, thereby preserving native recording configurations. Further details on filtering, channel handling, quality control, and segmentation are provided in the supplementary material. TaskDataset Ch. Duration Samples Orig. Hz Classes Abnormal detection TUAB1610s409,4552502 Event typeTUEV165s112,4912506 Motor imageryPhysioMI 644s9,8371604 Motor imageryBCI-2a224s5,1842504 Emotion recognition FACED3210s10,3322509 Sleep stagingISRUC630s89,2402005 Mental disorderMumtaz195s7,1432562 Mental stressMAT205s1,7075002 Imagined speechBCI20-3 643s6,0002565 Table 4: Full fine-tuning tasks and datasets. Evaluation Protocol We evaluate transfer under two regimes. For full fine-tuning, we follow the CBraMod protocol on nine datasets span- ning clinical, motor-imagery, affective, sleep, stress, and imagined-speech tasks (Table 4). Each dataset retains its pre- scribed channel set and ordering rather than being mapped to the canonical 19-channel pretraining montage. We compare EEG-JEPA with supervised EEG decoders and EEG foun- dation models (Lawhern et al. 2018; Song et al. 2023; Wang et al. 2025). External baseline results are taken from REVE (El Ouahidi et al. 2025). For frozen transfer, we follow the 14-task EEG-FM-Bench multitask protocol, which covers clinical, sleep, motor- imagery, affective, workload, seizure, depression, and vi- sual EEG tasks (Xiong et al. 2025). We compare against the released foundation-model baselines spanning contrastive, reconstruction, and hybrid pretraining. We also include matched controls on the CBraMod backbone that vary the target-encoder update rule and target support. The encoder and its normalization statistics remain fixed, while mean- pooled token representations are passed to dataset-specific MLP heads that are trained jointly across tasks. We report balanced accuracy (BA) averaged over five downstream runs. Scale ModelPretraining paradigmParams (M) Macro BA (%) Tiny EEG-JEPA-L3 Latent predictive1.2648.87 Small EEG-JEPA-L6 Latent predictive2.4749.06 Base BIOTContrastive3.1947.37 EEG-JEPA-L9 Latent predictive3.6849.70 BENDRContrastive3.9734.14 Large EEG-JEPA-final Latent predictive4.9252.94 CBraModMasked reconstruction4.9240.49 LaBraMMasked token prediction5.8244.91 CSBrainMasked reconstruction8.8645.44 Huge EEGPTMasked + latent alignment 25.2952.15 REVEMasked reconstruction69.1951.50 Table 3: Frozen 14-task macro BA versus encoder size and pretraining paradigm. Main Results Frozen multitask transfer. Frozen-transfer results are summarized in Table 3 and Figure 4c, while task-level BAs are shown in Figure 4a, with detailed results provided in MethodTUABTUEV PhysioMI BCI-2aFACEDISRUCMumtazMATBCI20-3Avg. EEGNet76.42±0.36 38.76±1.43 58.14±1.25 44.82±0.94 40.90±1.22 71.54±1.21 92.32±1.04 67.70±1.16 44.13±0.96 59.41±0.37 EEGConformer 77.58±0.49 40.74±1.64 60.49±1.04 46.96±1.06 45.59±1.25 74.00±1.33 93.08±1.17 68.05±1.23 45.06±1.33 61.28±0.44 SPaRCNet78.96±0.18 41.61±2.62 59.32±1.52 46.35±1.17 46.73±1.55 74.87±0.75 93.16±0.95 68.79±1.07 44.26±1.56 61.56±0.47 ContraWR77.46±0.41 43.84±3.49 58.92±1.33 46.78±1.25 48.87±0.78 74.02±1.26 91.95±1.15 66.31±0.97 42.57±1.62 61.19±0.53 CNN-Transformer 77.77±0.22 40.87±1.61 60.53±1.18 46.00±1.08 46.97±1.32 73.63±0.87 93.05±0.68 67.79±2.68 45.33±0.92 61.33±0.45 FFCL78.48±0.38 39.79±1.04 57.26±0.92 44.70±1.43 46.73±1.58 72.77±1.82 93.14±0.38 67.98±1.42 46.78±1.97 60.85±0.44 ST-Transformer 79.66±0.23 39.84±2.28 60.35±0.81 45.75±1.45 48.10±0.79 73.81±2.05 91.35±1.03 66.31±1.73 41.26±1.22 60.71±0.48 BIOT79.59±0.57 52.81±2.25 61.53±1.54 47.48±0.93 51.18±1.18 75.27±1.21 93.58±0.52 68.75±1.86 49.20±0.86 64.38±0.44 LaBraM-Base81.40±0.19 64.09±0.65 61.73±1.22 48.69±0.85 52.73±1.07 76.33±1.02 94.09±0.79 69.09±1.25 50.60±1.55 66.53±0.31 CBraMod82.89±0.2266.71±1.0764.17±0.9151.38±0.6655.09±0.8978.65±1.1095.60±0.5672.56±1.3253.73±1.0868.98±0.31 EEG-JEPA83.19±0.31 67.31±1.52 64.82±0.96 55.38±1.38 56.09±1.07 79.10±1.24 96.00±0.61 76.06±2.08 57.86±1.76 70.65±0.44 Table 5: Single-task full fine-tuning BA (%, mean± std). '40(- ! 2+1!5 3$/!**!*!,"$#""2/!"4 (&2/$-*(01(""-+.!/(0-,!"/-00,(,$%2**%(,$12,(,&1!0)0 '40(- ! 2+1!5 3$/!**!*!,"$#""2/!"4 (&2/$-*(01(""-+.!/(0-,!"/-00,(,$%2**%(,$12,(,&1!0)0 " "! " Balanced Accuracy over 14 tasks (%) Overall Balanced Accuracy ModelSize (M) L9 L6 L3 ļ¼aļ¼ ļ¼bļ¼ ļ¼cļ¼ Figure 4: Results. (a) Frozen 14-task EEG foundation-model comparison on EEG-FM-Bench. (b) Full fine-tuning 9-task comparison. Different task suites are used because evalua- tions follow standard benchmarks; see Sec. Evaluation Pro- tocol for details. (c) Parameter-efficiency comparison for a. the appendix. With the encoder and its normalization statis- tics fixed, EEG-JEPA achieves the highest 14-task macro BA of 52.94%, exceeding EEGPT and REVE by 0.79 and 1.44 percentage points, respectively, and ranks among the top two methods on 10 of the 14 tasks. Its performance across clinical, sleep, motor-imagery, affective, workload, seizure, depression, and visual EEG tasks shows that a single frozen representation transfers broadly without task-specific encoder adaptation. Single-task full fine-tuning. Protocol-matched full fine- tuning gives EEG-JEPA the highest mean BA on all nine tasks, improving the nine-task average from 68.98% for CBraMod to 70.65%, an absolute gain of 1.67 percentage points (Table 5). The largest improvements over CBraMod occur on BCI20-3 (+4.13 points), BCI-2a (+4.00), and MAT (+3.50), showing that the gains extend beyond clinical EEG to imagined speech, motor imagery, and stress. Parameter-efficient transfer. Hierarchical supervision makes intermediate encoder prefixes usable at different model sizes. From the same pretrained encoder, the 1.26M- parameter L3, 2.47M-parameter L6, and 3.68M-parameter L9 prefixes achieve macro BAs of 48.87%, 49.06%, and 49.70%, respectively, while the full 4.92M-parameter en- coder reaches 52.94%, compared with 40.49% for the parameter-matched CBraMod encoder (Table 3). A single pretrained model therefore provides an accuracyāsize trade- off as additional encoder blocks are used. Objective variantSpaceMaskTarget Multi-level Context VCReg 2-stage Macro BA CBraMod MAE controlRaw EEG Random patch ā40.49± 0.32 Naive EEG-I-JEPA control LatentRandom block Ind.44.79± 0.91 Naive EEG-V-JEPA control LatentRandom block EMAā45.12± 0.43 Naive EEG-V-JEPA control LatentN-METEMAā46.47± 0.36 + multi-level predictionLatentN-METEMAā50.13± 0.58 + context consistencyLatentN-METEMAā50.42± 0.58 - VCReg controlLatentN-METEMAā50.09± 0.32 Full EEG-JEPALatentN-METEMAā 52.94± 0.30 Table 6: Progressive ablation of the EEG-JEPA objective design; BA is the 14-task macro average. Ind. denotes an independently optimized target encoder without EMA; EMA denotes an exponential-moving-average target encoder. Ablation Study. Table 6 evaluates structured latent pre- diction under a common Stage-1 setting with the same backbone, corpus, encoder capacity, and training duration. The CBraMod MAE control serves as the masked-waveform baseline, while independent- and EMA-target variants pro- vide latent-prediction controls. We then replace random- block masking with N-MET to evaluate target support, add hierarchical prediction to evaluate target depth, and ex- amine context consistency and VCReg as auxiliary objec- tives. Stage-2 continuation is reported separately because it changes the pretraining corpus. The masked-waveform baseline achieves a 14-task macro BA of 40.49%, while the independent- and EMA-target con- trols reach 44.79% and 45.12%. Under the matched EMA set- 14812 Peak layer Complexity Spectral Time morphology Time- frequency Cross- frequency Cross- channel (a) Descriptor depth 03060 Peak in L9-L12 (%) 0 3 19 20 6 10 20 21 22 24 31 52 0 3 41 44 38 55 (b) Late-depth enrichment CBraMod (MAE)EEG-I-JEPA (Ours-control)EEG-JEPA (Ours) Signal statistics Within-signal interactions Cross-channel relations Figure 5: Matched descriptor organization across 12 frozen layers. (a) Median peak layer; curves connect families within each model and coincident markers share the same peak. (b) Fraction peaking in L9āL12. Bands denote signal statistics, within-signal interactions, and cross-channel relations. ting, N-MET improves performance from 45.12% to 46.47%, and hierarchical prediction provides the largest Stage-1 gain, raising it by 3.66 points to 50.13%. Context consistency further increases performance to 50.42%, while removing VCReg reduces it to 50.09%. Multisource continuation then adds 2.52 points, reaching 52.94%. These results show that EEG latent prediction benefits not only from learned tar- gets, but also from structured target support and supervision across representation depths. Layerwise Representation Organization We next examine whether the transfer gains of EEG-JEPA are accompanied by a different organization of information across encoder depth. Following the frozen representation audit of Tang et al. (2026), we train ridge probes to predict 63 descriptors from six families on matched samples from five EEG-FM-Bench tasks. For each taskādescriptor pair, the peak layer is selected after validation against shuffled and Gaussian controls, and comparisons use 158 pairs retained across the three models. We group the descriptor families into signal statistics, within-signal interactions, and cross- channel relations. Centered linear CKA complements the probe analysis by measuring changes in sample geometry across layers on matched inputs (Kornblith et al. 2019). Distinct depth profiles emerge across the three models (Figure 5a). CBraMod is front-loaded: complexity, spectral, time-morphology, time-frequency, and cross-frequency de- scriptors peak at L1āL2, whereas cross-channel relations peak at L5. EEG-I-JEPA shifts time-frequency descriptors to L7 and cross-channel relations to L9. EEG-JEPA shows a clearer ordering, with complexity and spectral descriptors peaking at L2, time morphology at L7, and within-signal interactions and cross-channel relations around L8āL9. Late-layer peak rates further show that EEG-JEPA does not uniformly shift information deeper (Figure 5b). Within L9āL12, the rates are 14.0% for signal statistics, 41.5% for 14812 1 4 8 12 Layer EEG-JEPA (Ours) First-to-last layer CKA: 0.61 Mean adjacent-layer CKA: 0.986 14812 CBraMod (MAE) First-to-last layer CKA: 0.75 Mean adjacent-layer CKA: 0.947 14812 Layer 1 4 8 12 Layer EEG-I-JEPA (Ours-control) First-to-last layer CKA: 0.92 Mean adjacent-layer CKA: 0.997 14812 Layer 0.0 0.1 0.2 0.3 0.4 0.5 1 - CKA(L1, Lk) Cumulative representation change EEG-JEPA (Ours) CBraMod (MAE) EEG-I-JEPA (Ours-control) 0.4 0.5 0.6 0.7 0.8 0.9 1.0 Linear CKA Figure 6: Task-averaged layerwise CKA and cumulative rep- resentation change on matched inputs. within-signal interactions, and 54.8% for cross-channel rela- tions. Compared with EEG-I-JEPA, EEG-JEPA reduces late peaks for signal statistics from 20.9% to 14.0%, while in- creasing them for within-signal interactions from 26.8% to 41.5% and cross-channel relations from 51.6% to 54.8%. CBraMod reaches only 14.6% and 9.7% for the latter two groups. This pattern indicates selective late-depth enrich- ment of interaction and cross-channel information. The CKA analysis clarifies how this depth organization de- velops (Figure 6). EEG-I-JEPA changes little across depth, with a mean adjacent-layer CKA of 0.997 and an L1āL12 CKA of 0.92. CBraMod exhibits larger changes (0.947) but a smaller first-to-last change (0.75). EEG-JEPA combines high adjacent-layer similarity (0.986) with the largest first-to-last change (0.61), as its distance from L1 accumulates through- out the encoder. Thus, small changes between neighboring layers build into substantial reorganization across depth. Together, the probe and CKA results show that EEG-JEPA develops a selective depth organization: basic signal statistics remain most accessible in shallow layers, whereas within- signal interactions and cross-channel relations become more accessible in deeper representations. Conclusion We introduced EEG-JEPA, a structured latent-prediction framework that jointly defines contextual target content, neurotopology-aware electrodeātime support, and represen- tation depth. N-MET structures prediction across multiple electrodeātime supports, while hierarchical prediction su- pervises multiple encoder depths. Ablations identify target support and hierarchical prediction as the main sources of improvement, and layerwise analysis reveals increasing ac- cessibility of interaction and cross-channel information at depth. EEG-JEPA achieves the highest frozen average among the foundation models evaluated on EEG-FM-Bench and im- proves full fine-tuning across nine tasks, demonstrating broad transfer across heterogeneous EEG applications. References Assran, M.; Duval, Q.; Misra, I.; Bojanowski, P.; Vincent, P.; Rabbat, M.; LeCun, Y.; and Ballas, N. 2023. Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Baevski, A.; Hsu, W.-N.; Xu, Q.; Babu, A.; Gu, J.; and Auli, M. 2022. data2vec: A General Framework for Self- Supervised Learning in Speech, Vision and Language. In Proceedings of the 39th International Conference on Ma- chine Learning, volume 162 of Proceedings of Machine Learning Research, 1298ā1312. PMLR. Banville, H.; Chehab, O.; HyvƤrinen, A.; Engemann, D. A.; and Gramfort, A. 2021. Uncovering the Structure of Clinical EEG Signals with Self-Supervised Learning. Journal of Neural Engineering, 18(4): 046020. Bardes, A.; Ponce, J.; and LeCun, Y. 2022. VICReg: Variance-Invariance-Covariance Regularization for Self- Supervised Learning. In International Conference on Learn- ing Representations. Devlin, J.; Chang, M.-W.; Lee, K.; and Toutanova, K. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. In Proceedings of the North Amer- ican Chapter of the Association for Computational Linguis- tics. El Ouahidi, Y.; Lys, J.; Thƶlke, P.; Farrugia, N.; Pasdeloup, B.; Gripon, V.; Jerbi, K.; and Lioi, G. 2025. REVE: A Foundation Model for EEGāAdapting to Any Setup with Large-Scale Pretraining on 25,000 Subjects. In Advances in Neural Information Processing Systems, volume 38. Eldele, E.; Ragab, M.; Chen, Z.; Wu, M.; Kwoh, C.-K.; Li, X.; and Guan, C. 2021. Time-Series Representation Learning via Temporal and Contextual Contrasting. In Proceedings of the International Joint Conference on Artificial Intelligence. Farwell, L. A.; and Donchin, E. 1988. Talking off the top of your head: toward a mental prosthesis utilizing event-related brain potentials. Electroencephalography and clinical Neu- rophysiology, 70(6): 510ā523. Grill, J.-B.; Strub, F.; AltchĆ©, F.; Tallec, C.; Richemond, P.; Buchatskaya, E.; Doersch, C.; Pires, B. A.; Guo, Z.; Azar, M. G.; Piot, B.; Kavukcuoglu, K.; Munos, R.; and Valko, M. 2020. Bootstrap Your Own Latent: A New Approach to Self- Supervised Learning. In Advances in Neural Information Processing Systems. Grootswagers, T.; Zhou, I.; Robinson, A. K.; Hebart, M. N.; and Carlson, T. A. 2022. Human EEG Recordings for 1,854 Concepts Presented in Rapid Serial Visual Presen- tation Streams. Scientific Data, 9: 3. He, K.; Chen, X.; Xie, S.; Li, Y.; DollĆ”r, P.; and Girshick, R. 2022. Masked Autoencoders Are Scalable Vision Learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Jiang, W.-B.; Wang, Y.; Lu, B.-L.; and Li, D. 2025. Neu- roLM: A universal multi-task foundation model for bridging the gap between language and EEG signals. In Interna- tional conference on learning representations, volume 2025, 55436ā55457. Jiang, W.-B.; Zhao, L.-M.; and Lu, B.-L. 2024. Large Brain Model for Learning Generic Representations with Tremen- dous EEG Data in BCI. In International Conference on Learning Representations. Jiang, X.; Bian, G.-B.; and Tian, Z. 2019. Removal of arti- facts from EEG signals: a review. Sensors, 19(5): 987. Khalighi, S.; Sousa, T.; Santos, J. M.; and Nunes, U. 2016. ISRUC-Sleep: A Comprehensive Public Dataset for Sleep Researchers. Computer Methods and Programs in Biomedicine, 124: 180ā192. Kornblith, S.; Norouzi, M.; Lee, H.; and Hinton, G. 2019. Similarity of Neural Network Representations Revisited. In Proceedings of the 36th International Conference on Machine Learning, volume 97 of Proceedings of Machine Learning Research, 3519ā3529. PMLR. Kostas, D.; Aroca-Ouellette, S.; and Rudzicz, F. 2021. BENDR: Using Transformers and a Contrastive Self- Supervised Learning Task to Learn from Massive Amounts of EEG Data. Frontiers in Human Neuroscience, 15: 653659. Lawhern, V. J.; Solon, A. J.; Waytowich, N. R.; Gordon, S. M.; Hung, C. P.; and Lance, B. J. 2018. EEGNet: A Com- pact Convolutional Neural Network for EEG-Based Brain- Computer Interfaces. Journal of Neural Engineering, 15(5): 056013. Lotte, F.; Bougrain, L.; Cichocki, A.; Clerc, M.; Congedo, M.; Rakotomamonjy, A.; and Yger, F. 2018. A review of classification algorithms for EEG-based brainācomputer in- terfaces: a 10 year update. Journal of neural engineering, 15(3): 031005. Ma, Z.; Li, Z.; Qiu, Z.; Li, J.; Meng, L.; Zhang, X.; Liu, Y.; Shen, X.; and Song, S. 2026. Dsainet: An efficient dual-scale attentive interaction network for general eeg decoding. arXiv preprint arXiv:2604.18095. Melnik, A.; Legkov, P.; Izdebski, K.; KƤrcher, S. M.; Hairston, W. D.; Ferris, D. P.; and Kƶnig, P. 2017. Systems, subjects, sessions: to what extent do these factors influence EEG data? Frontiers in human neuroscience, 11: 150. Mohammadi Foumani, N.; Mackellar, G.; Ghane, S.; Irtza, S.; Nguyen, N.; and Salehi, M. 2024. Eeg2rep: enhancing self- supervised eeg representation through informative masked inputs. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 5544ā5555. Mohsenvand, M. N.; Izadi, M. R.; and Maes, P. 2020. Con- trastive representation learning for electroencephalogram classification. In Machine learning for health, 238ā253. PMLR. Obeid, I.; and Picone, J. 2016. The Temple University Hos- pital EEG Data Corpus. Frontiers in Neuroscience, 10: 196. Oord, A. v. d.; Li, Y.; and Vinyals, O. 2018. Representation learning with contrastive predictive coding. arXiv preprint arXiv:1807.03748. Pfurtscheller, G.; and Da Silva, F. L. 1999. Event-related EEG/MEG synchronization and desynchronization: basic principles. Clinical neurophysiology, 110(11): 1842ā1857. Schalk, G.; McFarland, D. J.; Hinterberger, T.; Birbaumer, N.; and Wolpaw, J. R. 2004. BCI2000: A General-Purpose Brain-Computer Interface (BCI) System. IEEE Transactions on Biomedical Engineering, 51(6): 1034ā1043. Song, Y.; Zheng, Q.; Liu, B.; and Gao, X. 2023. EEG Con- former: Convolutional Transformer for EEG Decoding and Visualization. IEEE Transactions on Neural Systems and Rehabilitation Engineering, 31: 710ā719. Tang, L.; Chen, Q.; Mei, J.; Xu, H.; Zhang, Q.; Shao, J.; Zou, N.; Hu, X.; and Liu, D. 2026. What Do EEG Foundation Models Capture from Human Brain Signals? arXiv preprint arXiv:2605.11410. Vincent, P.; Larochelle, H.; Bengio, Y.; and Manzagol, P.- A. 2008. Extracting and Composing Robust Features with Denoising Autoencoders. In Proceedings of the International Conference on Machine Learning. Wang, G.; Liu, W.; He, Y.; Xu, C.; Ma, L.; and Li, H. 2024. Eegpt: Pretrained transformer for universal and reliable rep- resentation of eeg signals. Advances in Neural Information Processing Systems, 37: 39249ā39280. Wang, J.; Zhao, S.; Luo, Z.; Zhou, Y.; Jiang, H.; Li, S.; Li, T.; and Pan, G. 2025. CBraMod: A Criss-Cross Brain Founda- tion Model for EEG Decoding. In International Conference on Learning Representations. Xiong, W.; Li, J.; Li, J.; and Zhu, K. 2025. Eeg-fm-bench: A comprehensive benchmark for the systematic evaluation of eeg foundation models. arXiv preprint arXiv:2508.17742. Yang, C.; Westover, M. B.; and Sun, J. 2023. BIOT: Biosig- nal Transformer for Cross-Data Learning in the Wild. In Advances in Neural Information Processing Systems. Yue, Z.; Wang, Y.; Duan, J.; Yang, T.; Huang, C.; Tong, Y.; and Xu, B. 2022. TS2Vec: Towards Universal Representation of Time Series. In Proceedings of the AAAI Conference on Artificial Intelligence.