Paper deep dive
Foundation Models for EEG Are Blind to Long-Range Temporal Correlations: A Spectral-Temporal Dissociation Behind Their Cross-Population Fragility
Marzieh Zare
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 8/1/2026, 10:29:12 AM
Summary
This study evaluates five EEG foundation models (REVE, LaBraM, BENDR, CBraMod, BIOT) to determine if their frozen embeddings retain long-range temporal correlations (LRTC), quantified by the detrended-fluctuation-analysis (DFA) exponent. The results indicate a spectral-temporal dissociation: spectral-input models (CBraMod, BIOT) recover the static 1/f aperiodic slope but fail to encode LRTC, while raw-waveform models (REVE, LaBraM, BENDR) recover neither. The DFA exponent is orthogonal to the aperiodic slope and is site-robust, whereas FM embeddings are dominated by recording-site artifacts. Consequently, zero-shot cross-population transfer using frozen embeddings performs at chance, while the discarded DFA exponent shows directional transfer potential.
Entities (11)
Relation Signals (7)
Foundation Model Embeddings → dominatedby → Recording Site
confidence 95% · All five were dominated by a recording-site axis (decodable at 0.98-1.00 vs. 0.500 chance)
CBraMod → encodes → 1/f Aperiodic Slope
confidence 95% · spectral-input models (CBraMod, BIOT), which recovered 1/f strongly (R^2 = 0.59-0.73)
CBraMod → failstoencode → Long-Range Temporal Correlations
confidence 95% · spectral-input models (CBraMod, BIOT), which recovered 1/f strongly ... but not DFA across cohorts
REVE → failstoencode → Long-Range Temporal Correlations
confidence 95% · None of the five FMs represented the LRTC in the temporal order. Raw-waveform models (REVE, LaBraM, BENDR) recovered neither the DFA exponent nor the 1/f slope
Long-Range Temporal Correlations → isorthogonalto → 1/f Aperiodic Slope
confidence 90% · LRTC was orthogonal to the aperiodic slope (r = -0.06)
DFA Exponent → issiterobust → Recording Site
confidence 90% · the DFA exponent they discard is site-robust (0.71)
Foundation Model Embeddings → performsatchancein → Cross-Population Transfer
confidence 90% · the frozen REVE embedding did not beat chance (W to K, 0.45)
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Objective. Electroencephalography (EEG) foundation models (FMs) are trained to reconstruct or contrastively align short patches, then pooled into a fixed embedding. We tested whether these embeddings retained the long-range temporal correlations (LRTC) quantified by the detrended-fluctuation-analysis (DFA) exponent of the alpha-band envelope, and whether it governs cross-population transfer. Approach. We probed five EEG FMs spanning raw-waveform and spectral-input architectures (REVE, LaBraM, BENDR, CBraMod, BIOT) on two out-of-distribution cohorts, comparing recovery of the DFA exponent against the static 1/f aperiodic slope. Order-preserving and residualization controls tested for pooling or aperiodic shadowing. A montage-harmonized, zero-shot transfer task compared the frozen embedding with the DFA exponent across three cohorts (adding a Western reference). Main results. None of the five FMs represented the LRTC in the temporal order. Raw-waveform models (REVE, LaBraM, BENDR) recovered neither the DFA exponent nor the 1/f slope (R^2 <= 0.12); for these three the probe is uninformative, so the dissociation is specific to the spectral-input models (CBraMod, BIOT), which recovered 1/f strongly (R^2 = 0.59-0.73) but not DFA across cohorts. A classical DFA feature recovered the exponent (R^2 = 0.32-0.38 against a 0.64 reliability ceiling), and LRTC was orthogonal to the aperiodic slope (r = -0.06). On cross-population transfer, the frozen REVE embedding did not beat chance (W to K, 0.45) and the dimensionless DFA exponent transferred directionally but not at family-wise significance; the other four did not uniformly replicate it. All five were dominated by a recording-site axis (decodable at 0.98-1.00 vs. 0.500 chance), whereas the DFA exponent they discard is site-robust (0.71).
Tags
Links
- Source: https://arxiv.org/abs/2607.24834v1
- Canonical: https://arxiv.org/abs/2607.24834v1
Trouble viewing inline? Open PDF directly →
Full Text
71,477 characters extracted from source content.
Expand or collapse full text
Foundation Models for EEG Are Blind to Long-Range Temporal Correlations: A Spectral–Temporal Dissociation Behind Their Cross-Population Fragility Marzieh Zare1,2 1Université Laval, School of Psychology, Quebec City, QC, Canada 2NeuroGenis Inc., Toronto, ON, Canada Abstract Objective. Electroencephalography (EEG) foundation models (FMs) are trained to reconstruct or contrastively align short patches, then pooled into a fixed embedding. We tested whether these embeddings retained the long-range temporal correlations (LRTC) quantified by the detrended-fluctuation-analysis (DFA) exponent of the alpha-band envelope, and whether it governs cross-population transfer. Approach. We probed five EEG FMs spanning raw-waveform and spectral-input architectures (REVE, LaBraM, BENDR, CBraMod, BIOT) on two out-of-distribution cohorts, comparing recovery of the DFA exponent against the static 1/f1/f aperiodic slope. Order-preserving and residualization controls tested for pooling or aperiodic shadowing. A montage-harmonized, zero-shot transfer task compared the frozen embedding with the DFA exponent across three cohorts (adding a Western reference). Main results. None of the five FMs represented the LRTC in the temporal order. Raw-waveform models (REVE, LaBraM, BENDR) recovered neither the DFA exponent nor the 1/f1/f slope (R2≤0.12R^2\!≤\!0.12); for these three the probe is uninformative, so the dissociation is specific to the spectral-input models (CBraMod, BIOT), which recovered 1/f1/f strongly (R2=0.59R^2\!=\!0.59–0.730.73) but not DFA across cohorts. A classical DFA feature recovered the exponent (R2=0.32R^2\!=\!0.32–0.380.38 against a 0.640.64 reliability ceiling), and LRTC was orthogonal to the aperiodic slope (r=−0.06r\!=\!-0.06). On cross-population transfer, the frozen REVE embedding did not beat chance (W→ 0.450.45) and the dimensionless DFA exponent transferred directionally but not at family-wise significance; the other four did not uniformly replicate it. All five were dominated by a recording-site axis (decodable at 0.980.98–1.001.00 vs. 0.5000.500 chance), whereas the DFA exponent they discard is site-robust (0.710.71). Significance. These results localize the failure to the pretraining objective (compounded by local positional encoding in CBraMod and BENDR) and motivate an LRTC-aware auxiliary loss. For cold-start deployment, the discarded exponent transfers directionally where the frozen embedding is at chance. Keywords: EEG foundation models, detrended fluctuation analysis, long-range temporal correlations, cross-population transfer, scale-free dynamics, recording-site leakage 1 Introduction Electroencephalography (EEG) foundation models (FMs) learn a general representation by self-supervised pretraining on large unlabeled corpora, and then transfer them to downstream clinical tasks through a frozen embedding with a light probe or a parameter-efficient adapter. A representation is only as useful as the signal properties it preserves. Among the most clinically meaningful properties of resting EEG is their scale-free temporal structure. The amplitude envelope of the alpha rhythm exhibits long-range temporal correlations (LRTC) with a power-law autocorrelation, summarized by a single dimensionless detrended-fluctuation-analysis (DFA) exponent (Linkenkaer-Hansen et al., 2001; Hardstone et al., 2012). LRTC is a candidate biomarker that is altered in Alzheimer’s disease, major depression, and across the psychiatric spectrum (Montez et al., 2009), and a proposed signature of near-critical cortical dynamics, though power-law-like envelope correlations are necessary but not sufficient evidence for criticality, so we treat criticality as motivating context and make our claims about the operational DFA measure itself, not about the presence of a critical state. Two facts make LRTC an unusually good probe for what a foundation model keeps. First, the DFA exponent is dimensionless; it is invariant to amplitude scaling, reference montage, and acquisition gain, which are the very factors that confound cross-site EEG. Second, LRTC is a temporal-ordering property: it exists only in the order of the samples; therefore, a representation that discards the temporal order cannot encode it. We ask the following direct question: Do EEG foundation models retain LRTC? The answer, across the five models and two out-of-distribution cohorts, is no, and the way in which they fail is specific and mechanistic. EEG FMs preserve the static spectral shape of the signal (the 1/f1/f aperiodic slope) when their front end ingests a spectral representation, but none of the five encode the temporal correlation structure the exponent measures. This is not a pooling artifact and is not a shadow of the aperiodic slope; therefore, we rule out both. The consequence is practical: when a Western-pretrained encoder is deployed zero-shot to a new population, the frozen embedding collapses to chance, while the single dimensionless DFA exponent, the property the FM discarded, shows a directional cross-population trend without any harmonization, a trend the data are consistent with, but that we are currently underpowered to confirm as a family-wise significant effect (§4.4). Contributions. 1. A DFA/LRTC probe of frozen EEG-FM embeddings across five architectures, establishing a spectral-vs-temporal dissociation, in which the aperiodic slope is (partly) decodable from spectral-input models, whereas the temporal-scaling exponent is decodable from none (§4.1). 2. An order-preserving control (learned probe on the pre-pool token sequence, ordered vs. shuffled) showing that, for REVE and CBraMod, the blindness is a property of the objective, not of mean-pooling, and that CBraMod’s apparent DFA recovery is an order-independent static proxy (BIOT’s gap is small but nominally reliable, so this claim is weaker for BIOT); for LaBraM and BENDR the control is uninformative because the probe underperforms on those embeddings (§4.2). 3. A dissociation from the aperiodic slope: envelope-DFA and the 1/f1/f slope are empirically orthogonal, so the FM’s failure to encode LRTC is not reducible to “FMs over-encode 1/f1/f” (§4.3). 4. A deployment-cost illustration (not a primary result): a dimensionless DFA exponent transfers directionally where the frozen FM collapses to chance, reported honestly as underpowered rather than family-wise significant (§4.4). 5. Novel site-dominance evidence and a proposed fix: on a leak-free two-cohort task, all five frozen FM embeddings decode the recording site near-perfectly (0.980.98–1.001.00 vs 0.5000.500 chance, ≈2×≈ 2× a matched null; §4.5); the DFA exponent they fail to encode is comparatively site-robust (site decodable at 0.710.71); site encoding is a universal property of the embeddings, not a REVE-specific artifact. We propose an LRTC-aware pretraining objective to improve DFA recovery and transfer (§5). 2 Related Work Spectral-bias audits of EEG foundation models. Kommineni et al. (2026) showed that reconstruction-based EEG-FM embeddings linearly decode the aperiodic exponent and offset but not oscillatory information, and attributed this to the reconstruction objective, in which 1/f1/f power dominates the loss. This is the closest prior work, and it differs from ours in scope. They measured the static, frequency-domain aperiodic component; they never touched the DFA, the alpha envelope, scale-free temporal dynamics, or cross-population transfer. Instantaneous spectral power is not a long-range temporal correlation. Bindra and Panwar (2026) reached a similar aperiodic-reliance conclusion for EEG/ECG models with the task-conditioned nuance that we adopted. The “Identity Trap” (Lin et al., 2026) shows FM embeddings are dominated by a subject-identity axis and notes REVE carries no measurable aperiodic dependence, consistent with our raw-waveform-blind finding, but frames 1/f1/f as a nuisance to erase rather than dynamics to preserve. Batch effects and leakage. Two 2026 studies (Tai, 2026; Tao et al., 2026) have shown that FM embeddings encode subject, site, and acquisition identity over neural semantics, and that leakage survives freezing. Tai (2026) showed that subject-linkable attributes transfer across independently pretrained EEG-FM encoders via a linear bridge and are largely unaffected by differential-privacy defenses, and Tao et al. (2026) showed that batch-related variability dominates diagnosis-related information in frozen fMRI foundation-model embeddings across sites. These findings support our cross-population-collapse result; our contribution is distinct in identifying the underlying mechanism (specifically LRTC) and quantifying the transfer contrast, rather than documenting that FMs generalize poorly across sites. A concurrent EEG-FM audit (Tang et al., 2026) applies a related LEACE-erasure diagnostic to a 6363-feature lexicon of hand-crafted neuro-features rather than to subject identity, and reports that erasing those features degrades label decoding. Crucially, that lexicon includes a DFA-style exponent (their feature C005, the clipped OLS slope of logF(s) F(s) vs. logs s over dyadic windows s∈16,…,512s∈\16,…,512\ samples, i.e. 0.080.08–2.562.56 s at 200200 Hz) and they find that their complexity family, C005 included, is broadly encoded and representation-causal in CBraMod and LaBraM. This is not in tension with our null result; it is the same dissociation observed from the other side. Their C005 is computed on the raw signal, where the DFA exponent is tied to the aperiodic slope by α≈(β+1)/2α\!≈\!(β+1)/2, so it is substantially a 1/f1/f proxy, exactly the static-spectral quantity our spectral-input models do encode (§4.1). We instead probe the narrowband alpha-envelope DFA over 0.50.5–3030 s, which is empirically orthogonal to that slope in our data (r=−0.06r\!=\!-0.06, §4.3); their positive result on a raw-signal (1/f-linked) exponent and our null on the envelope (1/f-orthogonal) exponent therefore corroborate rather than contradict. Separately, we measured whether the recording-site axis itself is decodable and whether harmonizing it away rescues cross-population disease transfer (§4.5), a different diagnostic question. Classical-competitive benchmarks. Classical features that can outperform the FMs cross-dataset have already been published (Kastrati et al., 2025). Their feature sets include the fractal dimension but not envelope-DFA/LRTC, use no single dimensionless measure, and offer no mechanism. Therefore, our contribution is not that classical features outperform FMs; it is the LRTC probe, the dimensionless zero-shot transfer, and the temporal-scaling mechanism. Why do the sequence models lose their temporal order? Attention is permutation-invariant absent positional signal (Zeng et al., 2023) and contrastive SSL destroys autocorrelation unless explicitly regularized (Jiao et al., 2025). Consistent with an objective-level cause, Lehn-Schiøler et al. (2026) recover signal amplitude but not phase from frozen EEG-FM embeddings (phase cosine ≤0.22≤ 0.22), attributing the lost temporal morphology to the time-translation invariance of the pretraining objective. We further hypothesize, without direct evidence here, that flat patch tokenization discards multi-scale temporal structures and that deep networks may underestimate the scaling exponent of fractional processes; we flag these as motivating conjectures rather than established results. Together, they motivated our mechanistic hypothesis and proposed fix. 3 Methods Cohorts. We used five resting-state clinical EEG cohorts spanning four continents (Table 1). The encoding probe (§4.1) used two out-of-distribution dementia cohorts, CAUEEG (Korean dementia/normal) (Kim et al., 2023) and BrainLat (Latin-American AD/bvFTD/HC, 128-channel) (Prado et al., 2023), the latter a replication cohort whose shorter recordings support only the short DFA range (§3, Targets). The cross-population transfer task (§4.4) additionally used ds004504 (Greek/Western AD/FTD/HC) (Miltiadous et al., 2023) as a Western training source, so it spans three cohorts while the encoding probe spans two. ASZED-153 (Nigerian schizophrenia) (Mosaku and others, 2025) is included in the released cohort set for future scale-boundary and population-invariance extensions (its median recording length, ∼ 40 s, is too short to support the same 0.50.5–3030 s DFA scale range used for the encoding and transfer analyses on the other four cohorts). ds004504 and CAUEEG were confirmed absent from REVE’s pretraining corpus and are the two cohorts used for the recording-site probe (§4.5); verification method: we cross-checked the 92-dataset pretraining enumeration in the REVE paper’s Appendix B (El Ouahidi et al., 2025) and neither cohort appears under its own name or any listed alias, whereas TDBRAIN does. TDBRAIN (Dutch psychiatric) (van Dijk et al., 2022) is listed in Table 1 for completeness but is excluded from the recording-site probe: it is named in REVE’s own pretraining corpus (El Ouahidi et al., 2025, Appendix B), so a REVE-embedding result on TDBRAIN cannot distinguish genuine site-encoding from memorization of specific recordings (§7). Table 1: Cohorts. All non-Western cohorts are out-of-distribution for the FMs. TDBRAIN is listed for completeness but excluded from all REVE-embedding analyses (§3, Cohorts). Cohort Population Task N Recording ds004504 Greek (Western) AD/FTD/HC 88 ≥ 240 s CAUEEG Korean Dementia/MCI/Normal 1187 ≥ 240 s BrainLat Latin-American AD/bvFTD/HC 79 long (HEP) TDBRAIN Dutch (psychiatric) excluded (REVE pretrain leak) 300 ∼ 120 s ASZED-153 Nigerian schizophrenia 122 ∼ 40 s Foundation models and a front-end taxonomy. We probed five EEG FMs, grouped by the representation their front end ingests, which turns out to be the axis that predicts what they preserve: raw-waveform models: REVE (69.4M, masked reconstruction) (El Ouahidi et al., 2025), LaBraM (patch) (Jiang et al., 2024), and BENDR (wav2vec-style contrastive) (Kostas et al., 2021); and spectral-input models: CBraMod (Wang et al., 2025) and BIOT (FFT/STFT front end) (Yang et al., 2023). All were frozen, and embeddings were subject-mean pooled, unless the pre-pool sequence was explicitly used. Beyond the raw-vs-spectral split, the five models differ along several further design axes that we did not factorially separate: pretraining target (LaBraM predicts a discrete vector-quantized code per patch, reconstructed by a separate decoder against Fourier amplitude and phase; CBraMod and REVE reconstruct the raw patch directly); reconstruction loss (CBraMod uses MSE; REVE uses L1, motivated in its release as more robust to the high-amplitude outliers common in raw EEG; BENDR uses a wav2vec2.0-style contrastive objective rather than reconstruction); positional encoding, which sets whether an architecture can represent long-range temporal order at all (REVE’s 4D Fourier code over electrode position and patch index, LaBraM’s learned per-patch temporal embedding, and BIOT’s sinusoidal segment code each assign a distinct code to every time index and can in principle represent order across their input window, whereas CBraMod’s ±3± 3 s conditional convolution and BENDR’s wav2vec-style grouped convolution provide only a local signal and cannot); and pretraining corpus scale (REVE: ∼ 60,000 h across 92 datasets and 25,000 subjects; CBraMod: a ∼ 9,000 h subset of the Temple University EEG Corpus; LaBraM: ∼ 2,500 h of mixed EEG). Because these axes co-vary across models, differences among the three raw-waveform models are reported descriptively (§4.1), rather than being attributed to any single design choice. REVE as the primary probe; the others test generality. REVE is the largest and most recent model in this set: its pretraining corpus is several-fold larger than any other’s (see above), so it is the strongest a priori candidate to have captured scale-free temporal structure, and a null result on REVE is correspondingly the hardest to dismiss. We therefore probe REVE most deeply, including the finer token-level pre-pool control (§4.2) and the original recording-site and transfer-rescue analyses (§4.5), then ask whether each finding generalizes across the four architecturally diverse FMs that span the raw-waveform (LaBraM, BENDR) and spectral-input (CBraMod, BIOT) front ends. We flag throughout where a result is REVE-specific rather than universal: the encoding blindness (§4.1) and near-perfect recording-site decoding (§4.5) hold for all five, whereas the cross-population transfer collapse is demonstrated for REVE, reproduces for LaBraM and BENDR, is sign-inverted for BIOT, and does not hold for CBraMod (§4.4, §4.5). Targets. The LRTC target is the DFA exponent of the alpha-band (8–13 Hz) amplitude envelope (Hardstone et al., 2012), computed per channel and then averaged across the 19 channels to a single scalar for the encoding-probe regression target used in §4.1–4.3; the transfer task (§4.4, Table 5) instead uses the un-averaged 19-channel DFA vector directly as a classifier feature, without collapsing it to a scalar. The envelope is the Hilbert-transform magnitude of a zero-phase order-4 Butterworth 8–13 Hz bandpass. The DFA integrates the mean-subtracted envelope via cumulative sum, splits the result into non-overlapping windows at each scale, removes a first-order (linear) trend per window by least-squares, and computes the root-mean-square residual; the exponent is the slope of a log-log linear fit of the RMS residual against scale. The full scale range is 16 log-spaced scales from 0.50.5–3030 s, used for all cohorts except BrainLat, where recordings are too short and a short range of 8 log-spaced scales from 0.50.5–22 s is used instead (labeled accordingly wherever it appears, e.g. §4.1); the two ranges are not interchangeable and cross-cohort DFA magnitudes should not be compared directly (we test this directly in §4.1). We flag one caveat on the short range: the bandpass has a nominal 55 Hz bandwidth, so the filtered envelope carries filter-induced autocorrelation on the order of its reciprocal (≈0.2≈ 0.2 s); a fit starting at 0.50.5 s is only ≈2.5×≈ 2.5× that, closer to the filter’s own correlation time than the full range’s starting scale is. We tested this directly with three surrogate controls on the BrainLat envelopes (N=79N=79, 0.50.5–22 s range). The log-log fit is highly linear (R2=0.99±0.004R^2=0.99± 0.004), so it is a well-defined slope over the 0.60.6-decade range; the exponent (0.90±0.080.90± 0.08) exceeds a matched filtered-white-noise null (0.76±0.0250.76± 0.025) by 5.65.6 null-SD (i.e. (0.898−0.757)/0.025(0.898-0.757)/0.025); and a temporally-shuffled envelope collapses to 0.50±0.020.50± 0.02. The filter inflates the short-range baseline (the noise null sits at 0.760.76, not 0.500.50); therefore, short-range exponents should be read against that floor, not against 0.50.5. (That the null lands at 0.760.76, the low end of the full-range cohort exponents in §3, is a scale-range coincidence, not evidence of comparability; this null was computed only over the 0.50.5–22 s range, and we never compare exponent magnitudes across scale ranges.) The classical feature’s between-subject recovery on BrainLat (§4.1) is nonetheless not a filter artifact, for variance reasons: the observed exponents vary ≈3.4×≈ 3.4× more widely than the filtered-noise null (SD 0.0840.084 vs. 0.0250.025), so the filter floor accounts for at most ∼ 9% of the exponent variance ((0.025/0.084)2(0.025/0.084)^2). We report the short-range measure with this caveat, rather than as an unqualified scaling exponent. Two further caveats apply to the full range: at the top scale (3030 s) a 240240 s recording yields only 8 non-overlapping windows, so the highest-leverage point on the log-log fit is also its noisiest; and our windows are non-overlapping rather than the sliding, overlapping-window convention some DFA implementations use, which trades a simpler, unbiased-by-construction estimator for a somewhat noisier one at a given recording length. The comparison target is the aperiodic 1/f1/f slope from the spectral parameterization. DFA quality control across the full-range cohorts: exponent ≈0.76≈ 0.76–0.800.80, 9696–99%99\% of subjects in [0.5,1.0][0.5,1.0], log-log fit R2≈0.98R^2≈ 0.98–0.990.99 (BrainLat’s short-range exponent is higher, ≈0.90≈ 0.90, and is not comparable across scale ranges; see the surrogate controls above): the measure is scale-free and population-invariant, a precondition for using it as a transfer target. This narrow target range (most subjects within a 0.50.5-wide band) does not by itself cap R2R^2, which is scale-invariant; what bounds achievable R2R^2 is measurement reliability, which we estimate directly in §4.1. Encoding probe. For each (FM, target) pair, we fitted a probe from the frozen embedding to the scalar target and reported the best out-of-sample 5-fold cross-validated R2R^2 across a ridge, gradient-boosting, and random-forest probe. A classical DFA/1/f1/f feature vector serves as the positive control (it must recover the target, and does). Using the best of the three probes is deliberately generous to the FM, ensuring that the probe choice cannot account for the observed failure to recover the LRTC. Pre-pool order control. To separate an objective-level failure from a pooling artifact, we read the pre-pool token/epoch sequence (a genuine time axis for every architecture) with a learned 1D-CNN and compared the true temporal order against a per-subject shuffle of the same sequence (identical capacity). LRTC is a temporal-ordering property; therefore, any representation of it must make ordered >> shuffled. We reported the mean ± SD over the five seeds. This control also speaks to the concern that poor frozen-embedding performance is largely a mean-pooling artifact, recoverable by keeping token-level rather than pooled embeddings (Širca et al., 2026): our probe reads the pre-pool token sequence directly, so an LRTC representation that pooling had merely hidden would surface here, and its absence (§4.2) locates the failure in the objective, not the readout. Aperiodic dissociation. To test whether any FM DFA recovery is a shadow of the 1/f1/f slope, we residualize the envelope-DFA exponent on the aperiodic slope β and re-probe. (The α≈(β+1)/2α\!≈\!(β+1)/2 identity applies to the raw-signal DFA, not the narrowband envelope DFA used here, so orthogonality is expected and is what we test.) Transfer protocol. Cohorts are harmonized to a common 19-channel 10–20 montage, 200 Hz, common-average reference, 1–45 Hz, first 240 s. Transfer is zero-shot train-one/test-other with a StandardScaler(train)+RBF-SVM, scored by AUROC with a repeated-resample 95% CI, a 5000-permutation label null, and Holm correction across the DFA/1/f1/f family. 4 Results 4.1 The spectral-vs-temporal dissociation Table 2 and Figure 1 report how well each frozen embedding predicts the DFA exponent and the 1/f1/f slope on two out-of-distribution dementia cohorts. On CAUEEG, we use the full available sample with cached embeddings and raw-EEG-derived targets (N=770N=770, up from an earlier N=200N=200 subsample; BENDR was not part of this rerun and remains at N=200N=200, footnoted in the table) rather than an arbitrary cap, because for a not-decodable claim, the standard rebuttal is an underpowered probe. N=770N=770 is not a further exclusion from CAUEEG’s full 1187 subjects: it is exactly the Normal (459459) plus Dementia (311311) subset that defines this binary task, while the remaining 417417 carry the intermediate MCI label and are out of scope for a Normal-vs-Dementia target by design, not dropped for data quality. These two patterns hold true across both cohorts. First, along the aperiodic axis the front end decides everything: spectral-input models (CBraMod, BIOT) recover the 1/f1/f slope strongly (R2=0.59R^2=0.59–0.630.63 on CAUEEG at N=770N=770; 0.640.64–0.730.73 on BrainLat), while all three raw-waveform models (REVE, LaBraM, BENDR) recover it near-zero (R2≤0.12R^2≤ 0.12 across both cohorts; the largest single raw-waveform cell is LaBraM on BrainLat at 0.120.12, roughly five-fold below the spectral models). Second, along the temporal axis, every model fails: no FM recovers the DFA exponent on both cohorts, while the classical feature recovers it (R2=0.32R^2=0.32 on CAUEEG at N=770N=770, using the full 0.50.5–3030 s scale range; 0.380.38 on BrainLat using a short 0.50.5–22 s scale range, since BrainLat’s shorter recordings cannot support the full range). Therefore, these two classical numbers are not computed on identical scale ranges and are not directly comparable in magnitude; both nonetheless clear the FMs’ R2≈0R^2≈ 0 for the same signals. On CAUEEG specifically, CBraMod and BIOT reach 0.590.59 and 0.630.63 on 1/f1/f against a classical ceiling of 0.590.59, essentially matching it, while reaching only 0.180.18 and 0.250.25 on DFA against a classical 0.320.32: the same embeddings, on the same subjects, close the gap to classical on the aperiodic axis and fall well short of it on the temporal axis. The weak DFA recovery that CBraMod/BIOT show on CAUEEG (0.180.18–0.250.25 at N=770N=770) does not replicate on BrainLat (≤0.06≤ 0.06), whereas their 1/f1/f recovery replicates cleanly. Because BrainLat’s DFA target necessarily uses the short 0.50.5–22 s range while CAUEEG’s headline number above uses the full 0.50.5–3030 s range (§3, Targets), the non-replication could, in principle, reflect a change in target definition rather than a change in cohort. We rule this out directly: re-probing CBraMod/BIOT on CAUEEG against DFAshort_short (the same 0.50.5–22 s range used on BrainLat) yields 0.1750.175/0.2420.242, within the noise of the full-range values (0.1780.178/0.2510.251) reported above. Their CAUEEG recovery is therefore scale-range-invariant, so the BrainLat collapse is a genuine cohort effect, not an artifact of comparing two different DFA definitions, and the CAUEEG DFA signal is a cohort-specific artifact rather than a representation of LRTC. Quadrupling the CAUEEG sample left both patterns essentially unchanged (DFA recovery flat for every model; 1/f1/f recovery for the spectral models and classical strengthened rather than weakened), which is the signature of a real, sample-size-independent dissociation rather than an artifact of the earlier N=200N=200 subsample. To establish what R2R^2 is achievable at all, rather than assume a ceiling of 1.01.0, we estimated the split-half reliability of DFAfull_full on CAUEEG: computed independently on the first and second half of each recording (0–120120 s vs. 120120–240240 s, N=770N=770) and correlated across subjects, giving r=0.80r=0.80 (5-fold cross-validated R2=0.64R^2=0.64 predicting one half from the other; Spearman-Brown-corrected full-length reliability r=0.89r=0.89). The classical probe’s R2=0.32R^2=0.32 is therefore roughly half of the empirically attainable ceiling, not close to a ceiling of 1.01.0 as a naive reading would suggest; the FMs’ R2≈0R^2≈ 0 means that they recover essentially none of a target with substantial recoverable structure, which is the stronger and more precise statement than “the target is hard.” Table 2: Encoding probe (best out-of-sample 5-fold CV R2R^2 over ridge/gboost/RF). 1/f1/f replicates by front end; no FM recovers DFA on both cohorts (BIOT and CBraMod recover it on CAUEEG only, and it does not replicate on BrainLat). †BENDR’s CAUEEG columns use the original N=200N=200 subsample; all other CAUEEG columns use the full N=770N=770 rerun (§4.1). CAUEEG BrainLat (N=79N=79) Model front end DFA 1/f1/f DFA 1/f1/f REVE raw-waveform −0.01-0.01 0.000.00 −0.03-0.03 −0.06-0.06 LaBraM raw-patch 0.000.00 0.010.01 −0.03-0.03 +0.12+0.12 BENDR† raw-wav2vec 0.010.01 0.100.10 −0.02-0.02 −0.03-0.03 CBraMod spectral-input 0.180.18 0.590.59 +0.06+0.06 +0.64+0.64 BIOT spectral-input 0.250.25 0.630.63 +0.03+0.03 +0.73+0.73 Classical hand-computed 0.320.32 0.590.59 +0.38+0.38 +0.86+0.86 Figure 1: The dissociation, both cohorts. Bars show the best cross-validated R2R^2 for each frozen embedding on CAUEEG (N=770N=770) and BrainLat (N=79N=79). Along the aperiodic axis the front end decides everything (spectral-input models recover 1/f1/f; raw-waveform models do not). Along the temporal axis every FM fails: the DFA recovery spectral models show on CAUEEG collapses on BrainLat, while the classical feature recovers it on both. Red = raw-waveform, blue = spectral-input, green = classical. BENDR’s CAUEEG bars use the original N=200N=200 subsample, as it was not part of the N=770N=770 rerun; its near-zero recovery is unaffected by sample size. Because the DFA target is narrow (§3 sec:methods, Targets), R2R^2 alone can be difficult to interpret; Figure 2 shows cross-validated predictions against ground truth directly, at the full available CAUEEG sample (N=770N=770, up from the N=200N=200 used for Table 2). The classical positive control tracks the identity line (R2=0.31R^2=0.31); REVE’s predictions are flat at ≈0.80≈ 0.80 regardless of the true value (R2=−0.06R^2=-0.06), the visual signature of a model predicting the sample mean rather than recovering subject-level structure. This replicates the R2≈0R^2≈ 0 result at nearly 4×4× the sample size used elsewhere in this section. Figure 2: Predicted vs. true alpha-envelope DFAfull exponent on CAUEEG (N=770N=770, 5-fold CV, ridge probe). Left: the classical feature tracks the identity line. Right: REVE’s predictions collapse to a narrow band near the sample mean irrespective of the true value, visually confirming that R2≈0R^2≈ 0 reflects a genuine failure to recover subject-level LRTC structure rather than an artifact of the target’s narrow variance. 4.2 The blindness is in the objective, not the pooling A natural rebuttal is that mean-pooling discards an LRTC representation that the model computes. We tested this directly by reading the pre-pool sequence with a learned order-aware probe and comparing the true order to a per-subject shuffle over 5 seeds (Table 3), reporting the paired ordered−-shuffled gap per seed with its standard error and a paired t-test rather than treating “gap ≈0≈ 0” as self-evident. This control is informative only where the probe is above the noise floor. For REVE, using a finer token-level probe, the ordered condition recovers no more than the shuffle and no more than the pooled embedding (both ≈0≈ 0), and there is no pre-pool temporal LRTC to discard, so the failure is at the level of the objective (there is nothing to pool away), not the pooling. For LaBraM and BENDR, both ordered and shuffled R2R^2 are negative (−0.12-0.12 and −0.15-0.15, respectively), meaning the shared 1D-CNN probe underperforms a mean predictor on these embeddings; a negative-vs-negative comparison cannot distinguish “no LRTC to discard” from “the probe was too weak to detect it.” LaBraM’s paired gap is in fact significantly negative (−0.090±0.019-0.090± 0.019 SEM, paired t-test p=0.009p=0.009), which no genuine LRTC signal would produce and which is itself evidence for probe instability rather than a resolved absence of LRTC; BENDR’s gap shows no directional bias (−0.001±0.025-0.001± 0.025, p=0.98p=0.98) despite both conditions sitting well below the noise floor. We treat the objective-vs-pooling question as unresolved for both models using this control (their encoding-probe failure in Table 2 stands on its own evidence). For the spectral-input models, the probe does recover the DFA-correlated information (0.060.06–0.180.18). CBraMod’s gap is not distinguishable from zero (+0.008±0.038+0.008± 0.038, p=0.84p=0.84): shuffling the epoch order does not reduce recovery, supporting an order-independent static proxy. BIOT’s gap is small but nominally reliable (+0.029±0.010+0.029± 0.010, p=0.049p=0.049, ≈9%≈\!9\% of the classical baseline): we cannot claim that BIOT’s apparent DFA recovery is fully order-independent, only that any order-sensitive component is far too small to account for its recovery magnitude. This is exactly why the CAUEEG spectral-model signal is cohort-fragile and did not replicate on BrainLat (§4.1). No FM recovers the DFA exponent (Table 2), and for REVE and CBraMod specifically, this control locates the failure in the objective rather than in pooling (Figure 3). Read together with the positional taxonomy (§3), this places the binding constraint on the objective rather than the architecture: REVE, LaBraM, and BIOT can represent long-range order, yet none reproducibly recovers the DFA exponent (Table 2), so the capacity to represent order does not translate into encoding it. CBraMod and BENDR are additionally constrained, their local positional encoding precluding long-range order in principle. Table 3: Pre-pool order control on CAUEEG (DFA, best CV R2R^2; mean± over 5 seeds; REVE from its finer token-level probe). Gap == paired ordered−-shuffled difference, mean± with a paired t-test p. This control is informative only where ordered/shuffled R2R^2 clear the noise floor: REVE and CBraMod cleanly support order-independence; BIOT’s gap is small but nominally significant; LaBraM’s gap is significantly negative (probe failure) and BENDR’s shows no bias despite both conditions being deeply negative. Each model’s ordered/shuffled/pooled R2R^2 is read against the classical DFA positive control (R2=0.32R^2=0.32; dashed line in Fig. 3). Model front end ordered shuffled gap (SEM) p pooled REVE raw-waveform −0.03-0.03 −0.02-0.02 ≈0≈ 0 – ≈0≈ 0 LaBraM raw-patch −0.12±0.11-0.12±0.11 −0.03±0.12-0.03±0.12 −0.090±0.019-0.090±0.019 0.0090.009 −0.01-0.01 BENDR raw-wav2vec −0.15±0.05-0.15±0.05 −0.15±0.05-0.15±0.05 −0.001±0.025-0.001±0.025 0.980.98 −0.02-0.02 CBraMod spectral-input +0.18±0.07+0.18±0.07 +0.17±0.10+0.17±0.10 +0.008±0.038+0.008±0.038 0.840.84 +0.09+0.09 BIOT spectral-input +0.06±0.07+0.06±0.07 +0.03±0.07+0.03±0.07 +0.029±0.010+0.029±0.010 0.0490.049 +0.08+0.08 Figure 3: Pre-pool order control (CAUEEG, DFA, 5 seeds). REVE and CBraMod show ordered≈ with no reliable paired gap (p=0.84p=0.84 for CBraMod), supporting order-independence; BIOT’s gap is small but nominally reliable (p=0.049p=0.049); LaBraM and BENDR are both negative, indicating probe underperformance (LaBraM’s gap is significantly negative, p=0.009p=0.009) rather than a resolved absence of pre-pool LRTC. The classical DFA feature (dashed) recovers the exponent. 4.3 LRTC is not a shadow of the aperiodic slope The strongest alternative explanation is that DFA and the 1/f1/f slope are the same axis, so “FMs encode 1/f1/f but not DFA” would be incoherent. They do not have the same axis on CAUEEG at the full N=770N=770 sample used in §4.1, and the envelope-DFA exponent and the aperiodic slope are empirically orthogonal in this cohort (r=−0.06r=-0.06, p=0.07p=0.07; β explains 0.4%0.4\% of the variance in α). Residualizing the DFA on β and re-probing leaves every encoder’s recovery essentially unchanged (classical 0.32→0.310.32→ 0.31; CBraMod 0.18→0.180.18→ 0.18; BIOT 0.25→0.240.25→ 0.24; LaBraM and REVE stay ≈0≈ 0). Therefore, the DFA variance that spectral FMs recover in this cohort is slope-independent (variance in the DFA target proper, not a shadow of the aperiodic component) rather than 1/f1/f re-labeled. This is all the residualization licenses: it does not make the recovery a representation of temporal LRTC, which the order control (§4.2) rules out, nor cohort-general, which §4.1 rules out. Similarly, the classical baseline carries a strong slope-independent DFA variance (0.320.32). The spectral-vs-temporal dissociation is therefore robust to this alternative explanation on CAUEEG, and we have not rerun the residualization on BrainLat. Therefore, we do not claim that the orthogonality itself generalizes across cohorts, only that it holds in the cohort where the dissociation is cleanest (§4.1). 4.4 Cross-population transfer: the dimensionless exponent carries what the FM drops This arm illustrates the cost of the encoding failure and is secondary to it: the frozen-embedding collapse reported below is a clean result (the classifier falls to chance), whereas the favorable contrast with the DFA feature is a directional trend on a small Western training set (n=53n=53), not a family-wise-significant effect, and the mechanism (§4.1) does not depend on it. With that scope fixed, this dissociation incurs a high deployment cost. On zero-shot cross-population transfer across the three long-recording cohorts (Table 4), the frozen REVE embedding, tested for transfer only on the W↔ pair (Table 5), did not beat chance there (raw W→ 0.450.45; K→ 0.500.50, a degenerate constant-classifier cell). This frozen-embedding transfer was not run for the BrainLat or leave-one-cohort-out directions, nor for the other four FMs (§7). The dimensionless DFA exponent transfers directionally (Figure 4): the sign of the class effect (dementia lowers DFA) is conserved across all three populations, and point AUROCs sit above chance in every one of the six pairwise and three leave-one-cohort-out (LOCO) directions we tested (0.5750.575–0.7400.740). We tested this pattern against the full DFA/1/f1/f family rather than reporting the single most favorable cell: two feature sets (DFA-only, DFA+1/f+1/f) × 9 directions (6 pairwise, 3 LOCO) == 18 cells, Holm-corrected. No DFA-only direction clears Holm-adjusted α=0.05α=0.05, including K→ , which has the smallest raw permutation p in the family (p=0.011p=0.011) but does not survive the correction (pholm=0.187p_holm=0.187). The one cell that does survive in the full 18-cell family is DFA+1/f+1/f in the K→ direction (0.8070.807, pholm=0.018p_holm=0.018), which is not the population-invariant axis the paper argues for, since the same 1/f1/f component inverts sign on transfer into BrainLat (0.0950.095–0.130.13, below chance, on both the pairwise and LOCO directions into BrainLat), so its significance is plausibly carried by the non-invariant aperiodic component rather than by DFA. Therefore, we report the DFA transfer honestly as directionally consistent across all tested cells, but statistically unconfirmed at family-wise-corrected strength: the small Western training set (n=53n=53) limits power, not effect size. This also isolates the dimensionless temporal-scaling exponent, not the aperiodic amplitude, as the candidate population-invariant axis; thus, a claim the present transfer analysis is underpowered to confirm, but does not contradict. Table 4: Zero-shot cross-population transfer, DFA feature. This is the model-free DFA feature vector (no FM embedding); per-model frozen-embedding transfer is in Table 5. All nine DFA-only directions (six pairwise, above the rule; three leave-one-cohort-out, below), so “directionally consistent across all nine” is verifiable here. AUROC; raw permutation p; Holm-adjusted pholmp_holm across the full 18-cell DFA/1/f1/f family; W=Greek, K=Korean, B=BrainLat. Every direction sits above chance (0.5750.575–0.7400.740); no cell clears pholm<0.05p_holm<0.05. For comparison, the frozen REVE embedding was tested for transfer only on the W↔ pair (Table 5), where it does not beat chance (raw W→ 0.450.45; K→ 0.500.50, a degenerate constant-classifier cell); it was not run on the BrainLat or leave-one-cohort-out directions. Direction DFA AUROC p pholmp_holm K→ 0.7400.740 0.0110.011 0.1870.187 B→ 0.6910.691 0.0850.085 0.7940.794 W→ 0.6800.680 0.0970.097 0.7940.794 B→ 0.6300.630 0.0620.062 0.7930.793 K→ 0.5770.577 0.0690.069 0.7930.793 W→ 0.5750.575 0.1360.136 0.9040.904 K++B→ 0.7060.706 0.0300.030 0.4860.486 W++B→ 0.6720.672 0.0480.048 0.7230.723 W++K→ 0.5810.581 0.0610.061 0.7930.793 Figure 4: Cross-population DFA transfer (point AUROC, resample 95% CI, raw permutation p, Holm-adjusted pholmp_holm). All point AUROCs sit above chance, though some resample CIs dip toward chance; the sign of the class effect is conserved across all three populations. K→ has the smallest raw p (0.0110.011) but does not survive Holm correction across the 18-cell DFA/1/f1/f family (pholm=0.187p_holm=0.187): the effect is directional but power-limited on three cohorts, not family-wise confirmed. Scope note. The cross-population reversal itself (FM vs. classical on non-Western clinical tasks) is the subject of a companion clinical benchmarking manuscript (Systematic Benchmarking of EEG Foundation Models on Clinical Tasks: Negative Controls Invalidate Most Apparent Gains), whose random-initialisation control — pretrained REVE underperforming a randomly-initialised encoder on a non-Western cohort — is the predicted empirical signature of the encoding failure characterised here; that paper’s within-cohort frozen scores (e.g. 0.5890.589 on CAUEEG) are not directly comparable to the zero-shot cross-population transfer values (≤0.44≤ 0.44) reported in this work. Here we use transfer only to establish the cost of the encoding failure, and keep the claim scoped to the DFA-vs-frozen-FM contrast. 4.5 The frozen embedding is dominated by recording site Site-dominance is a symptom of the same objective failure, not the mechanism of the collapse. As the harmonization test below shows, removing the site axis does not recover a disease axis that was never encoded. With that framing fixed, the frozen embedding is nonetheless strikingly dominated by a recording-site axis. A subject-level probe recovered the two-way recording site (ds004504/Greek, CAUEEG/Korean, both confirmed to be absent from REVE’s pretraining corpus; TDBRAIN is excluded here for the reason given in §3) from the frozen REVE embedding at a balanced accuracy of 0.9940.994 (chance 0.5000.500; 5-fold, PCA-50 + logistic), near-perfect site decoding consistent with the identity-leakage observation of Lin et al. (2026), who evaluated the same ds004504 cohort (the Miltiadous AD/FTD dataset). To rule out probe capacity alone as the explanation, a matched random-Gaussian null (20 draws of noise shaped identically to the pre-PCA embedding, N=176N=176, D=38,912D=38,912, put through the same StandardScaler→ -50→ pipeline) recovers only 0.523±0.0490.523± 0.049, close to chance: the real embedding’s decodability is a genuine 1.9×1.9× excess over this null, not an artifact of fitting a 50-dimensional probe on 176 samples. Is the dimensionless DFA exponent correspondingly harder to decode site from, or does it merely transfer disease signal despite site differences? We tested this directly with the identical two-way site probe (ds004504 vs. CAUEEG, logistic, matched null) applied to the 1919-d DFA vector used as the transfer feature (Table 5). The DFA exponent decodes site at balanced accuracy 0.710.71 (1.4×1.4× its matched null), far below the frozen embeddings’ 0.980.98–1.001.00: the exponent is markedly more site-robust, though not perfectly site-invariant. The comparison also guards against an obvious rebuttal: that any feature separates the two cohorts recorded on different continents and hardware. It does not distinguish the useful features: the 173173-d classical feature decodes site at 1.001.00, at ceiling like the FMs. Therefore, the discriminating variable is not whether site is decodable, but whether the transferable disease signal survives alongside it: present-but-masked for classical (rescued by harmonization below), comparatively site-light for DFA, and absent for the frozen FM. Site dominance is not a REVE idiosyncrasy but a property of all five frozen embeddings. Running the identical probe on the four other FMs, on the same leak-free two-cohort binary task used for REVE (ds004504 vs. CAUEEG, chance 0.5000.500), every model decodes the recording site nearly perfectly: LaBraM 0.9890.989, CBraMod 1.0001.000, BIOT 0.9770.977, BENDR 1.0001.000 (all at ≈2×≈ 2× their matched random-Gaussian null, 0.500.50–0.510.51), alongside REVE’s 0.9940.994. Unlike the graded 1/f1/f-encoding split of §4.1, site decodability does not separate raw-waveform from spectral-input models; it is uniformly at ceiling across the front-end taxonomy. (A three-way variant that adds TDBRAIN as a third site shows CBraMod/BIOT decoding more weakly, but that comparison mixes a harder three-class task with TDBRAIN’s pretraining-leak status for REVE and is not the apples-to-apples baseline; the leak-free binary task above is.) Site encoding is, therefore, a universal property of frozen EEG-FM embeddings, not a REVE-specific artifact, and not an architecture-graded one. We have directly answered the operative questions. A transductive, label-free ComBat on the reduced embedding (site as batch, no diagnosis covariate) removes the linear site axis; site decodability falls from 0.9940.994 to below chance (0.440.44); therefore, in-sample, decodable site is a first/second-moment batch effect. However, removing a site axis is not the same as recovering a disease axis, so we ran the deciding test: cross-population dementia-vs-control transfer of the frozen embedding under leakage-safe harmonization (label-free per-site standardization and ComBat; disease labels never used), against the two controls (Table 5). Harmonization does not rescue the frozen REVE embedding: no regime clears chance at the bootstrap CI lower bound, and in the well-powered direction (testing on the large Korean cohort), it stays at or below chance (0.450.45–0.460.46); the lone nominal uptick (K→W≈0.61K→W≈ 0.61) is the underpowered direction (n=53n=53 test) with a CI that spans chance. The identical harmonization does rescue the classical feature (whose raw zero-shot cell is a degenerate constant-output classifier, Table 5‡; harmonized to 0.770.77, CI above chance), a disease signal that was present but site-masked, and the dimensionless DFA exponent transfers without any harmonization (0.680.68–0.740.74). The three-way contrast is the result: leakage-safe harmonization recovers cross-population disease transfer when the signal is present but masked (classical), is unnecessary when the signal is scale-free (DFA), and cannot rescue the frozen REVE embedding, because the transferable disease signal is not encoded there, only the recording-site identity is. Removing the site axis does not conjure a disease axis that was never represented. The same test on the other four FMs (fresh n=65n=65/770 W/K cohort; §4.5) does not uniformly repeat the REVE pattern. LaBraM and BENDR match it: no regime clears chance at the CI lower bound in either direction (Table 5). BIOT’s raw W→KW→K cell sits significantly below chance (0.360.36, CI excludes 0.50.5, on the lower side) and stays there after harmonization (0.410.41). An AUROC significantly below chance is itself informative and consistent with our thesis: the frozen embedding does carry a disease-relevant structure, but a structure whose sign inverts across populations, exactly the non-invariance expected of a site-entangled representation, and a failure of transfer rather than a rescue. CBraMod is the one genuine exception: its raw W→KW→K cell already clears chance before any harmonization (0.590.59, CI lower bound 0.550.55) and stays there after ComBat (0.580.58), essentially unchanged by harmonization either way. We report this descriptively and do not have a mechanism for it. The natural explanation, that CBraMod’s spectral front end simply encodes a more transferable signal, is refuted by BIOT: BIOT encodes the 1/f1/f slope at least as strongly (Table 2: 0.630.63 vs. 0.590.59 on CAUEEG, 0.730.73 vs. 0.640.64 on BrainLat) yet transfers below chance, so stronger static-spectral encoding coincides with opposite transfer behavior and cannot be the cause. Instead, we offer a falsifiable hypothesis: CBraMod’s criss-cross attention explicitly mixes the channel and time axes, so its order-independent “static proxy” may carry cross-channel spatial covariance that is more population-stable than pure 1/f1/f amplitude, and it is this spatial component, not a temporal one, that transfers. This predicts two concrete tests: the above-chance CBraMod transfer should survive shuffling the pre-pool token order before re-pooling and re-running transfer (the signal is not temporal) but should be abolished by ablating the channel-mixing (spatial-covariance) pathway; we leave both to future work. Accordingly, we bound the practical punchline: the dimensionless exponent is a better cold-start representation than four of the five frozen embeddings, with CBraMod being an unexplained exception whose above-chance transfer we can neither attribute to LRTC nor yet explain. Thus, this arm’s irrecoverability conclusion is demonstrated for REVE, reproduces for LaBraM and BENDR, and holds in a sign-inverted form for BIOT. The frozen embedding also cannot be rescued by combining it with the model-free DFA feature: fusing the frozen CBraMod embedding with the alpha-envelope DFA exponent does not beat the DFA feature alone—decision-level fusion matches it, whereas naive feature concatenation degrades it (the K→ direction collapses to the degenerate chance value)—so the embedding adds no cross-population disease signal beyond the model-free marker. Table 5: Leakage-safe transfer-rescue (target AUROC; harmonized == label-free per-site standardization / ComBat, disease labels never used). Bold == 95% bootstrap CI lower bound >0.5>0.5. Harmonization rescues the site-masked classical feature and is unnecessary for the dimensionless DFA exponent, but does not lift the frozen REVE, LaBraM, or BIOT embedding above chance; CBraMod is the exception (§4.5). †CI spans chance (n=53n=53 test cohort). ‡Degenerate cell: under this zero-shot covariate shift the RBF-SVM emits a constant output for every test subject (class-probability and decision-function margin identical across subjects), so AUROC is undefined and reported as chance (0.50.5) with zero-width CI, the expected collapse of a high-dimensional embedding classifier under severe shift, resolved in the harmonized column where a rescue exists. REVE’s cohort is n=53n=53/770 (W/K); the other four FMs use a freshly-built n=65n=65/770 cohort (§4.5). raw (zero-shot) harmonized (label-free) Feature W→ K→ W→ K→ Frozen REVE (PCA-50) 0.450.45 0.50‡0.50 0.460.46 0.61†0.61 Frozen LaBraM (PCA-50) 0.530.53 0.50‡0.50 0.480.48 0.450.45 Frozen CBraMod (PCA-50) 0.590.59 0.50‡0.50 0.580.58 0.59†0.59 Frozen BIOT (PCA-50) 0.360.36 0.50‡0.50 0.410.41 0.470.47 Frozen BENDR (PCA-50) 0.50‡0.50 0.50‡0.50 0.510.51 0.530.53 Classical (173-d) 0.50‡0.50 0.50‡0.50 0.770.77 0.50‡0.50 DFA exponent (19-d) 0.680.68 0.740.74 0.680.68 0.740.74 5 A Proposed Fix: LRTC-Aware Pretraining Our diagnosis is prescriptive. Because the gap is an objective that never rewards long-range temporal ordering, the natural remedy is an auxiliary objective that does, extending to the temporal-scaling axis the auxiliary-loss remedy Kommineni et al. (2026) proposed for oscillatory structure. The DFA exponent is not smoothly differentiable (its detrend → RMS-over-windows → log-log-fit pipeline is piecewise and rank-like); therefore, rather than back-propagating through the DFA, we provide a differentiable surrogate with two separable terms. Let utu_t be a scalar series read from the pre-pool token sequence (we use the per-patch channel-averaged norm; a learned linear read-out is an alternative) and let its cumulative-sum profile be Yt=∑k≤t(uk−u¯)Y_t= _k≤ t(u_k- u). For dyadic window sizes s∈2js∈\2^j\, compute the detrended fluctuation F(s)F(s) using a closed-form least-squares linear detrend, which is differentiable in Y. Fit logF(s)=α^logs+c F(s)= α s+ c by least squares with α α free, and define ℒLRTC=λ1∑s(logF(s)−[α^logs+c^])2+λ2(α^−αinput)2.L_LRTC= _1 _s ( F(s)-[ α s+ c] )^2+ _2\,( α- _input)^2. The first term is a scale-freeness penalty that drives logF(s) F(s) toward a straight line without constraining its slope. The second is an exponent-fidelity penalty, where αinput _input is the alpha-envelope DFA exponent of that recording, computed on-the-fly from the raw input; thus, the objective remains fully self-supervised and requires no labels. The separation is essential. A single-term loss that penalizes the deviation from a fixed target slope would regularize every recording toward the same exponent. Because the between-subject variation in α is precisely the clinically informative quantity (dementia lowers DFA; §4.4), such a loss would be homogenizing; it would improve log-log linearity while plausibly reducing DFA-recovery R2R^2, the failure it is meant to repair. Only the per-recording fidelity term makes the representation informative about LRTC, and the linearity term alone makes it scale-free but uninformative. Every operation (cumulative sum, least-squares detrend, RMS, log) is differentiable, so this attaches to any masked-reconstruction or contrastive backbone whose architecture can represent long-range temporal order—directly for encoders with an absolute temporal position (REVE, LaBraM, BIOT), whereas for encoders whose positional encoding is purely local (CBraMod, BENDR) the loss would additionally require enriching the positional pathway, since no training signal can exploit order the architecture cannot encode. Scale budget. The number of usable scales is set by the token count: a 240240 s recording at 44 s patches yields ≈60≈ 60 tokens and thus s∈2,4,8,16,32s∈\2,4,8,16,32\, five fit points over roughly 1.21.2 decades. This is a shorter range than the 0.50.5–3030 s used for our encoding target and inherits the same short-range caveat we flag in §3; a longer context or finer patching widens it, and the range should be reported with any result. The success criterion is concrete and falsifiable: a so-adapted encoder should move the DFA-recovery R2R^2 from ≈0≈ 0 toward >0.3>0.3 against its measured reliability ceiling, and improve the zero-shot cross-population transfer over its own frozen baseline. A negative control is equally informative; the fixed-target variant above should not produce this improvement. A synthetic validation harness (fractional-Gaussian-noise inputs of known Hurst exponent, with a fixed-target negative control) is released with the code, but a faithful test is harder to construct than it appears. It requires an input in which LRTC competes with the dominant spectral structure, as in real EEG (on pure fractional noise the exponent is the entire signal, so masked reconstruction already recovers it and the regime cannot isolate the effect) and an input in which the envelope exponent is orthogonal to the aperiodic slope, as it is in our cohorts (r=−0.06r=-0.06, §4.3); satisfying both at once is non-trivial, because envelope-modulating a carrier injects a broadband structure that recouples the exponent to the slope. We leave this, and a single-model head-retrained pilot on a pretrained encoder, to future work. The contribution here is the mechanism, its measurement across five models and three cohorts, and the concrete objective it prescribes. 6 Discussion These three findings converge on a single mechanism. The encoding probe shows no FM recovers DFA on both cohorts; the order control shows, for REVE and CBraMod, that the recovery that does appear is order-independent (for BIOT the order-sensitive component is too small, ≈9%≈\!9\% of the classical baseline, to account for its recovery), and so is not a temporal representation; and the transfer contrast shows that the property the FM drops is the one that shows directionally consistent, though not family-wise significant, cross-population transfer. The common cause of this is the pretraining objective. Reconstruction and patch-contrastive objectives reward matching short-window spectral content, which is why spectral-input models preserve the static 1/f1/f slope for free, but neither rewards preserving long-range temporal correlation, which lives in the ordering of windows across tens of seconds. Mean-pooling is a red herring for REVE, where the pre-pool order control shows that there is no LRTC to pool away in the first place. For LaBraM and BENDR, that control was uninformative (§4.2): their failure to recover DFA from the pooled embedding (Table 2) is established, but whether their pre-pool representations carry LRTC that mean-pooling then discards remains untested, since the shared probe could not resolve it either way. The practical implication is immediate. For the cold-start deployment of an EEG model to a new population, a single dimensionless dynamical exponent is a better representation of the transferable disease signal than a frozen foundation-model embedding, because a scale-free exponent is invariant to the amplitude, reference, and acquisition factors that a site-encoding embedding entangles. This does not argue against foundation models; it argues for an objective that preserves the temporal-scaling structure. Our proposed fix (§5) adds an auxiliary LRTC-prediction loss during pretraining. If a small pilot moves DFA recovery from ≈0≈ 0 to >0.3>0.3 and improves cross-population transfer, this finding would extend from identifying the encoding failure to validating a remedy for it. More broadly, the blind spot is not cosmetic. LRTC is a dimensionless, disease-altered property of resting EEG (Montez et al., 2009) defined purely by temporal ordering, and the foundation models do not encode it: the spectral-input models reproduce the static 1/f1/f spectrum yet discard the temporal correlations, capturing the shape of the signal but not its dynamics. This matters for how such models are used: in the standard frozen-embedding-plus-light-probe or adapter setting, a downstream user targeting an LRTC-linked endpoint (dementia staging, depression, or another scale-free dynamical marker) would build on a representation from which that signal has already been removed. LRTC is only the sharpest instance we test; the same objective has no incentive to preserve any feature that lives in the ordering of windows across tens of seconds, so we read the result as a general caution about what current pretraining keeps, and §5 as one concrete way to change it. 7 Limitations • The transfer process is underpowered. The DFA transfer effect is directionally consistent across all nine tested directions (Table 4), but no DFA-only cell survives Holm correction across the 18-cell DFA/1/f1/f family; the one cell that does survive (DFA+1/f+1/f, K→ ) relies on the non-invariant aperiodic component, which itself inverts the sign on BrainLat (§4.4). The limitation is the small Western training set (power), not its effect. We do not present “criticality transfers cross-population” as our primary claim; we report that the frozen FM is at chance, while the dimensionless DFA exponent is directionally consistent but not yet family-wise significant. • “Blindness” is precise, not blanket. For raw-waveform FMs, both DFA and 1/f1/f are near-zero (R2≤0.12R^2≤ 0.12, with the largest single cell being LaBraM’s 1/f1/f on BrainLat, several-fold below the spectral models on 1/f1/f). For spectral-input FMs, the precise statement is “no temporal-order LRTC”: they carry an order-independent static proxy only. • The pre-pool order control is uninformative for LaBraM and BENDR. Both ordered and shuffled R2R^2 are negative for these two models (§4.2), indicating that the shared 1D-CNN probe underperforms a mean predictor rather than confirming an absence of pre-pool LRTC; the objective-vs-pooling distinction is established only for REVE (finer token-level probe) and the two spectral models. • The site-dominance result is confirmed for all five FMs; however, the transfer-rescue conclusion is not universal. In Table 4 (§4.4), the DFA-transfer cells are model-free (they transfer the dimensionless 1/f1/f+DFA feature vector, not an FM embedding), and the frozen REVE embedding was tested for transfer only on the W↔ pair (Table 5), and not on the BrainLat or LOCO directions in Table 4. Site decoding (§4.5) is now reported for all five FMs on the same leak-free binary task and is uniformly near-perfect (0.980.98–1.001.00, chance 0.5000.500), with no raw-vs-spectral gradient. The subsequent ComBat-harmonization and transfer-rescue arm (Table 5) is also extended to all five FMs (on a freshly-built n=65n=65/770 W/K cohort, because the original REVE analysis’s exact 53-subject subset could not be reconstructed from its cache); the irrecoverability conclusion holds for REVE, LaBraM, and BENDR, but not for CBraMod, whose frozen embedding already carries an above-chance transferable signal before any harmonization (§4.5). • TDBRAIN is excluded from every REVE-embedding analysis. TDBRAIN (Van Dijk et al., 2022) is named in REVE’s own pretraining corpus (El Ouahidi et al., 2025, Appendix B), so a REVE embedding evaluated on TDBRAIN subjects cannot distinguish genuine site- or disease-encoding from memorization of specific recordings seen during pretraining. Therefore, the recording-site probe (§4.5) uses only ds004504 and CAUEEG, both confirmed absent from REVE’s pretraining corpus; the reported 0.9940.994 balanced accuracy (chance 0.5000.500) is unaffected by this exclusion. TDBRAIN’s REVE-based MDD-vs-healthy classification (a separate, earlier analysis, not used in this study) carries the same caveat and is not part of any claim here. • The DFA measurement reliability, and not the target variance, bounds the achievable R2R^2. The alpha-envelope DFA exponent concentrates in [0.5,1.0][0.5,1.0] across cohorts (§3, Targets); R2R^2 itself is scale-invariant, so narrowness alone does not cap it, but split-half measurement reliability does: we estimate an empirical ceiling of R2=0.64R^2=0.64 on CAUEEG (§4.1). The classical baseline’s R2=0.32R^2=0.32–0.380.38 is roughly half this ceiling; the FMs’ R2≈0R^2≈ 0 means that they recover none of a target with substantial recoverable structure, not merely a hard one. • Per-site standardization is transductive, not zero-shot, and is labeled as such wherever it appears. Data availability statement The datasets analyzed in this study were third-party clinical EEG recordings that were not deposited by the authors. ds004504 (Greek AD/FTD/HC cohort) is openly available from OpenNeuro (Miltiadous et al., 2023). ASZED-153 (Nigerian schizophrenia cohort) is openly available via OpenNeuro/Zenodo (Mosaku and others, 2025). TDBRAIN (Dutch psychiatric cohort) and BrainLat (Latin-American AD/bvFTD/HC cohort) are available from the Synapse repository following registration (van Dijk et al., 2022; Prado et al., 2023). CAUEEG (Korean dementia cohort) requires a signed data-use agreement with the dataset custodians (Kim et al., 2023). The analysis code supporting the findings of this study is openly available at https://github.com/MarziehzZare/eeg-fm-criticality-probe (Apache-2.0). No EEG recordings, derived embeddings, or pretrained model weights are redistributed; all model weights are downloaded at runtime from their upstream hosts, and each cohort must be obtained from its custodian under that custodian’s own access terms. Funding statement This research received no external funding, and the work was self-funded by the corresponding author. Conflict of interest statement The corresponding author is affiliated with NeuroGenis Inc., which may have a commercial interest in EEG-based biomarker technology related to this work. The author declares that they have no other competing interests. Ethical statement This study is a secondary analysis of previously collected and de-identified EEG datasets. No new data were collected from human participants. Ethical approval for the original data collection was obtained by the respective data providers and is described in the original publications cited in Table 1. No additional ethical approval was required for this secondary analysis. References J. S. Bindra and S. Panwar (2026) A spectral audit framework reveals task-dependent aperiodic reliance across eeg and ecg deep learning. arXiv preprint arXiv:2606.08583. Cited by: §2. Y. El Ouahidi, J. Lys, P. Thölke, N. Farrugia, B. Pasdeloup, V. Gripon, K. Jerbi, and G. Lioi (2025) REVE: a foundation model for eeg — adapting to any setup with large-scale pretraining on 25,000 subjects. In Advances in Neural Information Processing Systems (NeurIPS), Note: arXiv:2510.21585 Cited by: §3, §3, 5th item. R. Hardstone, S. Poil, G. Schiavone, R. Jansen, V. V. Nikulin, H. D. Mansvelder, and K. Linkenkaer-Hansen (2012) Detrended fluctuation analysis: a scale-free view on neuronal oscillations. Frontiers in Physiology 3, p. 450. Cited by: §1, §3. W. Jiang, L. Zhao, and B. Lu (2024) Large brain model for learning generic representations with tremendous eeg data in bci. In International Conference on Learning Representations (ICLR), Note: spotlight; arXiv:2405.18765 Cited by: §3. Y. Jiao, S. van Cranenburgh, S. Calvert, and H. van Lint (2025) Structure-preserving contrastive learning for spatial time series. arXiv preprint arXiv:2502.06380. Note: contrastive SSL needs structure-preserving regularization to retain temporal/spatial similarity Cited by: §2. A. Kastrati, J. Bürki, J. Lauer, C. Xuan, R. Iaquinto, and R. Wattenhofer (2025) EEG-bench: a benchmark for eeg foundation models in clinical applications. arXiv preprint arXiv:2512.08959. Cited by: §2. M. Kim, Y. C. Youn, and J. Paik (2023) Deep learning-based eeg analysis to classify normal, mild cognitive impairment, and dementia: algorithms and dataset. NeuroImage 272, p. 120054. Note: CAUEEG: Chung-Ang University Hospital EEG dataset Cited by: §3, Data availability statement. A. Kommineni, E. Zhou, K. Avramidis, S. B. Segaard, J. R. Münster, A. P. J. Hansen, T. Medani, T. Feng, R. Leahy, and S. Narayanan (2026) Aperiodic and low-frequency spectral bias in reconstruction-based eeg foundation models. arXiv preprint arXiv:2605.26434. Cited by: §2, §5. D. Kostas, S. Aroca-Ouellette, and F. Rudzicz (2021) BENDR: using transformers and a contrastive self-supervised learning task to learn from massive amounts of eeg data. Frontiers in Human Neuroscience 15, p. 653659. Cited by: §3. W. Lehn-Schiøler, M. R. Kjær, R. Thapa, M. G. Pedersen, A. M. Storgaard, N. Williams, R. Gatej, T. Lehn-Schiøler, A. Brink-Kjær, S. Puthusserypady, S. Beniczky, J. Zou, and L. K. Hansen (2026) Mechanistic interpretability of EEG foundation models via sparse autoencoders. arXiv preprint arXiv:2605.13930. Note: amplitude recoverable from frozen EEG-FM embeddings but phase not, attributed to time-translation invariance of the pretraining objective Cited by: §2. J. Lin, Y. C. Wu, and T. Jung (2026) The identity trap in eeg foundation models: a diagnostic audit. arXiv preprint arXiv:2606.06647. Cited by: §2, §4.5. K. Linkenkaer-Hansen, V. V. Nikouline, J. M. Palva, and R. J. Ilmoniemi (2001) Long-range temporal correlations and scaling behavior in human brain oscillations. Journal of Neuroscience 21 (4), p. 1370–1377. Cited by: §1. A. Miltiadous, K. D. Tzimourta, T. Afrantou, P. Ioannidis, N. Grigoriadis, D. G. Tsalikakis, P. Angelidis, M. G. Tsipouras, E. Glavas, N. Giannakeas, and A. T. Tzallas (2023) A dataset of scalp eeg recordings of alzheimer’s disease, frontotemporal dementia and healthy subjects from routine eeg. Data 8 (6), p. 95. Note: OpenNeuro ds004504 Cited by: §3, Data availability statement. T. Montez, S. Poil, B. F. Jones, I. Manshanden, J. P. A. Verbunt, B. W. van Dijk, A. B. Brussaard, A. van Ooyen, C. J. Stam, P. Scheltens, and K. Linkenkaer-Hansen (2009) Altered temporal correlations in parietal alpha and prefrontal theta oscillations in early-stage alzheimer disease. Proceedings of the National Academy of Sciences 106 (5), p. 1614–1619. Cited by: §1, §6. K. Mosaku et al. (2025) An open-access eeg dataset from indigenous african populations for schizophrenia research. Data in Brief. Note: ASZED-153 / Nigerian Schizophrenia EEG Dataset; OAUTHC Ile-Ife. cf. NSzED arXiv:2311.18484 Cited by: §3, Data availability statement. P. Prado, V. Medel, R. Gonzalez-Gomez, A. Sainz-Ballesteros, V. Vidal, H. Santamaría-García, S. Moguilner, J. Mejia, A. Slachevsky, M. I. Behrens, et al. (2023) The brainlat project, a multimodal neuroimaging dataset of neurodegeneration from underrepresented backgrounds. Scientific Data 10, p. 889. Cited by: §3, Data availability statement. U. Širca, M. Alimardani, S. Zafeiriou, and K. Barmpas (2026) Beyond accuracy: robustness, interpretability and expressiveness of EEG foundation models. arXiv preprint arXiv:2605.17562. Note: argues poor head-only performance is largely a mean-pooling artifact recoverable with token-level embeddings; probes task accuracy, not temporal LRTC Cited by: §3. J. Tai (2026) Pretrained, frozen, still leaking: auditing cross-encoder attribute transfer in eeg foundation models. arXiv preprint arXiv:2606.09189. Cited by: §2. L. Tang, Q. Chen, J. Mei, H. Xu, Q. Zhang, J. Shao, N. Zou, X. Hu, and D. Liu (2026) What do eeg foundation models capture from human brain signals?. arXiv preprint arXiv:2605.11410. Cited by: §2. Y. Tao, B. T. Baker, Y. Wu, A. D. Sarwate, S. Panta, S. Plis, and V. D. Calhoun (2026) Batch effects in brain foundation model embeddings. arXiv preprint arXiv:2604.14441. Note: fMRI foundation models (BrainLM, SwiFT); cited for the batch-effect/harmonization argument Cited by: §2. H. van Dijk, G. van Wingen, D. Denys, S. Olbrich, R. van Ruth, and M. Arns (2022) The two decades brainclinics research archive for insights in neurophysiology (tdbrain) database. Scientific Data 9, p. 333. Note: DOI 10.1038/s41597-022-01409-z Cited by: §3, Data availability statement. J. Wang, S. Zhao, Z. Luo, Y. Zhou, H. Jiang, S. Li, T. Li, and G. Pan (2025) CBraMod: a criss-cross brain foundation model for eeg decoding. In International Conference on Learning Representations (ICLR), Note: arXiv:2412.07236 Cited by: §3. C. Yang, M. B. Westover, and J. Sun (2023) BIOT: biosignal transformer for cross-data learning in the wild. In Advances in Neural Information Processing Systems (NeurIPS), Cited by: §3. A. Zeng, M. Chen, L. Zhang, and Q. Xu (2023) Are transformers effective for time series forecasting?. In Proceedings of the AAAI Conference on Artificial Intelligence, Note: permutation-invariance of attention absent positional signal Cited by: §2.