Paper deep dive
Q-DIVER: Integrated Quantum Transfer Learning and Differentiable Quantum Architecture Search with EEG Data
Junghoon Justin Park, Yeonghyeon Park, Jiook Cha
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 97%
Last extracted: 3/31/2026, 2:22:10 AM
Summary
Q-DIVER is a hybrid framework that integrates a large-scale pretrained EEG encoder (DIVER-1) with a differentiable quantum classifier. It utilizes Differentiable Quantum Architecture Search (DiffQAS) to optimize circuit topologies for EEG classification, achieving performance comparable to classical multi-layer perceptrons while significantly reducing task-specific head parameters.
Entities (5)
Relation Signals (4)
Q-DIVER â employs â DiffQAS
confidence 100% ¡ we employ Differentiable Quantum Architecture Search to autonomously discover task-optimal circuit topologies
Q-DIVER â evaluatedon â PhysioNet Motor Imagery dataset
confidence 100% ¡ On the PhysioNet Motor Imagery dataset, our quantum classifier achieves predictive performance
Q-DIVER â utilizes â DIVER-1
confidence 100% ¡ Q-DIVER, a hybrid framework combining a large-scale pretrained EEG encoder (DIVER-1)
QTSTransformer â componentof â Q-DIVER
confidence 95% ¡ The QTSTransformer head introduces substantially fewer task-specific parameters
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Integrating quantum circuits into deep learning pipelines remains challenging due to heuristic design limitations. We propose Q-DIVER, a hybrid framework combining a large-scale pretrained EEG encoder (DIVER-1) with a differentiable quantum classifier. Unlike fixed-ansatz approaches, we employ Differentiable Quantum Architecture Search to autonomously discover task-optimal circuit topologies during end-to-end fine-tuning. On the PhysioNet Motor Imagery dataset, our quantum classifier achieves predictive performance comparable to classical multi-layer perceptrons (Test F1: 63.49\%) while using approximately \textbf{50$\times$ fewer task-specific head parameters} (2.10M vs. 105.02M). These results validate quantum transfer learning as a parameter-efficient strategy for high-dimensional biological signal processing.
Tags
Links
- Source: https://arxiv.org/abs/2603.28122v1
- Canonical: https://arxiv.org/abs/2603.28122v1
Trouble viewing inline? Open PDF directly â
Full Text
35,795 characters extracted from source content.
Expand or collapse full text
Q-DIVER: Integrated Quantum Transfer Learning and Differentiable Quantum Architecture Search with EEG Data Junghoon Justin Park* Seoul National University Seoul, Korea utopie9090@snu.ac.kr Yeonghyeon Park* Seoul National University Seoul, Korea mandy1002@snu.ac.kr Jiook Cha Seoul National University Seoul, Korea connectome@snu.ac.kr * These authors contributed equally AbstractâIntegrating quantum circuits into deep learning pipelines remains challenging due to heuristic design limitations. We propose Q-DIVER, a hybrid framework combining a large- scale pretrained EEG encoder (DIVER-1) with a differentiable quantum classifier. Unlike fixed-ansatz approaches, we employ Differentiable Quantum Architecture Search to autonomously discover task-optimal circuit topologies during end-to-end fine- tuning. On the PhysioNet Motor Imagery dataset, our quan- tum classifier achieves predictive performance comparable to classical multi-layer perceptrons (Test F1: 63.49%) while using approximately 50Ă fewer task-specific head parameters (2.10M vs. 105.02M). These results validate quantum transfer learning as a parameter-efficient strategy for high-dimensional biological signal processing. Index TermsâQuantum Machine Learning, Quantum Transfer Learning, Differentiable Quantum Architecture Search, EEG Classification, Quantum Time-series Transformer. I. INTRODUCTION Recent progress in quantum machine learning has demon- strated that quantum circuits can be integrated into classical deep learning pipelines in various ways. These integration- focused approaches have been instrumental in establishing the feasibility of classicalâquantum hybrid models. However, they leave a more fundamental question open: at this stage, which functional components of a classical learning pipeline should quantum models meaningfully replace rather than merely augment? From a practical standpoint, classical machine learning pipelines are already highly mature, while near-term quantum devices operate under significant constraints. As a result, replacing entire classical systems with quantum modelsâ or introducing quantum components without a clearly de- fined functional roleâis neither necessary nor well-motivated. These considerations call for a more deliberate and structurally informed approach to integrating quantum models, in which their functional role within the classical pipeline is explicitly specified rather than implicitly assumed. In particular, this motivates a systematic exploration of where and how quantum components should be incorporated, rather than relying on ad hoc circuit choices. This perspective aligns with a long-standing view in clas- sical machine learning that learning systems can be decom- posed into representation learning and task-specific decision making [1], [2]. More recently, the emergence of foundation models has elevated representation reuse into a central de- sign paradigm, explicitly treating pretrained representations as reusable assets across downstream tasks [3]. On a technical level, this paradigm is fundamentally enabled by transfer learn- ing, where large-scale pretraining decouples representation learning from task-specific decision making. Among the many application domains that have adopted this paradigm, electrophysiological signal analysis has applied large-scale self-supervised pretraining to learn transferable spatiotemporal representations [4], [5]. However, EEG signals exhibit substantial inter- and intra-subject variability and are non-stationary, leading to pronounced distribution shifts across subjects and sessions that challenge downstream generaliza- tion [6], [7]. Recent studies have suggested that downstream readout mechanisms can play an important role in EEG transfer learning, and that lightweight or generic classifiers may not fully exploit pretrained representations [8], [9]. Nevertheless, the readout stage is typically treated as part of a broader optimization pipeline rather than examined as a primary object of analysis. In this work, we investigate the downstream readout stage as a standalone modeling component and explore the use of quantum models as alternative readout mechanisms acting on fixed classical representations. Motivated by prior studies on quantum time-series transformers, we repurpose this architecture as a quantum readout module. To avoid attributing potential performance gains to arbitrary circuit choices, we adopt differentiable quantum architecture search (DiffQAS) as a systematic optimization framework, enabling controlled analysis of the structural role of the quantum readout. I. BACKGROUND A. Quantum Transfer Learning Transfer learning considers settings in which the source and target domains and/or tasks differ [10]. The transfer learning paradigm has been extended to hybrid classicalâ quantum architectures [11], where transfer is defined by the reuse of representations across stages, independent of whether arXiv:2603.28122v1 [quant-ph] 30 Mar 2026 the source or target models are classical or quantum. Among various settings, the classicalâquantum configurationâwhere a classical backbone provides representations to a quantum readoutâhas received particular attention in the Near-Term Intermediate Scale Quantum (NISQ) era due to its practical- ity [11]. Recent quantum transfer learning studies explore different ways of realizing transfer within this hybrid framework. Tseng et al. [12] formulate transfer learning directly at the level of the variational quantum circuit (VQC), modeling domain adaptation as a one-step algebraic estimation of parameter shifts under a fixed circuit and measurement. In contrast, many hybrid approaches freeze a deep classical network pretrained on large-scale datasets (e.g., ImageNet) as a feature extractor and train a VQC as a task-specific classification head on the target domain [11], [13]. B. DIVER-1 Recent advances in electrophysiology foundation models show that large-scale pretraining on heterogeneous neural recordings yields representations that generalize across sub- jects, recording setups, and downstream tasks. DIVER-1 [14] exemplifies this through scale, training strategy, and architec- tural design. DIVER-1 is pretrained on a diverse corpus of EEG and iEEG data from over 17,700 subjects, spanning substantial inter-subject and cross-modality variability [14]. Its data- constrained scaling analysis shows that representation quality depends not only on parameter count, but also on training duration and data diversity, with smaller models trained longer sometimes outperforming larger models trained briefly. Prior work suggests that transfer gains in data-limited sci- entific domains may reflect overparameterization or training dynamics rather than genuine feature reuse [15]. The scaling behavior observed in DIVER-1 supports the view that transfer- able electrophysiological representations cannot be attributed to model size alone. Architecturally, DIVER-1 incorporates permutation equiv- ariance and support for heterogeneous sensor configurations to promote cross-subject and cross-setup robustness [14]. These choices reduce sensitivity to electrode ordering and recording systems, making DIVER-1 a suitable backbone for studying how pretrained electrophysiological representations are leveraged by downstream prediction heads. C. Quantum Time-series Transformer To effectively capture long-range temporal dependencies within the high-dimensional EEG embeddings while mini- mizing parameter overhead, we employ the Quantum Time- series Transformer (QTSTransformer) [16]. Unlike classical transformers that rely on computationally intensive quadratic attention mechanisms (O(L 2 )), the QTSTransformer leverages quantum mechanical properties to process sequence data with polylogarithmic complexity (O(polylog(L))). The modelâs ar- chitecture is defined by four distinct operational stages: 1) Unitary Temporal Embedding: The first stage maps classical temporal data into the quantum Hilbert space. Let X = x 0 ,...,x nâ1 denote the sequence of latent feature vectors extracted by the classical backbone. Each feature vector x j at time-step j is mapped to a unique Variational Quantum Circuit (VQC), which defines a specific unitary transformation U j (θ j ). This process creates a quantum repre- sentation where the temporal information is encoded directly into the operation of the quantum gate sequence. 2) Time Sequence Mixing via LCU: To model temporal correlations, we employ the Linear Combination of Unitaries (LCU) primitive. This serves as a quantum-native attention mechanism. Instead of explicitly computing a pairwise at- tention matrix, the model prepares a control register in a superposition state defined by coefficients Îą. This control state directs the simultaneous application of the sequence unitaries U 0 ,...,U nâ1 onto a target register. The resulting operation creates a weighted superposition operator M , defined as: M = nâ1 X j=0 e iÎł j |a j | 2 U j (1) where e iÎł j represents the phase and|a j | 2 represents the magni- tude of the attention weight for the j-th time step. This allows the model to âattendâ to multiple time-steps simultaneously through quantum interference. 3) Non-linearity via QSVT: To introduce the non-linearity required for complex decision boundaries, we utilize the Quan- tum Singular Value Transformation (QSVT). While classical transformers rely on activation functions like Softmax or GELU, the QTSTransformer applies a polynomial transfor- mation P c to the singular values of the mixed operator M . The transformed operator takes the form: P c (M ) = c d M d + c dâ1 M dâ1 +¡ + c 1 M + c 0 I(2) where c 0 ,...,c d are trainable polynomial coefficients and d is the degree of the polynomial. This transformation allows the model to capture richer, higher-order interactions between different time-steps without collapsing the quantum state. 4) Readout and Classical Processing: In the final stage, the transformed quantum state is measured to extract clas- sical features via Pauli expectation values. These values are subsequently processed by a compact classical feed-forward neural network to produce the final downstream prediction (e.g., classification of motor imagery tasks). This architecture serves as the decision head of our hybrid framework. By processing the pre-trained features through this quantum-native attention mechanism, we achieve high repre- sentational power with significantly fewer trainable parameters than a comparable classical multi-layer perceptron (MLP). D. Differentiable Quantum Architecture Search Performance of variational quantum circuits can be highly sensitive to architectural design and initialization, making it difficult to disentangle intrinsic quantum limitations from incidental circuit choices [17], [18]. To enable principled eval- uation and task-specific adaptation, we adopt a differentiable quantum architecture search (DiffQAS) framework [19], which relaxes discrete circuit design into a continuous, end-to-end optimizable formulation. 1) Search Space Construction: We construct a modular search space in which a quantum circuit C consists of L sequential units S 1 ,...,S L . Each unit S l is selected from a predefined candidate library B l . The total number of possible circuit configurations is N = L Y l=1 |B l |.(3) The candidate library includes different entanglement pat- terns and parameterized single-qubit rotation gates (e.g., R x ,R y ,R z ), allowing exploration of varying expressivity and correlation structures. 2) Differentiable Optimization Framework: Rather than se- lecting a single discrete architecture, we assign a learnable structural weight w j to each candidate configuration C j . Let f C j (x;θ j ) denote the output of candidate C j with parameters θ j . The effective model output is defined as the weighted ensemble f ens (x) = N X j=1 w j f C j (x;θ j ).(4) This continuous relaxation enables joint optimization of circuit parameters and architecture. 3) Training and Discretization: We jointly optimize the variational parameters Î = θ 1 ,...,θ N and structural weights W =w 1 ,...,w N by minimizing min Î,W L(f ens (x),y).(5) After convergence, the final discrete architecture is obtained by selecting the candidate with the largest structural weight: C final = arg max j w j .(6) In our implementation, DiffQAS is applied in a factorized manner by optimizing the Timestep Modeling and QFF blocks separately and composing the selected candidates. I. Q-DIVER Our Q-DIVER is a hybrid framework that integrates the classical and quantum components into a unified pipeline designed for data-efficient fine-tuning on downstream EEG tasks. A. Q-DIVER Hybrid Architecture The Q-DIVER pipeline effectively bridges the gap between high-dimensional classical data and the quantum feature space. As illustrated in Fig. 1, the architecture operates as a two-stage transfer learning system. Pretrained Foundation Model Resampling & Tokenization â resampled to 500 Hz/Window = 4 s â T = 2000 â N = 4 tokens (patch=500) DIVER-1 Backbone ⢠Multi-scale feature extraction ⢠Patch embedding (500 samples) ⢠Transformer encoder (12 layers) Reshape for Quantum Input ⢠Feature projection â rotation params Class Logits QTSTransformer Timestep Circuit (DiffQAS, 24candidates) ⢠QSVT polynomial evolution ⢠LCU timestep mixing QFF Circuit. (DiffQAS, 24 candidates) Quantum Head Raw EEG Signal PhysioNet-MI (fs=160 Hz) Fig. 1: Overview of the Q-DIVER hybrid architecture. Raw EEG signals are processed by the pretrained DIVER-1 back- bone and passed to a QTSTransformer-based quantum head. DiffQAS is applied to the Timestep and QFF circuit blocks (24 candidates each). 1) The Hybrid Pipeline: The data flow proceeds sequen- tially through the classical backbone and the quantum classi- fication head: 1) Input Processing: The pipeline accepts raw EEG sig- nals denoted as X âR BĂCĂT , where B is the batch size, C is the number of channels, and T represents time points. 2) Classical Backbone (DIVER-1): The pre-trained DIVER-1 encoder processes the input X to extract high- level spatio-temporal representations. Unlike standard transfer learning where the backbone is often frozen, we perform end-to-end fine-tuning, allowing the backbone to adapt its latent space specifically for the quantum manifold. The output is a latent representation H class â R BĂNĂd , where N is the sequence length (patches) and d is the feature dimension. 3) Quantum Projection Interface: To encode the classical features into the quantum circuit, we employ a learnable projection layer with regularization. For each timestep t, the classical feature vector x t âR d feat is mapped to quantum rotation parameters θ t â [0, 1] n params via: θ t = Ď(W ¡ Dropout(x t ) + b)(7) where Ď(¡) is the sigmoid function, ensuring parameters remain within the normalized range. These parameters are then applied to the quantum gates as rotation angles scaled by Ď/2: R Îą (θ t,i ) = exp âi¡ Ď¡ θ t,i 2 ¡ Ď Îą (8) where Îą â X,Y,Z corresponds to the selected Pauli rotation gate. 4) Quantum Classification Head: The projected features drive the QTSTransformer. The quantum head applies the optimal circuit ansatz discovered via DiffQAS to perform unitary evolution using LCU and QSVT. Fi- nally, quantum measurement (expectation values of Pauli operators) yields the logits Y out âR BĂK , where K is the number of classes. B. Parameter Breakdown and Efficiency The QTSTransformer head introduces substantially fewer task-specific parameters than a classical dense MLP head applied to flattened spatio-temporal features. The DIVER-1 backbone contains 51.36M parameters. In the end-to-end fine- tuning setting, both the backbone and the classifier head are trainable. TABLEI:Parameterbreakdowncomparisonbetween DIVER+MLP and DIVER+QTSTransformer (PhysioNet-MI, 4-second trial setting, 8 qubits). Both models are trained end- to-end. ComponentDIVER+MLP DIVER+QTSTransformer DIVER-1 Backbone51.36M51.36M Feature Projectionâ2.10M Classifier105.02M143 Total156.38M53.46M As shown in Table I, the DIVER+MLP configu- ration contains 156.38M trainable parameters, whereas DIVER+QTSTransformer contains 53.46M parameters under end-to-end training. This corresponds to an overall parameter reduction of approximately 2.9Ă at the full-model level. The difference becomes substantially more pronounced when focusing only on the task-specific classifier head. The classical MLP head introduces 105.02M parameters, while the QTSTransformer head requires only 2.10M parameters, yield- ing an approximate 50Ă reduction in head-specific trainable parameters. Importantly, in the quantum configuration, the majority of head parameters reside in the classical feature projection layer(2.10M), while the quantum circuit itself introduces only 43 trainable parameters. This corresponds to approximately 0.002% of the head parameters and about 0.00008% of the total model parameters, confirming that the quantum compo- nent contributes a negligible fraction of the overall parameter count. IV. EXPERIMENT A. Pretraining The backbone encoder used in this study follows the DIVER-1 self-supervised pretraining procedure [14] and is pretrained on a large-scale corpus of EEG and iEEG record- ings aggregated from multiple publicly available datasets. Raw EEG and iEEG signals were standardized following the DIVER-1 preprocessing protocol: resampling to 500 Hz and minimal filtering with a 0.3â0.5 Hz high-pass and a 60 Hz notch filter, without low-pass filtering. Each training sample consisted of a 30-second continuous window tokenized into temporal patches of either 0.1 s or 1.0 s duration. The pretraining data include intracranial EEG datasets such as AJILE12 [20] and the Penn Electrophysiology of Encoding and Retrieval Study (PEERS) [21], a self-collected iEEG dataset from epilepsy patients, as well as large-scale scalp EEG datasets including the Temple University Hospital EEG Corpus (TUEG) [22], the Healthy Brain Network EEG dataset (HBN-EEG) [23], and the Nationwide Childrenâs Hospital Sleep DataBank (NCHSDB) [24]. These datasets span diverse recording modalities and subject populations, supporting ro- bust representation learning across heterogeneous electrophys- iological signals. The pretraining objective operates on multiple signal repre- sentations. In addition to reconstructing the raw time-domain signal, masked patches are reconstructed in complementary domains such as spectral representations, and the correspond- ing errors are aggregated into a single loss. To improve robustness to heterogeneous recording configurations, random subsampling was applied during pretraining: for each 30- second window, up to 32 channels and up to 30 temporal patches were selected. For the 0.1 s patch configuration, the number of patches was capped at 30. Unmasked patches are processed by a lightweight patch- wise convolutional neural network to extract local temporal features and project them into token representations compat- ible with the transformer backbone. The resulting pretrained encoder is reused for downstream tasks with different predic- tion heads. B. Downstream Task We evaluate Q-DIVER on a publicly available EEG dataset not used during pre-training: the PhysioNet Motor Imagery (PhysioNet-MI) dataset [25]. PhysioNet-MI includes EEG recordings from 109 subjects acquired with 64 channels at 160 Hz. The dataset comprises four motor imagery classes (left fist, right fist, both fists, both feet), and we evaluate a 4-class motor imagery classification task. C. QAQC and Preprocessing Quality assessment and minimal preprocessing followed the same protocol used in DIVER-1 pretraining, including amplitude normalization and clipping-based rejection. TABLE I: Search Space for Differentiable Quantum Archi- tecture Search ComponentOptions InitializationHadamard, None Entanglement Linear CNOT, Ring CNOT, CRX Forward, CRX Backward Variational GatesRX, RY, RZ D. Experimental Setting We evaluate the framework on the downstream EEG dataset (PhysioNet-MI) using a rigorous fine-tuning protocol. Raw EEG recordings (originally sampled at 160 Hz for PhysioNet- MI) were re-referenced and minimally filtered using a 0.3â0.5 Hz high-pass filter and a 60 Hz notch filter. All signals were then resampled to 500 Hz to maintain consistency with the DIVER-1 pretraining configuration. Each trial in the PhysioNet-MI dataset corresponds to a 4-second motor imagery interval. We directly used these trial segments as input samples without introducing additional temporal windowing. 1) Hyperparameters and Configuration: The hybrid model is trained end-to-end using the Cross-Entropy loss function. We utilize the AdamW optimizer to update both the classical backbone weights and the quantum circuit parameters. The specific hyperparameters used for fine-tuning are: batch size 32, learning rate 5e â5 , weight decay 1e â2 , and dropout 0.1 (applied to the classical projection layer). 2) DiffQAS Search Space: Instead of relying on stochastic Monte Carlo sampling [19], we employ a structured factorized search over the same candidate space. Each of the two circuit blocks (Timestep Modeling and QFF) is optimized indepen- dently as a 24-way softmax mixture, introducing 24 + 24 = 48 learnable structural weights while implicitly representing all 24Ă 24 = 576 joint configurations. Given the moderate search size, all candidates within each block are evaluated deterministically at every step. After a short warmup with uniform mixing (1/24), structural weights are optimized via softmax relaxation, and the final architecture is obtained by argmax selection in each block. V. RESULTS We present the PhysioNet-MI experiment results demon- strating the modelâs classification performance, architectural convergence, and parameter efficiency. A. Classification Performance As shown in Table I, Q-DIVER achieved a Test F1- score of 63.49%, exceeding the validation F1-score of 59.81%, which suggests stable performance on the held-out test set. TABLE I: Classification performance metrics for the Q- DIVER model on the PhysioNet-MI dataset SplitAccuracyF1-ScoreKappaPrecisionRecall Validation59.77%59.81%0.463659.89%59.77% Test63.39%63.49%0.511863.91%63.38% B. Optimal Quantum Architecture Discovery Q-DIVER successfully navigated the search space of 576 potential architecture combinations (24 candidates per com- ponent) to identify a task-optimal circuit configuration. The search process converged on specific structural motifs that maximize expressivity for EEG signal classification. ⢠Timestep Circuit: The search selected a 2-layer architec- ture utilizing a CRX Backward Ring entangling pattern and RZ variational gates (Fig. 2a). ⢠Quantum Feed-Forward (QFF) Circuit: The search selected a 1-layer architecture utilizing a CRX Forward Ring entangling pattern and RY variational gates (Fig. 2b). q0 ⢠⢠q6 q7 q1 q2 q3 q4 q5 q6 ⢠⢠RZ RZ RZ RZ RZ RZ RZ RZ ⢠RX RZ ⢠RX RZ ⢠RX RZ ⢠RX RZ ⢠RX RZ ⢠RX RZ RZ RX RXRZ RX RX ⢠RX ⢠RX ⢠RX ⢠RX ⢠RX ⢠RX (a) Timestep Circuit q0 q7 q1 q2 q3 q4 q5 q6 RY RY RY RY RY RY RY RX RX RX RX RX RX RX RY ⢠⢠⢠⢠⢠⢠RX ⢠⢠⢠⢠⢠⢠⢠⢠(b) Quantum Feed Forward Fig. 2: Optimal Quantum Circuit Architecture Obtained from DiffQAS. Three distinct patterns emerged from the optimal architec- ture: ⢠Preference for Parametric Entanglement: Both optimal circuits selected CRX gates over static CNOT gates. This suggests that the additional learnable parameters in the entangling layers are crucial for capturing the complex correlations in high-dimensional EEG data. ⢠No Initialization Layer: Neither the timestep nor the QFF circuit selected the Hadamard initialization layer. This indicates that for this specific motor imagery task, starting from the computational basis state |0⊠ân and relying solely on variational rotations provides sufficient expressivity. ⢠Complementary Variational Gates: The pipeline uti- lizes complementary rotation axesâRZ (phase) for temporal modeling and RY (amplitude) for feature extractionâenhancing the diversity of transformations within the quantum Hilbert space. C. Architecture Weight Convergence The differentiable search process exhibited stable conver- gence behavior. Initialized uniformly at w = 1/24 â 0.0417 during the 5-epoch warmup phase, the architecture weights for the optimal circuits gradually increased throughout training. By epoch 99, the weights for the optimal Timestep and QFF circuits reached approximately 0.063 and 0.065, respectively. This gradual divergence, rather than a sharp winner-takes-all collapse, suggests that the soft selection mechanism allows the model to robustly explore the search space before settling on the optimal configuration. D. Parameter Efficiency For PhysioNet-MI with 4-second motor imagery trials (8 qubits), DIVER+QTSTransformer has 53.46M parameters (51.36M backbone + 2.10M head), whereas DIVER+MLP has 156.38M (51.36M + 105.02M). With end-to-end fine-tuning (unfrozen backbone), this yields an overall reduction of 2.9Ă at the full-model level. Focusing on the task-specific readout head, QTSTrans- former uses 2.10M parameters versus 105.02M for the MLP head, corresponding to an approximate 50Ă reduction. Most head parameters in the quantum configuration come from the classical feature projection (2.097M); the quantum circuit contributes 43 trainable parameters and the post-measurement linear classifier adds 100. This compact decision head suggests that the optimized quantum ansatz can provide an expressive readout under severe parameter constraints. VI. CONCLUSION In this paper, we introduced Q-DIVER, a hybrid frame- work synergizing the DIVER-1 foundation model with the QTSTransformer to enable quantum transfer learning for high-dimensional EEG data. By extending DiffQAS, our ap- proach autonomously discovered task-optimal circuit topolo- gies, eliminating the reliance on heuristic ansatz design. Empirical results on the PhysioNet-MI dataset highlight three key contributions. First, the quantum classifier matched classical MLP performance while using approximately 50Ă fewer task-specific head parameters, with detailed parameter comparisons provided in Table I and the Results section. This demonstrates that substantial parameter savings arise primarily from the decision-making module, supporting the high effec- tive dimension of quantum feature spaces [26]. Second, the architecture search consistently converged on expressive para- metric entangling gates (CRX) and eschewed initialization lay- ers, revealing distinct, task-specific inductive biases [19], [27]. Finally, the model exhibited robust generalization with test performance exceeding validation metrics (Test F1: 63.49%), consistent with theoretical predictions regarding the favorable generalization bounds of constrained quantum models [28]â [30]. Q-DIVER offers a viable pathway for deploying power- ful, interpretability-friendly neuroimaging models in resource- constrained environments. Beyond empirical performance, our findings cautiously suggest that the structural constraints of quantum circuits may induce a task-relevant inductive bias, enabling effective decision boundaries under severe parameter limitations. Future work will expand this framework to diverse neurological domains and integrate quantum error mitigation to enhance robustness on NISQ hardware. ACKNOWLEDGMENT Special thanks to the members of the SNU Connectome Lab, particularly Taeyang Lee, for their invaluable support to implementing hybrid quantum transfer learning of theDIVER-1model.Thisworkwassupportedby theNationalResearchFoundationofKorea(NRF) grant funded by the Korea government (MSIT) (No. 2021R1C1C1006503,RS-2023-00266787,RS-2023- 00265406,RS-2024-00421268,RS-2024-00342301,RS- 2024-00435727, NRF-2021M3E5D2A01022515, and NRF- 2021S1A3A2A02090597),bytheCreative-Pioneering Researchers Program through Seoul National University (No. 200-20240057, 200-20240135). Additional support was provided by the Institute of Information & Communications Technology Planning & Evaluation (IITP) grant funded by the Korea government (MSIT) [No. RS-2021-I211343, 2021-0-01343,ArtificialIntelligenceGraduateSchool Program, Seoul National University] and by the Global Research Support Program in the Digital Field (RS-2024- 00421268). This work was also supported by the Artificial Intelligence Industrial Convergence Cluster Development Project funded by the Ministry of Science and ICT and Gwangju Metropolitan City, by the Korea Brain Research Institute (KBRI) basic research program (25-BR-05-01), by the Korea Health Industry Development Institute (KHIDI) and the Ministry of Health and Welfare, Republic of Korea (HR22C1605), and by the Korea Basic Science Institute (National Research Facilities and Equipment Center) grant funded by the Ministry of Education (RS-2024-00435727). We acknowledge the National Supercomputing Center for providing supercomputing resources and technical support (KSC-2023-CRE-0568). An award for computer time was provided by the U.S. Department of Energyâs (DOE) ASCR Leadership Computing Challenge (ALCC). This research used resources of the National Energy Research Scientific Computing Center (NERSC), a DOE Office of Science User Facility, under ALCC award m4750-2024, and supporting resources at the Argonne and Oak Ridge Leadership Computing Facilities, U.S. DOE Office of Science user facilities at Argonne National Laboratory and Oak Ridge National Laboratory. REFERENCES [1] J. Yosinski, J. Clune, Y. Bengio, and H. Lipson, âHow transferable are features in deep neural networks?,â Advances in neural information processing systems, vol. 27, 2014. [2] Y. Bengio, A. Courville, and P. Vincent, âRepresentation learning: A review and new perspectives,â IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 35, no. 8, p. 1798â1828, 2013. [3] R. Bommasani, D. Hudson, E. Adeli, R. Altman, S. Arora, S. Arx, M. Bernstein, J. Bohg, A. Bosselut, E. Brunskill, E. Brynjolfsson, S. Buch, D. Card, R. Castellon, N. Chatterji, A. Chen, K. Creel, J. Davis, D. Demszky, and P. Liang, On the Opportunities and Risks of Foundation Models. 2021. [4] G. Wang, W. Liu, Y. He, C. Xu, L. Ma, and H. Li, âEEGPT: pretrained transformer for universal and reliable representation of EEG signals,â 2024. [5] W. Jiang, L. Zhao, and B.-l. Lu, âLarge brain model for learning generic representations with tremendous EEG data in bci,â in The Twelfth International Conference on Learning Representations. [6] A. Apicella, P. Arpaia, G. DâErrico, D. Marocco, G. Mastrati, N. Moc- caldi, and R. Prevete, âToward cross-subject and cross-session general- ization in eeg-based emotion recognition: Systematic review, taxonomy, and methods,â Neurocomputing, vol. 604, p. 128354, 2024. [7] J. Wang, S. Zhao, Z. Luo, Y. Zhou, H. Jiang, S. Li, T. Li, and G. Pan, âCBramod: A criss-cross brain foundation model for EEG decoding,â in The Thirteenth International Conference on Learning Representations, 2025. [8] S. Liang, L. Li, W. Zu, W. Feng, and W. Hang, âAdaptive deep feature representation learning for cross-subject EEG decoding,â BMC Bioinformatics, vol. 25, no. 1, p. 393, 2024. [9] J. Jeon, S. Jeong, Y. Shon, and H.-I. Suk, âParameter-efficient transfer learning for EEG foundation models via task-relevant feature focusing,â 2025. [10] S. J. Pan and Q. Yang, âA survey on transfer learning,â IEEE Transac- tions on knowledge and data engineering, vol. 22, no. 10, p. 1345â1359, 2009. [11] A. Mari, T. R. Bromley, J. Izaac, M. Schuld, and N. Killoran, âTransfer learning in hybrid classical-quantum neural networks,â Quantum, vol. 4, p. 340, 2020. [12] H.-H. Tseng, H.-Y. Lin, S. Y.-C. Chen, and S. Yoo, âTransfer learning analysis of variational quantum circuits,â in 2025 IEEE International Conference on Acoustics, Speech, and Signal Processing Workshops (ICASSPW), p. 1â5, IEEE. [13] A. Khatun and M. Usman, âQuantum transfer learning with adver- sarial robustness for classification of high-resolution image datasets,â Advanced Quantum Technologies, vol. 8, no. 1, p. 2400268, 2025. [14] D. D. Han, Y. Gwon, A. L. Lee, T. Lee, S. J. Lee, J. Choi, S. Lee, J. Bang, S. Lee, and D. K. Park, âDIVER-1: Deep integra- tion of vast electrophysiological recordings at scale,â arXiv preprint arXiv:2512.19097, 2025. [15] M. Raghu, C. Zhang, J. Kleinberg, and S. Bengio, âTransfusion: Un- derstanding transfer learning for medical imaging,â Advances in neural information processing systems, vol. 32, 2019. [16] J. J. Park, J. Seo, S. Bae, S. Y.-C. Chen, H.-H. Tseng, J. Cha, and S. Yoo, âResting-state fMRI analysis using quantum time-series transformer,â in 2025 IEEE International Conference on Quantum Computing and Engineering (QCE), vol. 1, p. 2352â2363, IEEE, 2025. [17] J. R. McClean, S. Boixo, V. N. Smelyanskiy, R. Babbush, and H. Neven, âBarren plateaus in quantum neural network training landscapes,â Nature Communications, vol. 9, no. 1, p. 4812, 2018. [18] Z. Holmes, K. Sharma, M. Cerezo, and P. J. Coles, âConnecting ansatz expressibility to gradient magnitudes and barren plateaus,â PRX Quantum, vol. 3, p. 010313, Jan 2022. [19] S.-X. Zhang, C.-Y. Hsieh, S. Zhang, and H. Yao, âDifferentiable quantum architecture search,â Quantum Science and Technology, vol. 7, p. 045023, aug 2022. [20] S. M. Peterson, S. H. Singh, B. Dichter, M. Scheid, R. P. N. Rao, and B. W. Brunton, âAJILE12: Long-term naturalistic human intracranial neural recordings and pose,â Scientific Data, vol. 9, no. 1, p. 184, 2022. [21] M. J. Kahana, J. H. Rudoler, L. J. Lohnas, K. Healey, A. Aka, A. Broitman, E. Crutchley, P. Crutchley, K. H. Alm, B. S. Katerman, N. E. Miller, J. R. Kuhn, Y. Li, N. M. Long, J. Miller, M. D. Paron, J. K. Pazdera, I. Pedisich, and C. T. Weidemann, âPenn Electrophysiology of Encoding and Retrieval Study (PEERS),â 2023. [22] I. Obeid and J. Picone, âThe Temple University Hospital EEG Data Corpus,â Frontiers in Neuroscience, vol. 10, p. 196, 05 2016. [23] S. Y. Shirazi, A. Franco, M. S. Hoffmann, N. B. Esper, D. Truong, A. Delorme, M. P. Milham, and S. Makeig, âHBN-EEG: The FAIR implementation of the Healthy Brain Network (HBN) electroencephalog- raphy dataset,â bioRxiv, p. 2024.10.03.615261, 2024. [24] H. Lee, B. Li, S. DeForte, M. L. Splaingard, Y. Huang, Y. Chi, and S. L. Linwood, âA large collection of real-world pediatric sleep studies,â Scientific Data, vol. 9, no. 1, p. 421, 2022. [25] A. L. Goldberger, L. A. N. Amaral, L. Glass, J. M. Hausdorff, P. C. Ivanov, R. G. Mark, J. E. Mietus, G. B. Moody, C.-K. Peng, and H. E. Stanley, âPhysiobank, Physiotoolkit, and Physionet: components of a new research resource for complex physiologic signals,â Circulation, vol. 101, no. 23, p. e215âe220, 2000. doi: 10.1161/01.CIR.101.23.e215. [26] A. Abbas, D. Sutter, C. Zoufal, A. Lucchi, A. Figalli, and S. Woerner, âThe power of quantum neural networks,â Nature Computational Sci- ence, vol. 1, no. 6, p. 403â409, 2021. [27] S. Y.-C. Chen and P. Tiwari, âQuantum long short-term memory with differentiable architecture search,â in 2025 IEEE International Confer- ence on Quantum Artificial Intelligence (QAI), p. 13â18, 2025. [28] M. C. Caro, H.-Y. Huang, M. Cerezo, K. Sharma, A. Sornborger, L. Cincio, and P. J. Coles, âGeneralization in quantum machine learning from few training data,â Nature Communications, vol. 13, no. 1, p. 4919, 2022. [29] E. Gil-Fuster, J. Eisert, and C. Bravo-Prieto, âUnderstanding quantum machine learning also requires rethinking generalization,â Nature Com- munications, vol. 15, no. 1, p. 2277, 2024. [30] T. Wu, A. Bentellis, A. Sakhnenko, and J. M. Lorenz, âGeneralization bounds in hybrid quantum-classical machine learning models,â arXiv preprint, 2025.