Paper deep dive
A Drift Stable Quantum Federated Learning for Intelligent Services
Shanika Iroshi Nanayakkara, Shiva Raj Pokhrel
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 7/28/2026, 3:30:15 AM
Summary
The paper introduces DUQFL-Prox, a drift-stable quantum federated learning (QFL) framework designed for heterogeneous intelligent services. It addresses client drift and unstable local updates in QNN training by employing deep-unfolded local optimization with adaptive SPSA updates guided by a lightweight controller. A proximal regularization term keeps local models close to the global model, while a meta-objective refines the controller based on post-aggregation performance. Experiments on financial fraud and genomic classification demonstrate improved stability, generalization, and fairness compared to standard QFL baselines.
Entities (8)
Relation Signals (6)
DUQFL-Prox → uses → SPSA
confidence 95% · each client performs adaptive unfolded SPSA updates
DUQFL-Prox → addresses → Client Drift
confidence 92% · DUQFL-Prox aims to stabilize their local QNN optimization trajectories so that their updates remain useful for global aggregation
DUQFL-Prox → applies → Proximal Regularization
confidence 90% · DUQFL-Prox therefore introduces proximal regularization into the client objective to penalize excessive deviation from the broadcast global model
DUQFL-Prox → evaluatedon → financial fraud detection
confidence 88% · Experiments on financial fraud and genomic classification tasks show that DUQFL-Prox improves stability
DUQFL-Prox → evaluatedon → Genomic Classification
confidence 88% · Experiments on financial fraud and genomic classification tasks show that DUQFL-Prox improves stability
Quantum Federated Learning → enables → Privacy-Aware Services
confidence 85% · Quantum federated learning enables distributed clients to train quantum neural networks without sharing local data, making it promising for privacy-aware intelligent services
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Quantum federated learning enables distributed clients to train quantum neural networks without sharing local data, making it promising for privacy-aware intelligent services. Intelligent services in this context refer to privacy-sensitive distributed decision systems, such as fraud detection and genomic classification, where reliable and fair client-level learning is as important as the accuracy of the aggregate model. However, heterogeneous client data and noisy quantum optimization often cause unstable local updates, client drift, and unfair performance between clients. This paper proposes DUQFL-Prox, a drift-stable quantum federated learning framework based on deep-unfolded local optimization. Instead of using a fixed local optimizer, each client performs adaptive unfolded SPSA updates, while a proximal term keeps the local model close to the global model. A lightweight controller learns step-specific optimization parameters to improve post-aggregation performance. Experiments on financial fraud and genomic classification tasks show that DUQFL-Prox improves stability, generalization, and client fairness compared with standard QFL baselines. The results suggest that deep-unfolded quantum federated learning can support more reliable and fair intelligent services in heterogeneous distributed environments.
Tags
Links
- Source: https://arxiv.org/abs/2607.21647v1
- Canonical: https://arxiv.org/abs/2607.21647v1
Trouble viewing inline? Open PDF directly →
Full Text
93,830 characters extracted from source content.
Expand or collapse full text
A Drift-Stable Quantum Federated Learning for Intelligent Services Shanika Iroshi Nanayakkara and Shiva Raj Pokhrel Shanika Iroshi Nanayakkara is with School of IT, Deakin University, VIC 3125, Burwood, Australia, (e-mail: s222112938@deakin.edu.au). Shiva Raj Pokhrel is with School of IT, Deakin University, VIC 3125, Burwood, Australia, (e-mail: shiva.pokhrel@deakin.edu.au). Abstract Quantum federated learning enables distributed clients to train quantum neural networks without sharing local data, making it promising for privacy-aware intelligent services. Intelligent services in this context refer to privacy-sensitive distributed decision systems, such as fraud detection and genomic classification, where reliable and fair client-level learning is as important as the accuracy of the aggregate model. However, heterogeneous client data and noisy quantum optimization often cause unstable local updates, client drift, and unfair performance between clients. This paper proposes DUQFL-Prox, a drift-stable quantum federated learning framework based on deep-unfolded local optimization. Instead of using a fixed local optimizer, each client performs adaptive unfolded SPSA updates, while a proximal term keeps the local model close to the global model. A lightweight controller learns step-specific optimization parameters to improve post-aggregation performance. Experiments on financial fraud and genomic classification tasks show that DUQFL-Prox improves stability, generalization, and client fairness compared with standard QFL baselines. The results suggest that deep-unfolded quantum federated learning can support more reliable and fair intelligent services in heterogeneous distributed environments. I Introduction Federated learning (FL) has emerged as a promising framework for collaborative model training, where multiple clients optimize a shared model without centralizing raw local data. By allowing clients to train locally and communicate only model updates to a coordinating server, FL provides an attractive paradigm for privacy-aware and communication-efficient intelligence [16]. As quantum machine learning continues to develop, this distributed setting becomes increasingly relevant in quantum contexts as well, giving rise to quantum federated learning (QFL), in which multiple clients collaboratively train quantum or quantum-enhanced models while keeping local data decentralized [36, 7, 23]. In this setting, however, optimization becomes substantially more challenging than in conventional classical FL, because local training of quantum neural networks (QNNs) is often noisy, nonconvex, shot-sensitive, and highly dependent on optimizer configuration [10, 36]. In this work, the phrase “intelligent services” refers to distributed AI-enabled service environments in which data-driven decisions must be learned from decentralized, privacy-sensitive, and heterogeneous client data. Such services include financial fraud detection, genomic classification, healthcare analytics, cyber-physical monitoring [34], and edge intelligence, where data are naturally distributed across institutions, devices, or service providers and cannot be freely centralized due to privacy, regulatory, or ownership constraints [5]. These settings require not only high aggregate accuracy, but also stable optimization, fair client-level performance, and reliable generalization across heterogeneous participants [22]. Figure 1: High-level DUQFL setup with adaptive local QNN training, selected-model aggregation, and controller-guided global optimization. In heterogeneous federated environments, one common strategy for improving training efficiency is to exclude, down-weight, delay, or selectively sample slow clients, often referred to as stragglers. Previous FL studies have shown that system heterogeneity, including differences in computation and communication capabilities, can substantially slow synchronous federated training [16, 12, 28]. Client-selection methods such as Oort therefore improve time-to-accuracy by prioritizing clients with favorable statistical and system utility [15]. However, selection or straggler-mitigation strategies can also create a representativeness issue when slow, resource-limited, or statistically distinctive clients are repeatedly underrepresented. Recent work on biased client selection and participation imbalance shows that non-uniform client participation can bias the learned global model and affect fairness or client-level performance [2, 29]. In practical QFL settings, this issue is undesirable because heterogeneous clients may contain rare, domain-specific, or clinically important data distributions. Rather than excluding such clients, DUQFL-Prox aims to stabilize their local QNN optimization trajectories so that their updates remain useful for global aggregation. A central limitation in existing QFL pipelines is that local optimization is often treated as a fixed procedure. In a typical setting, each client receives the broadcast global model, performs a predetermined number of local optimization steps using a fixed optimizer schedule, and returns the resulting parameters to the server. This design implicitly assumes that the same local training rule is suitable for all clients, data distributions, and communication rounds. Such an assumption is restrictive in QFL, where QNN optimization is sensitive to learning rate, perturbation scale, measurement noise, and the non-convex geometry of variational quantum loss landscapes. In non-IID data, fixed local optimization can therefore produce uneven client updates, excessive client drift, and unstable post-aggregation behavior. SPSA is attractive for QNN training because it estimates an update direction using only two objective-function evaluations, independent of the number of trainable parameters [32]. This property makes SPSA practical for variational quantum models, particularly when analytic gradients are costly, unavailable, or noisy [8]. However, in federated QNN training, a fixed SPSA schedule can be too rigid. Conservative settings may slow local improvement, whereas aggressive settings may amplify client drift and reduce the compatibility of local updates during aggregation. Therefore, the central challenge is not merely to use SPSA, but to make local SPSA-based QNN optimization adaptive while keeping client updates aligned with the global federated objective. To address this challenge, we propose Deep-Unfolded Quantum Federated Learning with Proximal Regularization (DUQFL-Prox), a drift-stable QFL framework for heterogeneous intelligent services. DUQFL-Prox decomposes local SPSA-based QNN training into K unfolded optimization blocks. At each unfold step, a lightweight shared controller generates step-specific SPSA learning rates and perturbation scales from optimization-state features, including unfold progress, recent loss behavior, parameter displacement, and client context. This enables local QNN optimization to adapt across clients, unfold steps, and communication rounds, rather than relying on a fixed optimizer schedule. However, adaptive local optimization alone can amplify client drift under non-IID federated data. DUQFL-Prox therefore introduces proximal regularization into the client objective to penalize excessive deviation from the broadcast global model. This encourages each client to improve their local QNN objective while maintaining aggregation-compatible updates. In addition, each client uploads the validation-preferred unfolded checkpoint rather than necessarily returning the final unfolded state, preventing over-aggressive later updates from degrading global aggregation. DUQFL-Prox follows a bilevel learning structure. At the inner level, each client performs deep-unfolded proximal SPSA optimization of its local QNN parameters. At the outer level, the shared controller is periodically refined using a meta-objective defined on post-aggregation validation behaviour. Thus, DUQFL-Prox learns not only the QNN parameters but also how local QNN optimization should evolve so that client updates become more stable, fair, and useful after aggregation. The main contributions of this work are summarized as follows: • We propose DUQFL-Prox, a drift-stable deep-unfolded QFL framework for heterogeneous intelligent services, where local QNN training is modeled as an adaptive multi-step optimization process rather than a fixed optimizer routine. • We design a controller-driven SPSA mechanism that generates unfold-specific learning rates and perturbation scales from local optimization-state features, enabling adaptive QNN training across clients, communication rounds, and unfolded steps. • We incorporate proximal regularization and validation-based best-unfold selection to reduce client drift, prevent over-specialized local updates, and improve aggregation compatibility. Importantly, we introduce an outer meta-loss-guided controller adaptation mechanism that aligns local optimization behavior with post-aggregation global performance. We evaluate DUQFL-Prox across financial fraud detection and genomic classification tasks using global accuracy, client-level generalization, train–test gap, fairness gap, and imbalance-aware classification metrics. The remainder of this paper is organized as follows. Section I reviews related work on QFL, QNN optimization, federated heterogeneity, and deep unfolding. Section I presents the problem formulation and the DUQFL-Prox methodology. Section IV describes the experimental setup and reproducibility protocol. Section V presents the empirical results and ablation analysis. Section LABEL:sec:discussion discusses implications, limitations, and future research directions, and Section VI concludes the paper. I Related Work TABLE I: Research gap. ✓ indicates explicit support, △ indicates partial support, and × indicates limited or no support. Research stream References QFL QNN Adaptive Client drift L2L Generalization Observed Gap Secure / privacy-preserving QFL [17, 3, 37, 6] ✓ △ × × × × Focuses mainly on privacy/security rather than optimiser stability or heterogeneous client drift. QFL for financial fraud detection [9] ✓ ✓ △ △ × △ Applies quantum federated neural networks to financial fraud detection, but does not explicitly formulate local QNN optimization as a deep-unfolded SPSA trajectory with proximal client-drift control or outer meta-loss-guided controller adaptation. Quantum-secured fraud detection / blockchain-assisted quantum applications [18] × △ × × × △ Studies quantum-state verification, blockchain-assisted verification, and quantum-enhanced fraud detection, but does not address federated QNN optimization, client drift, or learning-to-learn. Post-quantum data integrity and secure outsourcing [38] × × × × × × Addresses post-quantum secure data outsourcing and public integrity verification for cloud storage, but does not focus on QFL, QNN training, or heterogeneous federated optimization. QFL for autonomous / distributed systems [35, 21] ✓ ✓ △ △ × △ Studies distributed or application-specific QFL, but local optimiser trajectories are not meta-controlled. QFL for healthcare and biomedical learning [7, 25, 39] ✓ ✓ △ △ × △ Demonstrates application feasibility, but Often relies on global accuracy without detailed client-level analysis. Quantum natural-gradient / optimization-based QFL [26, 1, 33] ✓ ✓ △ △ × △ Improves local quantum optimization but is usually not designed for federated non- IID settings. Classical privacy-preserving FL with adaptive pruning [24] × × ✓ △ × △ Addresses privacy-preserving FL, adaptive model pruning, communication cost, and non-IID data in classical federated learning, but does not consider QNN training, quantum optimization, deep unfolding, or quantum client drift. Classical deep unfolding / learning-to-optimise [20, 19] × × ✓ △ ✓ △ Provides learnable optimisation principles, but is not designed for quantum neural networks or federated QNN training. Proposed DUQFL-Prox This work ✓ ✓ ✓ ✓ ✓ ✓ Deep-unfolded SPSA-based local QNN optimization with proximal client-drift regularisation, outer/meta validation guidance, and multi-domain evaluation using global accuracy, client accuracy, fairness, and generalization metrics. Table I demonstrates that prior QFL studies have mainly focused on secure quantum communication, privacy-preserving learning, application-specific QFL implementations, and server-side aggregation mechanisms [17, 3, 37, 6, 7, 25, 39, 26, 1, 33]. While these studies establish the feasibility of QFL. However, local QNN optimization is typically treated as a fixed inner routine. This assumption becomes restrictive when client data are non-independent and identically distributed (non-IID) and quantum measurements are stochastic. In such cases, local updates may become unstable, diverge from the global model, and reduce client-level generalization. QNN training commonly relies on optimizers such as gradient descent, Adam, COBYLA, and SPSA. SPSA is particularly suitable for variational quantum models because it estimates a stochastic update direction using only two objective function evaluations, independent of the number of trainable parameters [27, 31]. However, in federated QNN training, a fixed SPSA schedule may be inadequate: conservative settings can slow local improvement, whereas aggressive settings can amplify client drift and reduce aggregation compatibility. Therefore, the key challenge is not merely to use SPSA, but to adapt SPSA-based local QNN optimization while preserving global federated consistency. Deep unfolding provides a principled way to convert iterative optimization into a structured and learnable update process [30]. In parallel, classical FL methods such as FedProx and SCAFFOLD show that client drift is a central obstacle under heterogeneous data and that proximal or correction-based mechanisms can improve stability [16, 13]. Although these methods are not quantum-specific, they motivate the development of adaptive and drift-aware local optimization strategies for QFL. Motivated by these gaps, we propose Deep-Unfolded Quantum Federated Learning with Proximal Regularization (DUQFL-Prox). DUQFL-Prox unfolds local SPSA-based QNN training into multiple controller-guided optimization blocks. The controller generates step-specific learning rates and perturbation scales from local optimization-state features, while a proximal term penalizes excessive deviation from the broadcast global model. The controller is further refined using a post-aggregation meta-objective, linking local optimization behaviour to global validation performance. The proposed framework differs from prior QFL work in three aspects. First, local QNN optimization is modeled as a controller-driven unfolded process rather than a fixed client routine. Second, proximal regularization is incorporated to improve aggregation compatibility under client drift. Third, the controller is updated using post-aggregation validation behaviour, rather than relying only on local training signals. To the best of our knowledge, this combination of unfolded SPSA-based QNN optimization, proximal client-drift control, and outer meta-loss-guided controller adaptation has not been explicitly developed in existing QFL literature. Unlike client-filtering or straggler-removal strategies, DUQFL-Prox does not discard heterogeneous clients. Instead, it stabilizes their local optimization trajectories so that their updates remain useful for aggregation. This is important in biomedical, genomic, fraud-detection, and remote-sensing settings, where difficult clients may contain rare but important patterns. Accordingly, our evaluation reports global accuracy together with mean client test accuracy, train–test gap, client fairness gap, and imbalance-aware metrics such as F1-score, precision, recall, ROC-AUC, PR-AUC, and MCC. I Problem Formulation and Methodology I-A Problem Statement The objective of this work is to improve the stability, generalization, and client-level reliability of QFL under heterogeneous client participation. In conventional QFL, each client receives the broadcast global QNN parameters, performs local optimization, and returns the updated parameters for server aggregation. However, under non-IID data and stochastic quantum measurements, fixed local optimization can produce unstable client trajectories and excessive drift from the global model. Consequently, the aggregated model may achieve reasonable global accuracy while still exhibiting poor client-level generalization or imbalanced performance across clients. Although adaptive optimizers such as Adam can adjust local parameter updates, they are not designed to explicitly account for federated client heterogeneity, quantum measurement noise, post-aggregation behaviour, or drift from the broadcast global model [14]. Similarly, straggler removal or client down-weighting can improve training efficiency, but may reduce representativeness by underutilizing clients with limited resources or statistically distinctive data [16, 12, 28]. This is undesirable in QFL applications such as genomics, biomedical analysis, fraud detection, and remote sensing, where difficult clients may contain rare but important data patterns. DUQFL-Prox addresses this problem by regulating local QNN optimization rather than excluding heterogeneous clients. The proposed method seeks a QFL training procedure in which: 1. local SPSA-based QNN updates adapt across unfolded optimization steps; 2. local trajectories remain sufficiently close to the broadcast global model to remain aggregation-compatible; and 3. the hyperparameter-generation policy is refined using post-aggregation validation performance. Thus, the central problem is to jointly control client-side unfolded QNN optimization and server-side controller adaptation so that the global model achieves stable post-aggregation performance, reduced client drift, and improved client-level generalization under heterogeneous federated data. I-B QFL Setting and Notation We consider a QFL system consisting of N distributed clients and one coordinating server. The clients collaboratively train a shared QNN model without exchanging raw local data. Let the client set be =1,2,…,N.C=\1,2,…,N\. (1) Training proceeds over T communication rounds indexed by t∈0,1,…,T−1.t∈\0,1,…,T-1\. (2) At communication round t, the server maintains a global QNN parameter vector (t)∈ℝP, θ^(t) ^P, (3) where P denotes the number of trainable parameters in the variational quantum model. These parameters define a parameterized quantum circuit, denoted by U(;(t)),U(x; θ^(t)), (4) where x is the classical input encoded into the quantum circuit and (t) θ^(t) parameterizes the trainable ansatz layers. The QNN output is obtained by measuring the resulting quantum state and applying classical post-processing to obtain prediction probabilities. The server broadcasts (t) θ^(t) to the participating clients, and each client initializes its local QNN training from this global parameter vector. Each client i∈i holds a private local dataset i=(i,j,yi,j)j=1ni,D_i= \ (x_i,j,y_i,j ) \_j=1^n_i, (5) where ni=|i|n_i=|D_i| is the number of local training samples. The client datasets may be statistically heterogeneous and non-identically distributed. Let (t)⊆S^(t) (6) denote the subset of clients participating in round t. Unless otherwise stated, all subsequent expressions are written for participating clients i∈(t)i ^(t). Figure 2: Overview of DUQFL-Prox. The server broadcasts global QNN parameters to clients, which perform deep-unfolded SPSA optimization with controller-generated hyperparameters, proximal drift control, and best-unfold selection. The selected local models are aggregated at the server, while a meta-objective periodically updates the shared controller using post-aggregation validation performance. I-C Local Quantum Model Each participating client trains a parameterized quantum model represented as a QNN with trainable parameter vector θ. Given an input x, the QNN applies a parameterized quantum circuit, followed by measurement and classical post-processing, to produce either a class-probability distribution or a predicted label. We denote the resulting client-side predictor by f(;).f(x; θ). (7) The empirical local objective at client i is defined as ℒi()=1ni∑(,y)∈iℓ(f(;),y),L_i( θ)= 1n_i _(x,y) _i (f(x; θ),y ), (8) where ℓ(⋅,⋅) (·,·) denotes the task-specific loss function, such as cross-entropy loss for classification. In standard QFL, client i minimizes ℒi()L_i( θ) using a fixed local optimizer initialized from the broadcast global model (t) θ^(t). In contrast, DUQFL-Prox treats local QNN training as a structured unfolded optimization process. The local optimizer is decomposed into a finite sequence of adaptive SPSA-based update blocks, allowing the learning rate and perturbation scale to vary across clients, unfold steps, and communication rounds. The proposed method is decomposed into three functional components. Algorithm 1 describes the client-side unfolded proximal QNN training procedure. Algorithm 2 describes one server-side communication round, including broadcast, local training, aggregation, and meta-loss evaluation. Algorithm 3 describes the outer SPSA-based refinement of the shared controller. This decomposition makes the bilevel structure explicit: the inner level optimizes local QNN parameters, while the outer level adapts the controller to improve post-aggregation global behaviour. I-D Deep-Unfolded Local Optimization Algorithm 1 Client-Side Deep-Unfolded Proximal QNN Training 0: Broadcast global parameters (t) θ^(t), client dataset iD_i, client validation set iV_i, controller ϕ(t) φ^(t), number of unfolds K, proximal coefficient μ 0: Selected local model i,⋆(t) θ_i, ^(t), optimization summaries i,k(t)k=1K\s_i,k^(t)\_k=1^K 1: Initialize local parameters: i,0(t)←(t) θ_i,0^(t)← θ^(t) 2: Initialize optimization summary i,0(t)s_i,0^(t) 3: Initialize Li,⋆(t)←∞L_i, ^(t)←∞ and ki⋆←0k_i ← 0 4: for k=0,1,…,K−1k=0,1,…,K-1 do 5: Construct optimization-state feature vector: i,k(t)←Ψ(k,t,i,k(t),|i|)z_i,k^(t)← (k,t,s_i,k^(t),|D_i| ) 6: Generate unfold-specific SPSA hyperparameters: ηi,k(t),δi,k(t)←Γ(i,k(t);ϕ(t)) _i,k^(t), _i,k^(t)← (z_i,k^(t); φ^(t) ) 7: Define the proximal local objective: ℒiprox()=ℒi()+μ2‖−(t)‖22L_i^prox( θ)=L_i( θ)+ μ2 \| θ- θ^(t) \|_2^2 8: Compute pre-update proximal loss: Li,kbefore←ℒiprox(i,k(t))L_i,k^before _i^prox ( θ_i,k^(t) ) 9: Apply one SPSA-based proximal QNN update: i,k+1(t)←SPSA(i,k(t);ηi,k(t),δi,k(t),ℒiprox) θ_i,k+1^(t) _SPSA ( θ_i,k^(t); _i,k^(t), _i,k^(t),L_i^prox ) 10: Compute post-update proximal loss: Li,kafter←ℒiprox(i,k+1(t))L_i,k^after _i^prox ( θ_i,k+1^(t) ) 11: Compute unfold-step displacement: Δi,k(t)←i,k+1(t)−i,k(t) θ_i,k^(t)← θ_i,k+1^(t)- θ_i,k^(t) 12: Evaluate validation loss: Li,kval←ℒi,val(i,k+1(t);i)L_i,k^val _i,val ( θ_i,k+1^(t);V_i ) 13: if Li,kval<Li,⋆(t)L_i,k^val<L_i, ^(t) then 14: Update best unfolded state: Li,⋆(t)←Li,kval,ki⋆←k+1L_i, ^(t)← L_i,k^val, k_i ← k+1 15: end if 16: Record optimization summary: i,k+1(t)←Summarize(Li,kbefore,Li,kafter,Li,kval,Δi,k(t))s_i,k+1^(t) (L_i,k^before,L_i,k^after,L_i,k^val, θ_i,k^(t) ) 17: end for 18: Select best unfolded local model: i,⋆(t)←i,ki⋆(t) θ_i, ^(t)← θ_i,k_i ^(t) 19: return i,⋆(t) θ_i, ^(t), i,k(t)k=1K\s_i,k^(t)\_k=1^K For each participating client i∈(t)i ^(t), the broadcast global model initializes the local unfolded trajectory, i,0(t)←(t). θ_i,0^(t)← θ^(t). (9) Local training is then unfolded into K SPSA-based update blocks indexed by k∈0,1,…,K−1k∈\0,1,…,K-1\. Each block updates the local QNN parameters according to i,k+1(t)=SPSA(i,k(t);ηi,k(t),δi,k(t),ℒiprox), θ_i,k+1^(t)=U_SPSA ( θ_i,k^(t); _i,k^(t), _i,k^(t),L_i^prox ), (10) where ηi,k(t) _i,k^(t) and δi,k(t) _i,k^(t) are the unfold-specific SPSA learning rate and perturbation scale generated by the controller. Rather than always returning the last unfolded state, the client selects the validation-preferred checkpoint: ki⋆=argmink∈1,…,Kℒi,val(i,k(t)),k_i = _k∈\1,…,K\L_i,val ( θ_i,k^(t) ), (11) and uploads i,⋆(t)=i,ki⋆(t). θ_i, ^(t)= θ_i,k_i ^(t). (12) The intermediate unfolded states remain local to the client and are used only for trajectory construction, validation-based checkpoint selection, and optimization-summary generation. I-E Controller-Driven Hyperparameter Generation The unfolded local optimizer requires step-specific SPSA hyperparameters for each client and unfold step. Instead of using a fixed learning rate and perturbation scale throughout local training, DUQFL-Prox uses a shared meta-controller to generate these quantities adaptively. The controller is parameterized by ϕ(t) φ^(t) and is shared across participating clients at communication round t. At unfold step k, client i constructs an optimization-state feature vector i,k(t)∈ℝd,z_i,k^(t) ^d, (13) which summarizes the current local optimization context. In our implementation, i,k(t)z_i,k^(t) includes normalized unfold progress, normalized communication-round progress, previous loss information, previous parameter displacement, client data fraction, and a client heterogeneity indicator. The controller maps this feature vector to the SPSA learning rate and perturbation scale: (ηi,k(t),δi,k(t))=Γ(i,k(t);ϕ(t)), ( _i,k^(t), _i,k^(t) )= (z_i,k^(t); φ^(t) ), (14) where Γ(⋅) (·) denotes the controller mapping. In this work, Γ(⋅) (·) is implemented as a clipped log-linear controller. Two parameter vectors, ϕη(t) φ_η^(t) and ϕδ(t) φ_δ^(t), generate the logarithmic learning rate and perturbation scale: logηi,k(t) _i,k^(t) =(ϕη(t))⊤i,k(t), = ( φ_η^(t) ) z_i,k^(t), (15) logδi,k(t) _i,k^(t) =(ϕδ(t))⊤i,k(t). = ( φ_δ^(t) ) z_i,k^(t). (16) The resulting values are exponentiated and clipped to predefined feasible intervals: ηi,k(t) _i,k^(t) =clip(exp((ϕη(t))⊤i,k(t)),ηmin,ηmax), =clip ( ( ( φ_η^(t) ) z_i,k^(t) ), _ , _ ), (17) δi,k(t) _i,k^(t) =clip(exp((ϕδ(t))⊤i,k(t)),δmin,δmax). =clip ( ( ( φ_δ^(t) ) z_i,k^(t) ), _ , _ ). (18) This controller provides the adaptive component of DUQFL-Prox by allowing the local SPSA behaviour to vary across clients, unfold steps, and communication rounds. The clipping bounds prevent numerically unstable hyperparameter values, while the proximal objective in Section I-F provides the stabilizing mechanism that restricts excessive client drift. I-F Proximal Drift Control Client drift is a central challenge in federated learning, particularly under non-IID data distributions. In QFL, this issue is further amplified by stochastic SPSA updates, finite-shot measurement effects, and the non-convex loss landscape of variational quantum circuits. A client may reduce its local training loss while producing a parameter update that becomes overly specialized to its own local distribution and less compatible with global aggregation. DUQFL-Prox addresses this issue by introducing a proximal penalty into the local client objective. Let (t) θ^(t) denote the global model broadcast by the server at communication round t. For client i, the proximal local objective is defined as ℒiprox()=ℒi()+μ2‖−(t)‖22,L_i^prox( θ)=L_i( θ)+ μ2 \| θ- θ^(t) \|_2^2, (19) where ℒi()L_i( θ) is the empirical local QNN loss and μ≥0μ≥ 0 controls the strength of proximal regularization. Within each unfolded local optimization step, SPSA minimizes ℒiproxL_i^prox rather than the unregularized local objective. Thus, the proximal unfolded update is written as i,k+1(t)=SPSA(i,k(t);ηi,k(t),δi,k(t),ℒiprox). θ_i,k+1^(t)=U_SPSA ( θ_i,k^(t); _i,k^(t), _i,k^(t),L_i^prox ). (20) The proximal term encourages local models to improve their client-specific objective while remaining close to the current global reference point. This is important in the unfolded setting because multiple adaptive local update blocks can improve flexibility, but may also increase the risk of excessive client movement. DUQFL-Prox therefore combines an adaptive component, provided by the controller-generated SPSA hyperparameters, with a stabilizing component, provided by proximal regularization. To monitor local movement, we record the unfold-step displacement Δi,k(t)=i,k+1(t)−i,k(t), θ_i,k^(t)= θ_i,k+1^(t)- θ_i,k^(t), (21) and its norm ‖Δi,k(t)‖2. \| θ_i,k^(t) \|_2. (22) This quantity is used as a diagnostic indicator of local client movement and can also be included in the controller state features for subsequent unfold steps. Overall, the proximal formulation helps produce adaptive yet aggregation-compatible local QNN updates under heterogeneous federated data. I-G Best-Unfold Model Selection The last unfolded state is not necessarily the best local model to upload. Later unfold steps may continue to reduce local training loss while degrading validation behaviour or increasing local specialization. Therefore, DUQFL-Prox uses validation-based checkpoint selection over the local unfolded trajectory. After generating the candidate states i,k(t)k=1K, \ θ_i,k^(t) \_k=1^K, (23) client i selects the unfold index with the lowest validation loss: ki⋆=argmink∈1,…,Kℒi,val(i,k(t)).k_i = _k∈\1,…,K\L_i,val ( θ_i,k^(t) ). (24) The selected local model is then i,⋆(t)=i,ki⋆(t). θ_i, ^(t)= θ_i,k_i ^(t). (25) Only i,⋆(t) θ_i, ^(t) is uploaded to the server. The remaining unfolded states are retained locally and are used only for trajectory construction, checkpoint selection, and optimization-summary generation. This selection mechanism prevents over-aggressive later unfold steps from dominating the uploaded client update. Algorithm 2 One Federated Round of DUQFL-Prox 0: Client set C, current global model (t) θ^(t), current controller ϕ(t) φ^(t), participating clients (t)S^(t), number of unfolds K, proximal coefficient μ 0: Updated global model (t+1) θ^(t+1), outer meta-loss ℒmeta(t)L_meta^(t) 1: Server broadcasts (t) θ^(t) to all clients in (t)S^(t) 2: for all clients i∈(t)i ^(t) in parallel do 3: Execute Algorithm 1: (i,⋆(t),i,k(t)k=1K)←LocalDUQFLProx((t),i,i,ϕ(t),K,μ) ( θ_i, ^(t),\s_i,k^(t)\_k=1^K ) ( θ^(t),D_i,V_i, φ^(t),K,μ ) 4: end for 5: Compute sample-size aggregation weights: wi(t)=|i|∑j∈(t)|j|w_i^(t)= |D_i| _j ^(t)|D_j| 6: Aggregate selected unfolded local models: (t+1)←∑i∈(t)wi(t)i,⋆(t) θ^(t+1)← _i ^(t)w_i^(t) θ_i, ^(t) 7: Evaluate post-aggregation meta-loss: ℒmeta(t)←ℒvalglobal((t+1))+λfairΩfair(t)+λcommΩcomm(t)+λstabΩstab(t)L_meta^(t) _val^global ( θ^(t+1) )+ _fair _fair^(t)+ _comm _comm^(t)+ _stab _stab^(t) 8: return (t+1) θ^(t+1), ℒmeta(t)L_meta^(t) Algorithm 3 Outer SPSA-Based Controller Update 0: Current controller ϕ(t) φ^(t), current global model (t) θ^(t), participating clients (t)S^(t), outer learning rate αout _out, outer perturbation radius coutc_out 0: Updated controller ϕ(t+1) φ^(t+1) 1: Sample Rademacher perturbation vector ϕ _φ 2: Construct perturbed controllers: ϕ+=ϕ(t)+coutϕ,ϕ−=ϕ(t)−coutϕ φ^+= φ^(t)+c_out _φ, φ^-= φ^(t)-c_out _φ 3: Evaluate one federated round under ϕ+ φ^+ using Algorithm 2 to obtain: ℒmeta+L_meta^+ 4: Evaluate one federated round under ϕ− φ^- using Algorithm 2 to obtain: ℒmeta−L_meta^- 5: Estimate the outer SPSA gradient: ^ϕ(t)←ℒmeta+−ℒmeta−2coutϕ−1 g_φ^(t)← L_meta^+-L_meta^-2c_out _φ^-1 6: Update the controller: ϕ(t+1)←ϕ(t)−αout^ϕ(t) φ^(t+1)← φ^(t)- _out g_φ^(t) 7: return ϕ(t+1) φ^(t+1) I-H Server Aggregation After receiving the selected unfolded local models from participating clients, the server constructs the next global model using sample-size-weighted FedAvg: (t+1)=∑i∈(t)wi(t)i,⋆(t), θ^(t+1)= _i ^(t)w_i^(t) θ_i, ^(t), (26) where wi(t)w_i^(t) denotes the aggregation weight of client i. The weights satisfy wi(t)≥0,∑i∈(t)wi(t)=1.w_i^(t)≥ 0, _i ^(t)w_i^(t)=1. (27) In this work, the sample-size-weighted coefficient is wi(t)=ni∑j∈(t)nj,w_i^(t)= n_i _j ^(t)n_j, (28) where ni=|i|n_i=|D_i| is the number of local training samples available to client i in the current round. Thus, the server does not aggregate every unfolded state. It aggregates only the validation-selected local checkpoint i,⋆(t) θ_i, ^(t) from each participating client. I-I Outer Meta-Objective and Controller Update The controller is not treated as a fixed hyperparameter generator. Instead, it is periodically updated using an outer meta-objective evaluated after server aggregation. Let valD_val denote the validation data used to measure post-aggregation global performance. At communication round t, the outer meta-loss is defined as ℒmeta(t)=ℒvalglobal((t+1))+λfairΩfair(t)+λcommΩcomm(t)+λstabΩstab(t).L_meta^(t)=L_val^global ( θ^(t+1) )+ _fair _fair^(t)+ _comm _comm^(t)+ _stab _stab^(t). (29) Here, ℒvalglobalL_val^global is the validation loss of the aggregated global model, Ωfair(t) _fair^(t) measures client-level performance imbalance, Ωcomm(t) _comm^(t) accounts for communication cost, and Ωstab(t) _stab^(t) penalizes unstable optimization trajectories. The coefficients λfair _fair, λcomm _comm, and λstab _stab control the relative contribution of these terms. To update the controller, DUQFL-Prox applies an outer SPSA step to ϕ(t) φ^(t). A Rademacher perturbation vector ϕ _φ is sampled, and two perturbed controllers are formed: ϕ+=ϕ(t)+coutϕ,ϕ−=ϕ(t)−coutϕ. φ^+= φ^(t)+c_out _φ, φ^-= φ^(t)-c_out _φ. (30) The corresponding perturbed meta-losses, ℒmeta+L_meta^+ and ℒmeta−L_meta^-, are obtained by evaluating virtual federated rounds under ϕ+ φ^+ and ϕ− φ^-, respectively. The outer SPSA gradient estimate is then ^ϕ(t)=ℒmeta+−ℒmeta−2coutϕ−1. g_φ^(t)= L_meta^+-L_meta^-2c_out _φ^-1. (31) Since each perturbation entry satisfies Δϕ,j∈−1,+1 _φ,j∈\-1,+1\, we have Δϕ,j−1=Δϕ,j _φ,j^-1= _φ,j. The controller is updated as ϕ(t+1)=ϕ(t)−αout^ϕ(t), φ^(t+1)= φ^(t)- _out g_φ^(t), (32) where αout _out is the outer learning rate. This establishes a bilevel learning structure. The inner level performs client-side deep-unfolded proximal SPSA optimization of QNN parameters, while the outer level adapts the shared controller so that future local optimization trajectories become more compatible with post-aggregation global performance. The outer update is applied periodically rather than necessarily at every round, which reduces computational overhead while still allowing the controller to improve across training. IV Experimental Setup This section describes the implementation environment, quantum model configuration, datasets, federated partitioning strategy, baseline methods, evaluation metrics, and reproducibility protocol used to evaluate DUQFL-Prox. The objective of the experiments is to assess not only final global accuracy, but also client-level generalization, train-test gap, client fairness, and stability under heterogeneous federated data. IV-A Implementation Details All experiments were implemented in Python using Visual Studio Code as the main development environment. Quantum models were implemented using Qiskit and Qiskit Machine Learning, with Qiskit Aer used for simulator-based experiments. The main Python libraries used were NumPy, pandas, scikit-learn, matplotlib, seaborn, Qiskit, Qiskit Aer, and Qiskit Machine Learning. Fixed random seeds were used for dataset splitting, client partitioning, QNN parameter initialization, and optimizer-related stochasticity wherever supported. The main experimental configuration is summarized in Table I. The reported values correspond to the default configuration used in the main experiments; dataset-specific adjustments were made when necessary due to dataset capacity or computational constraints. TABLE I: Main experimental configuration used in DUQFL-Prox experiments. Setting Value Programming language Python 3.11.14 Development environment Visual Studio Code Quantum framework Qiskit, Qiskit Machine Learning Simulator Qiskit Aer simulator Feature map ZZFeatureMap Ansatz RealAmplitudes Optimizer SPSA / DUQFL-Prox SPSA Number of clients 5 or 10 Communication rounds 10–50 depending on dataset capacity Unfold steps K=5K=5 SPSA iterations per unfold 5–10 Initial learning rate 0.13 Initial perturbation 0.13 Proximal coefficient μ=10−2μ=10^-2 Shots 1024 IV-B Quantum Neural Network Architecture Each client model was implemented as a parameterized QNN. The classical input features were first preprocessed into a low-dimensional representation and then encoded into a quantum circuit using a ZZFeatureMap. A RealAmplitudes ansatz was used as the trainable variational circuit. The number of qubits was determined by the number of retained input features after preprocessing. In the main experiments, we used either two or four QNN input features depending on the dataset and computational budget. Let ∈ℝdx ^d denote the preprocessed input vector. The QNN maps x to a quantum state through the feature map and then applies a trainable ansatz parameterized by θ. Measurement outcomes are classically post-processed to obtain class probabilities or predicted labels. All compared methods used the same QNN architecture for a given dataset to ensure a fair comparison. IV-C Datasets and Preprocessing The proposed DUQFL-Prox framework was evaluated on datasets from genomics and financial fraud detection. Since current QNN models are constrained by the number of available qubits and circuit depth, each dataset was transformed into a low-dimensional QNN-compatible representation before federated training. Table I summarizes the dataset usage, preprocessing pipeline, and final QNN input dimension. TABLE I: Datasets used for evaluating DUQFL-Prox across multiple domains. Dataset Domain Task Source QNN features BAF Finance Bank fraud detection [11] 4 Genome Genomics Genomic sequencing [4] 2 or 4 The BAF dataset is taken from the Bank Account Fraud Dataset Suite introduced by Jesus et al. [11]. The dataset suite was designed as a privacy-preserving, large-scale tabular benchmark for bank account-opening fraud detection, with realistic challenges including temporal dynamics, severe class imbalance, and bias/fairness-related distributional shifts. The Genome experiments use the DemoHumanOrWorm task from the Genomic Benchmarks suite [4]. Genomic Benchmarks provides curated datasets for genomic sequence classification and offers standardized access through common machine-learning and deep-learning interfaces. IV-C1 BAF Dataset For the BAF, Bank Account fraud-detection experiment, the target variable was fraud_bool. After removing missing values, the dataset was optionally subsampled using stratified sampling to preserve the class distribution under the QNN computational budget. The data were split into training, validation, and test sets before feature encoding and scaling to avoid data leakage. Categorical variables were encoded using ordinal encoding with support for unseen categories, while numerical variables were standardized. PCA was then applied to obtain a low-dimensional representation compatible with the number of QNN qubits. Finally, the PCA features were optionally scaled to the quantum angle range [0,π][0,π]. In the reported BAF run, the resulting split contained 2,999 training samples, 750 validation samples, and 1,250 test samples, using four QNN features/qubits. IV-C2 Genome Dataset For the Genome experiment, we used the DemoHumanOrWorm task from the Genomic Benchmarks suite. Each DNA sequence was converted into a numerical vector using a fixed word-size encoding strategy. Specifically, all unique words of length word_size were assigned integer identifiers, and each sequence was represented by the corresponding sequence of integer word indices. The resulting vectors were shuffled and scaled using MinMax scaling to obtain bounded QNN-compatible input features. In the implementation used in this work, 10,000 processed records were used for training and 2,000 records were used for testing. The BAF and Genome datasets are used as representative intelligent-service tasks because they capture two unique privacy-sensitive distributed decision setting challenges: financial fraud detection and genomic classification. V Experimental Results and Analysis (a) Global accuracy (b) Mean client test accuracy (c) Train–test gap (d) Client fairness gap Figure 3: Federated performance comparison on the BAF dataset. DUQFL-Prox improves final global accuracy and mean client test accuracy while substantially reducing the train–test gap and client fairness gap. Lower values are preferable for train–test gap and fairness gap. (a) Recall (b) precision (c) F1 (d) Specificity Figure 4: Epoch-wise ROC-AUC, PR-AUC, MCC, and specificity trajectories on the BAF dataset. DUQFL-Prox shows more stable late-epoch behaviour on test ROC-AUC, PR-AUC, and MCC, while Adam-QFL exhibits a recall-heavy operating regime with declining specificity. (a) rocauc (b) pr auc (c) mcc (d) Client Fairness Figure 5: BAF Data We evaluate DUQFL-Prox against representative QFL baselines and ablation variants. The analysis is organized around four questions: (i) whether DUQFL-Prox improves global and client-level performance, (i) whether the proposed method reduces train–test generalization gap and client fairness gap, (i) how DUQFL-Prox behaves under severe class imbalance, and (iv) whether the learned global QNN checkpoints remain executable on real IBM quantum hardware. We report both global and client-level metrics because global accuracy alone is insufficient to evaluate heterogeneous QFL. A method may obtain strong global accuracy while still producing unstable or imbalanced performance across clients. Therefore, in addition to global accuracy, we analyse mean client test accuracy, train–test gap, client fairness gap, precision, recall, F1-score, ROC-AUC, PR-AUC, MCC, and specificity. V-A Results on the BAF Dataset The BAF dataset represents a highly imbalanced financial fraud-detection task. This setting is challenging for QFL because the minority class is rare and the client partitions are non-IID. Consequently, high global accuracy alone does not necessarily imply good fraud-detection behaviour. We therefore evaluate BAF using both federated learning metrics and imbalance-aware classification metrics. Figure 3 compares DUQFL-Prox with Default-QFL and FedProx-QFL using global accuracy, mean client test accuracy, train–test gap, and client fairness gap. DUQFL-Prox achieves the strongest final global accuracy, reaching approximately 0.65040.6504, compared with approximately 0.53440.5344 for FedProx-QFL and 0.45920.4592 for Default-QFL. Although Default-QFL reaches a temporary peak during intermediate rounds, its late-round performance decreases substantially, indicating unstable post-aggregation behaviour under the non-IID BAF setting. The client-level results provide a clearer indication of the benefit of DUQFL-Prox. The proposed method obtains the highest final mean client test accuracy, approximately 0.64360.6436, while FedProx-QFL and Default-QFL obtain approximately 0.54800.5480 and 0.50220.5022, respectively. This suggests that DUQFL-Prox improves not only the aggregated global model, but also the generalization behaviour observed across distributed clients. The train–test gap further supports this conclusion. Default-QFL and FedProx-QFL show relatively large final train–test gaps, approximately 0.28540.2854 and 0.21160.2116, respectively. In contrast, DUQFL-Prox obtains a near-zero final gap. Since lower train–test gap is preferable, this indicates that DUQFL-Prox reduces local over-specialization and improves generalization under heterogeneous client data. The client fairness gap, measured as the P90−P10P90-P10 spread of client test accuracies, is also lowest for DUQFL-Prox. The final fairness gap is approximately 0.01930.0193 for DUQFL-Prox, compared with 0.09580.0958 for FedProx-QFL and 0.19100.1910 for Default-QFL. This result is important because a high global accuracy can hide poor performance on difficult or minority clients. The low fairness gap shows that DUQFL-Prox produces more balanced client-level performance. V-B Imbalance-Aware Classification Behaviour on BAF Because the BAF dataset is severely class-imbalanced, we further analyse precision, recall, F1-score, specificity, ROC-AUC, PR-AUC, and MCC on validation and test splits. In Figures 4 and 5, the prefixes F, D, and A denote FedAvg-default, DUQFL-Prox, and FedAvg-tuned-Adam, respectively, while Val and T denote validation and test splits. The results reveal different operating behaviours. FedAvg-tuned-Adam achieves high recall for much of training, indicating aggressive minority-class detection. However, this behaviour is accompanied by a decline in specificity, which suggests a larger number of false positives. Thus, Adam-QFL behaves as a recall-oriented baseline but provides a less selective operating point. DUQFL-Prox shows a more balanced late-epoch behaviour. It provides more consistent improvements in test F1, MCC, ROC-AUC, precision, and specificity. This suggests that DUQFL-Prox does not simply increase minority-class detection at the expense of false positives; instead, it maintains a more balanced recall–specificity trade-off. FedAvg-default exhibits stronger fluctuations across epochs, particularly for MCC, F1-score, and ROC-AUC, indicating weaker stability under severe class imbalance. The validation curves are noisier than the test curves, especially for PR-AUC, MCC, precision, and recall. This is expected because the validation split contains very few positive fraud samples. Therefore, epoch-wise trajectories are important in addition to checkpoint-based summaries, since they reveal the underlying training dynamics and the recall–specificity trade-off among the methods. V-C Results on the Genome Dataset The Genome dataset presents a more nuanced comparison. As shown in Figure 6, FedProx-QFL achieves the highest final global accuracy, reaching approximately 0.850.85. DUQFL-Prox remains competitive, with final global accuracy around 0.820.82–0.830.83, while Default-QFL ends with substantially lower global accuracy. This result shows that DUQFL-Prox is not always the single best method in terms of final global accuracy. However, the client-level metrics show a different and important trend. DUQFL-Prox achieves the highest final mean client test accuracy, approximately 0.800.80, outperforming the other compared methods. This indicates that DUQFL-Prox provides stronger generalization across distributed clients, even when another baseline obtains slightly higher final global accuracy. The train–test gap and client fairness gap further support this interpretation. DUQFL-Prox obtains the lowest final train–test gap, close to zero, suggesting reduced overfitting and improved client-level generalization. It also achieves the lowest final client fairness gap, indicating that its performance is more balanced across clients. DUQFL-best and DUQFL-drift also improve over Default-QFL, but DUQFL-Prox provides the strongest overall stability and fairness profile. Therefore, the Genome experiment supports a more federated-learning-relevant conclusion: DUQFL-Prox is not merely an accuracy-maximizing method, but a stability-aware and generalization-aware QFL method. It provides the best trade-off among global accuracy, mean client test accuracy, train–test gap, and fairness under heterogeneous client data. (a) Global accuracy (b) Mean client test accuracy (c) Train–test gap (d) Client fairness gap Figure 6: Federated performance comparison on the Genome dataset. FedProx-QFL achieves the strongest final global accuracy, while DUQFL-Prox provides the best client-level generalization, lowest train–test gap, and lowest fairness gap. TABLE IV: Qualitative comparison of QFL and DUQFL methods across Genome and BAF datasets. Here, ✓✓ denotes the strongest performance, ✓ denotes competitive or moderate performance, and −- denotes weak, unstable, or less consistent performance. For train–test gap, client drift/regularization, and fairness/balance, lower values are preferable. Dataset Method Final global accuracy Peak global accuracy Mean client test accuracy Low train–test gap Low client drift / regularization Client fairness / balance Genome Adam-QFL −- ✓✓ −- −- −- −- DUQFL-Prox ✓ −- ✓✓ ✓✓ ✓✓ ✓✓ Default-QFL −- −- −- −- −- −- FedProx-QFL ✓✓ ✓ ✓ ✓ ✓ ✓ BAF Adam-QFL −- −- −- −- −- −- DUQFL-Prox ✓✓ ✓✓ ✓✓ ✓✓ ✓✓ ✓✓ Default-QFL −- ✓ −- −- −- −- FedProx-QFL ✓ −- ✓ ✓ ✓ ✓ V-D Cross-Dataset Comparison Table IV summarizes the qualitative behaviour of the compared methods across the Genome and BAF datasets. The results show that DUQFL-Prox provides the most consistent stability–generalization trade-off across datasets. On BAF, DUQFL-Prox achieves the strongest performance across final global accuracy, mean client test accuracy, train–test gap, and client fairness gap. This indicates that proximal deep-unfolded QNN optimization is particularly effective under severe class imbalance and non-IID client partitions. On Genome, FedProx-QFL achieves the highest final global accuracy, whereas DUQFL-Prox achieves the strongest client-level generalization, lowest train–test gap, and lowest fairness gap. This distinction is important because a single global test metric does not fully characterize federated performance. The Genome results show that DUQFL-Prox improves the reliability and balance of client-level performance even when another method slightly improves the final centralized global accuracy. Overall, the cross-dataset results indicate that DUQFL-Prox should be interpreted as a stability-aware and generalization-aware QFL method. Its advantage lies in improving the quality and aggregation compatibility of local QNN updates through deep-unfolded optimization, validation-based checkpoint selection, and proximal drift control. (a) DUQFL-Prox (b) Default-QFL Figure 7: Simulator and real IBM quantum hardware validation of trained global QFL checkpoints. DUQFL-Prox shows a more stable hardware trajectory than Default-QFL across saved global parameter checkpoints. V-E Controller Behaviour and Hyperparameter Adaptation Figure 8: Controller behaviour and hyperparameter adaptation in DUQFL-Prox on the BAF dataset. The trace logs show the controller-generated SPSA learning rate η, perturbation scale δ, local loss reduction Δℒ , and selected best-unfold checkpoint k⋆k . These results verify that DUQFL-Prox performs adaptive local QNN optimization rather than using a fixed SPSA schedule. Figure 8 analyses the controller trace logs of DUQFL-Prox. The generated learning rate η and perturbation scale δ remain within stable numerical ranges while varying across the unfolded training process. The local loss-reduction curve shows that the unfolded SPSA blocks produce measurable improvement in the client objective, particularly in the early unfold steps. The selected-checkpoint distribution further shows that the last unfolded state is not always uploaded, supporting the use of validation-based best-unfold selection. Together, these results provide empirical evidence that DUQFL-Prox uses controller-guided adaptive optimization rather than a fixed local SPSA schedule. V-F Real IBM Quantum Hardware Validation To examine practical deployability, we evaluated selected trained global QNN checkpoints on IBM Quantum hardware. Due to the high queue time and execution cost of repeatedly training the full federated process on real quantum devices, hardware execution was used for post-training validation rather than full hardware-in-the-loop federated training. Specifically, saved global parameter checkpoints were assigned to the trained QNN circuit, transpiled for the selected IBM backend, and executed using finite-shot measurement. The resulting hardware accuracy was compared with noiseless simulator, shot-based Aer simulator, and backend-inspired noisy Aer simulator results. Figure 7 compares DUQFL-Prox and Default-QFL across saved global checkpoints. DUQFL-Prox shows a more stable and improving hardware trajectory than Default-QFL and remains reasonably aligned with the simulator curves in later checkpoints. The difference between simulator and hardware results is expected due to finite-shot uncertainty, device noise, calibration drift, transpilation constraints, and backend-specific routing overhead. These results should not be interpreted as full federated training on IBM hardware. Rather, they provide post-training hardware feasibility evidence, showing that the learned DUQFL-Prox global checkpoints remain executable on real quantum hardware and retain a more stable trajectory than the Default-QFL baseline. VI Conclusion The proposed DUQFL-Prox framework is supported by three theoretical observations. First, the optional step-projection mechanism bounds each unfolded local update and therefore bounds cumulative client drift across the unfolded trajectory. Second, the proximal regularization term acts as a drift-control mechanism by penalizing deviation from the broadcast global model, balancing local adaptation with aggregation compatibility. Third, under standard smoothness assumptions on the post-aggregation meta-objective, the outer SPSA update provides a zeroth-order controller-update direction that supports descent-oriented adaptation in expectation. Our results support the use of DUQFL-Prox as a drift-stable QFL‘ framework for intelligent services that require privacy preservation, fairness, and reliable generalization across heterogeneous clients. These results do not constitute a full convergence guaranty for arbitrary nonconvex QFL systems; instead, they characterize the stability, drift-control, and meta-optimization behavior induced by DUQFL-Prox. Detailed statements and proofs are provided in Appendix A. References [1] M. Chehimi, S. Y.-C. Chen, W. Saad, D. Towsley, and M. Debbah (2023) Foundations of quantum federated learning over classical and quantum networks. IEEE Network. Cited by: TABLE I, §I. [2] Y. J. Cho, J. Wang, and G. Joshi (2022) Towards understanding biased client selection in federated learning. In International Conference on Artificial Intelligence and Statistics, p. 10351–10375. Cited by: §I. [3] C. Chu, L. Jiang, and F. Chen (2023) Cryptoqfl: quantum federated learning on encrypted data. In Proc. IEEE International Conference on Quantum Computing and Engineering (QCE), Vol. 1, p. 1231–1237. Cited by: TABLE I, §I. [4] K. Grešová, V. Martinek, D. Čechák, P. Šimeček, and P. Alexiou (2023) Genomic benchmarks: a collection of datasets for genomic sequence classification. BMC Genomic Data 24 (1), p. 25. Cited by: §IV-C, TABLE I. [5] A. Hamdi et al. (2025) Drone-as-a-service: research challenges and directions. Proceedings of the IEEE 113 (5), p. 416–442. External Links: Document Cited by: §I. [6] Y. F. Hanna, A. A. Khater, M. El-Bardini, and A. M. El-Nagar (2023) Real time adaptive pid controller based on quantum neural network for nonlinear systems. Engineering Applications of Artificial Intelligence 126, p. 106952. Cited by: TABLE I, §I. [7] R. Huang, X. Tan, and Q. Xu (2022) Quantum federated learning with decentralized data. IEEE Journal of Selected Topics in Quantum Electronics 28 (4), p. 1–10. Cited by: §I, TABLE I, §I. [8] IBM Quantum SPSA Optimizer. Note: https://quantum.cloud.ibm.com/docs/api/qiskit/qiskit_algorithms.optimizers.SPSAAccessed: 2026-05-17 Cited by: §I. [9] N. Innan, A. Marchisio, M. Bennai, and M. Shafique (2025) Qfnn-ffd: quantum federated neural network for financial fraud detection. In 2025 IEEE International Conference on Quantum Software (QSW), p. 41–47. Cited by: TABLE I. [10] S. Jerbi, C. Gyurik, S. C. Marshall, R. Molteni, and V. Dunjko (2024) Shadows of quantum machine learning. Nature Communications 15 (1), p. 5676. Cited by: §I. [11] S. Jesus, J. Pombal, D. Alves, A. Cruz, P. Saleiro, R. Ribeiro, J. Gama, and P. Bizarro (2022) Turning the tables: biased, imbalanced, dynamic tabular datasets for ml evaluation. Advances in Neural Information Processing Systems 35, p. 33563–33575. Cited by: §IV-C, TABLE I. [12] P. Kairouz and H. B. McMahan (2021) Advances and open problems in federated learning. Foundations and trends in machine learning 14 (1-2), p. 1–210. Cited by: §I, §I-A. [13] S. P. Karimireddy, S. Kale, M. Mohri, S. Reddi, S. Stich, and A. T. Suresh (2020) Scaffold: stochastic controlled averaging for federated learning. p. 5132–5143. Cited by: §I. [14] D. P. Kingma and J. Ba (2014) Adam: a method for stochastic optimization. arXiv preprint arXiv:1412.6980. Cited by: §I-A. [15] F. Lai, X. Zhu, H. V. Madhyastha, and M. Chowdhury (2021) Oort: efficient federated learning via guided participant selection. In 15th \USENIX\ Symposium on Operating Systems Design and Implementation (\OSDI\ 21), p. 19–35. Cited by: §I. [16] T. Li, A. K. Sahu, M. Zaheer, M. Sanjabi, A. Talwalkar, and V. Smith (2020) Federated optimization in heterogeneous networks. Proceedings of Machine Learning and Systems 2, p. 429–450. Cited by: §I, §I, §I, §I-A. [17] W. Li, S. Lu, and D. Deng (2021) Quantum federated learning through blind quantum computing. Science China Physics, Mechanics & Astronomy 64 (10), p. 100312. Cited by: TABLE I, §I. [18] S. Majumder, S. Ray, M. Dasgupta, P. Bhattacharya, T. R. Gadekallu, and G. Srivastava (2026) QuanFraud: quantum state verification scheme for fraud detection in iot-assisted quantum-blockchain networks. IEEE Transactions on Services Computing 19 (1), p. 616–627. External Links: Document Cited by: TABLE I. [19] A. Nakai-Kasai and T. Wadayama (2024) Deep unfolding-based weighted averaging for federated learning under device and statistical heterogeneous environments. IEICE Transactions on Communications 108 (4), p. 411–420. Cited by: Appendix D, TABLE I. [20] S. I. Nanayakkara and S. R. Pokhrel (2025) New insights on unfolding and fine-tuning quantum federated learning. arXiv preprint arXiv:2506.20016. Cited by: TABLE I. [21] B. Narottama and S. Y. Shin (2023) Federated quantum neural network with quantum teleportation for resource optimization in future wireless communication. IEEE Transactions on Vehicular Technology. Cited by: TABLE I. [22] A. G. Neiat, A. Bouguettaya, and M. Bahutair (2022) A deep reinforcement learning approach for composing moving iot services. IEEE Transactions on Services Computing 15 (5), p. 2538–2550. External Links: Document Cited by: §I. [23] D. C. Nguyen, M. R. Uddin, S. Shaon, R. Rahman, O. Dobre, and D. Niyato (2025) Quantum federated learning: a comprehensive survey. arXiv preprint arXiv:2508.15998. Cited by: §I. [24] D. Ning, Y. Ge, E. Bertino, Z. Zheng, Y. Jiang, and H. Wang (2026) AdpFL: a privacy-preserving federated learning framework through adaptive model pruning on non-iid data. IEEE Transactions on Services Computing (), p. 1–15. External Links: Document Cited by: TABLE I. [25] S. R. Pokhrel, N. Yash, J. Kua, G. Li, and L. Pan (2024) Quantum federated learning experiments in the cloud with data encoding. arXiv preprint arXiv:2405.00909. Cited by: TABLE I, §I. [26] J. Qi, X. Zhang, and J. Tejedor (2023) Optimizing quantum federated learning based on federated quantum natural gradient descent. In ICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), p. 1–5. Cited by: TABLE I, §I. [27] Qiskit Community (2025) SPSA - qiskit algorithms documentation. Note: Accessed: 2026-05-10https://qiskit-community.github.io/qiskit-algorithms/stubs/qiskit_algorithms.optimizers.SPSA.html Cited by: §I. [28] A. Reisizadeh, I. Tziotis, H. Hassani, A. Mokhtari, and R. Pedarsani (2022) Straggler-resilient federated learning: leveraging the interplay between statistical accuracy and system heterogeneity. IEEE Journal on Selected Areas in Information Theory 3 (2), p. 197–205. Cited by: §I, §I-A. [29] K. Selialia, Y. Chandio, and F. M. Anwar (2024) Mitigating group bias in federated learning for heterogeneous devices. In Proceedings of the 2024 ACM conference on fairness, accountability, and transparency, p. 1043–1054. Cited by: §I. [30] N. Shlezinger, S. Segarra, Y. Zhang, D. Avrahami, Z. Davidov, T. Routtenberg, and Y. C. Eldar (2025) Deep unfolding: recent developments, theory, and design guidelines. arXiv preprint arXiv:2512.03768. Cited by: §I. [31] J. C. Spall (1998) An overview of the simultaneous perturbation method for efficient optimization. Johns Hopkins apl technical digest 19 (4), p. 482–492. Cited by: §I. [32] J. C. Spall (2002) Multivariate stochastic approximation using a simultaneous perturbation gradient approximation. IEEE transactions on automatic control 37 (3), p. 332–341. Cited by: §I. [33] Q. Xia and Q. Li (2021) Quantumfed: a federated learning framework for collaborative quantum training. In 2021 IEEE Global Communications Conference (GLOBECOM), p. 1–6. Cited by: TABLE I, §I. [34] J. Xu et al. (2024) A holistic and hybrid service selection strategy for mec-based uav last-mile delivery systems. IEEE Transactions on Services Computing 17 (6), p. 3022–3036. External Links: Document Cited by: §I. [35] W. Yamany, N. Moustafa, and B. Turnbull (2021) OQFL: an optimized quantum-based federated learning framework for defending against adversarial attacks in intelligent transportation systems. IEEE Transactions on Intelligent Transportation Systems. Cited by: TABLE I. [36] K. Yu, X. Zhang, Z. Ye, G.-D. Guo, and S. Lin (2022) Quantum federated learning based on gradient descent. arXiv preprint arXiv:2212.12913. Cited by: §I. [37] W. J. Yun, J. P. Kim, H. Baek, S. Jung, J. Park, M. Bennis, and J. Kim (2022) Quantum federated learning with entanglement controlled circuits and superposition coding. arXiv preprint arXiv:2212.01732. Cited by: TABLE I, §I. [38] X. Zhang, J. Zhao, C. Xu, H. Wang, and Y. Zhang (2022) DOPIV: post-quantum secure identity-based data outsourcing with public integrity verification in cloud storage. IEEE Transactions on Services Computing 15 (1), p. 334–345. External Links: Document Cited by: TABLE I. [39] H. Zhao (2023) Non-iid quantum federated learning with one-shot communication complexity. Quantum Machine Intelligence 5 (1), p. 3. Cited by: TABLE I, §I. . Table V summarizes the main notation used throughout the paper. TABLE V: Summary of notation used in DUQFL-Prox. Notation Description =1,…,NC=\1,…,N\ Set of all clients N Number of clients t∈0,…,T−1t∈\0,…,T-1\ Communication round index T Number of communication rounds (t)⊆S^(t) Participating client set at round t iD_i Local training dataset of client i ni=|i|n_i=|D_i| Number of local training samples at client i iV_i Local validation set of client i valD_val Global or held-out validation data used for meta-objective evaluation (t) θ^(t) Global QNN parameter vector at round t i,k(t) θ_i,k^(t) Local QNN parameter vector of client i at unfold step k of round t i,⋆(t) θ_i, ^(t) Selected best unfolded local model of client i at round t P Number of trainable QNN parameters K Number of unfolded local optimization steps k Unfold-step index f(;)f(x; θ) QNN prediction function ℒi()L_i( θ) Empirical local loss of client i ℒiprox()L_i^prox( θ) Proximal local objective of client i ηi,k(t) _i,k^(t) SPSA learning rate generated for client i at unfold step k δi,k(t) _i,k^(t) SPSA perturbation scale generated for client i at unfold step k i,k(t)z_i,k^(t) Optimization-state feature vector used by the controller ϕ(t) φ^(t) Shared controller parameter vector at round t ϕη(t) φ_η^(t) Controller parameters used to generate ηi,k(t) _i,k^(t) ϕδ(t) φ_δ^(t) Controller parameters used to generate δi,k(t) _i,k^(t) μ Proximal regularization coefficient wi(t)w_i^(t) Aggregation weight of client i at round t ℒmeta(t)L_meta^(t) Outer meta-loss at round t αout _out Outer controller learning rate coutc_out Outer SPSA perturbation radius ϕ _φ Rademacher perturbation vector for outer SPSA Appendix A Proximal Objective and Drift Control This appendix clarifies the drift-control role of the proximal component in DUQFL-Prox. In heterogeneous QFL, each client initializes local training from the broadcast global model (t) θ^(t). Under non-IID data and stochastic QNN optimization, repeated local updates may move the client model far from this global reference point, producing updates that are less compatible with server aggregation. DUQFL-Prox therefore replaces the unregularized local objective ℒi()L_i( θ) with ℒiprox()=ℒi()+μ2‖−(t)‖22,L_i^prox( θ)=L_i( θ)+ μ2 \| θ- θ^(t) \|_2^2, (33) where μ≥0μ≥ 0 controls the strength of proximal regularization. The first term promotes local empirical improvement, whereas the second term penalizes excessive deviation from the broadcast global model. Formally, the gradient of the proximal objective is ∇ℒiprox()=∇ℒi()+μ(−(t)). _i^prox( θ)= _i( θ)+μ ( θ- θ^(t) ). (34) Although DUQFL-Prox uses SPSA rather than analytic gradients, this expression shows that the proximal term introduces a drift-correcting component toward the current global model. Thus, local training is discouraged from moving arbitrarily far from (t) θ^(t). At unfold step k, SPSA is applied to the proximal objective: i,k+1(t)=SPSA(i,k(t);ηi,k(t),δi,k(t),ℒiprox). θ_i,k+1^(t)=U_SPSA ( θ_i,k^(t); _i,k^(t), _i,k^(t),L_i^prox ). (35) Therefore, the local update is influenced by both client-specific loss reduction and proximity to the broadcast global parameters. The unfold-step displacement is recorded as Δi,k(t)=i,k+1(t)−i,k(t), θ_i,k^(t)= θ_i,k+1^(t)- θ_i,k^(t), (36) with magnitude ‖Δi,k(t)‖2. \| θ_i,k^(t) \|_2. (37) This quantity provides a diagnostic measure of local movement and can be used as an optimization-state feature in subsequent unfold steps. Overall, the proximal term does not prevent local learning. Instead, it balances local improvement with aggregation compatibility. This is particularly important in DUQFL-Prox because deep unfolding increases the flexibility of local optimization, which can otherwise increase the risk of client over-specialization under heterogeneous data. Appendix B Outer SPSA Gradient Estimation The shared controller in DUQFL-Prox is parameterized by ϕ(t) φ^(t) and generates the SPSA learning rate and perturbation scale used during unfolded local QNN optimization. The purpose of the outer update is to refine ϕ(t) φ^(t) so that the induced local optimization trajectories improve post-aggregation global performance. At communication round t, the outer meta-loss is ℒmeta(t)=ℒvalglobal((t+1))+λfairΩfair(t)+λcommΩcomm(t)+λstabΩstab(t).L_meta^(t)=L_val^global ( θ^(t+1) )+ _fair _fair^(t)+ _comm _comm^(t)+ _stab _stab^(t). (38) The dependence of ℒmeta(t)L_meta^(t) on ϕ(t) φ^(t) is indirect: the controller determines (η,δ)(η,δ), which affects local QNN training, server aggregation, and finally the global validation loss. Since this mapping is expensive and not easily differentiable, DUQFL-Prox applies an outer SPSA update. A Rademacher perturbation vector ϕ∈−1,+1dim(ϕ) _φ∈\-1,+1\ ( φ) is sampled, and two perturbed controllers are formed: ϕ+=ϕ(t)+coutϕ,ϕ−=ϕ(t)−coutϕ. φ^+= φ^(t)+c_out _φ, φ^-= φ^(t)-c_out _φ. (39) Evaluating the federated process under these two controllers gives ℒmeta+L_meta^+ and ℒmeta−L_meta^-. For the j-th controller parameter, the SPSA finite-difference estimate is g^ϕ,j(t)=ℒmeta+−ℒmeta−2coutΔϕ,j. g_φ,j^(t)= L_meta^+-L_meta^-2c_out _φ,j. (40) Equivalently, in vector form, ^ϕ(t)=ℒmeta+−ℒmeta−2coutϕ−1. g_φ^(t)= L_meta^+-L_meta^-2c_out _φ^-1. (41) Since Δϕ,j∈−1,+1 _φ,j∈\-1,+1\, we have Δϕ,j−1=Δϕ,j _φ,j^-1= _φ,j. The controller is then updated as ϕ(t+1)=ϕ(t)−αout^ϕ(t), φ^(t+1)= φ^(t)- _out g_φ^(t), (42) where αout _out is the outer learning rate. This update does not directly overwrite client-local QNN models. Instead, it modifies the controller mapping Γ(i,k(t);ϕ(t))↦(ηi,k(t),δi,k(t)), (z_i,k^(t); φ^(t) ) ( _i,k^(t), _i,k^(t) ), (43) so that future local SPSA behaviour becomes more aligned with the post-aggregation global objective. Appendix C Stability and Convergence Discussion This appendix provides a stability-oriented interpretation of DUQFL-Prox. We do not claim a general convergence theorem for arbitrary non-convex QNN objectives, since variational quantum loss landscapes may be non-convex and affected by finite-shot measurement noise. Instead, we clarify how the proposed components are designed to improve optimization stability under heterogeneous federated data. Let the global objective be written as ℱ()=∑i=1Npiℒi(),pi=ni∑j=1Nnj.F( θ)= _i=1^Np_iL_i( θ), p_i= n_i _j=1^Nn_j. (44) Here, ℒi()L_i( θ) denotes the local objective of client i, and pip_i is its sample-size proportion. Under non-IID data, the local objectives may induce different descent directions. Therefore, a client model that improves its own local objective may still be poorly aligned with the global objective after aggregation. This mismatch is the source of client drift. DUQFL-Prox addresses this instability through three complementary mechanisms. First, the proximal penalty μ2‖−(t)‖22 μ2 \| θ- θ^(t) \|_2^2 (45) discourages excessive deviation from the broadcast global model (t) θ^(t). This encourages local updates to remain within a controlled neighbourhood of the global reference point, reducing the risk of over-specialized client updates. Second, validation-based best-unfold selection prevents the method from always uploading the final unfolded state. Instead, client i uploads i,⋆(t)=i,ki⋆(t),ki⋆=argmink∈1,…,Kℒi,val(i,k(t)). θ_i, ^(t)= θ_i,k_i ^(t), k_i = _k∈\1,…,K\L_i,val ( θ_i,k^(t) ). (46) This mechanism is important because later unfold steps may reduce training loss while increasing validation loss or local specialization. Third, the outer meta-objective evaluates the consequence of local optimization after aggregation: ℒmeta(t)=ℒvalglobal((t+1))+λfairΩfair(t)+λcommΩcomm(t)+λstabΩstab(t).L_meta^(t)=L_val^global ( θ^(t+1) )+ _fair _fair^(t)+ _comm _comm^(t)+ _stab _stab^(t). (47) Thus, the controller is guided by post-aggregation global behaviour rather than local training loss alone. The proximal coefficient μ controls the trade-off between local adaptation and aggregation compatibility. A larger μ more strongly penalizes movement away from (t) θ^(t), whereas a smaller μ gives clients greater freedom to adapt to their local data: local adaptivity↔global aggregation compatibility.local adaptivity aggregation compatibility. (48) Therefore, DUQFL-Prox should be interpreted as a stability-aware optimization framework rather than a method that guarantees global convergence for all QNN loss landscapes. Its purpose is to reduce harmful client drift, select more generalizable local unfolded checkpoints, and align the controller with post-aggregation validation performance. This interpretation is consistent with our experimental protocol, which reports global accuracy together with mean client test accuracy, train–test gap, fairness gap, and classification metrics. Appendix D Relation to Classical Deep-Unfolded Federated Learning Classical deep-unfolded federated learning and DUQFL-Prox share the same high-level principle: an iterative federated optimization process can be unfolded into structured steps, and selected components of this process can be made learnable. However, the level at which learning is introduced is different. In classical deep-unfolded weighted averaging for FL, the unfolded process is typically applied at the server aggregation level [19]. Standard FedAvg computes (t+1)=∑i=1Nni∑j=1Nnji(t),w^(t+1)= _i=1^N n_i _j=1^Nn_jw_i^(t), (49) where i(t)w_i^(t) denotes the local model from client i and nin_i is the corresponding local sample size. Deep-unfolded weighted averaging replaces these fixed aggregation weights with learnable coefficients: (t+1)=∑i=1NΘi(t)i(t),w^(t+1)= _i=1^N _i^(t)w_i^(t), (50) where Θi(t) _i^(t) determines the contribution of client i to the global model at round t. Thus, classical deep-unfolded FL primarily asks how client updates should be weighted during aggregation. DUQFL-Prox addresses a different question: how each client should optimize its local QNN before aggregation. Instead of learning aggregation weights Θi(t) _i^(t), DUQFL-Prox learns a controller parameterized by ϕ(t) φ^(t). This controller generates unfold-specific SPSA hyperparameters: (ηi,k(t),δi,k(t))=Γ(i,k(t);ϕ(t)). ( _i,k^(t), _i,k^(t) )= (z_i,k^(t); φ^(t) ). (51) The server aggregation remains sample-size-weighted: (t+1)=∑i∈(t)wi(t)i,⋆(t). θ^(t+1)= _i ^(t)w_i^(t) θ_i, ^(t). (52) The key difference is that the uploaded model i,⋆(t) θ_i, ^(t) is produced through controller-guided deep-unfolded proximal QNN optimization and validation-based best-unfold selection. Therefore, DUQFL-Prox transfers the deep-unfolding principle from server-side aggregation-weight learning to client-side quantum optimizer-policy learning. This distinction is important because QFL introduces optimization challenges not present in ordinary classical FL, including finite-shot measurement noise, SPSA perturbation sensitivity, and non-convex variational quantum loss landscapes. DUQFL-Prox is therefore complementary to classical deep-unfolded FL methods: a future extension could jointly learn client-side QNN optimization policies and server-side aggregation weights, whereas this work focuses on stabilizing local QNN optimization before aggregation. Appendix E Additional Ablation and Controller Trace Analysis (a) Global accuracy (b) Mean client test accuracy (c) Train–test gap (d) Client fairness gap Figure 9: Ablation analysis of DUQFL variants on the Genome non-IID setting. DUQFL-last uploads the final unfolded local state, DUQFL-best uploads the validation-preferred unfolded checkpoint, DUQFL-drift adds explicit displacement control, and DUQFL-Prox incorporates proximal drift regularization. DUQFL-Prox achieves the strongest client-level generalization, lowest train–test gap, and smallest fairness gap, indicating that best-unfold selection and proximal drift control improve stability and balance under heterogeneous QFL. E-A Compared Methods We compare DUQFL-Prox with representative QFL baselines and ablation variants. All methods use the same QNN architecture, dataset split, client partitioning, number of communication rounds, and evaluation protocol for each dataset. Thus, differences in performance are attributed to the local optimization strategy, checkpoint-selection mechanism, and drift-control design rather than to model capacity or data partitioning. Table VI summarizes the role of each method. TABLE VI: Interpretation of QFL baselines and DUQFL ablation variants. Method Deep unfold Controller Best unfold Drift control Role Default-QFL × × × × Baseline QFL using Qiskit’s default SPSA optimizer and FedAvg aggregation. Adam-QFL × × × × Generic adaptive optimizer baseline used to assess whether standard local adaptivity is sufficient under QFL heterogeneity. FedProx-QFL × × × ✓ Proximal QFL baseline that tests drift control without deep unfolding or controller-guided SPSA adaptation. DUQFL-last ✓ ✓ × × Tests whether uploading the final unfolded local checkpoint is sufficient. DUQFL-best ✓ ✓ ✓ × Tests validation-based best-unfold checkpoint selection. DUQFL-drift ✓ ✓ ✓ ✓ Tests explicit displacement-based drift control within unfolded local QNN optimization. DUQFL-Prox ✓ ✓ ✓ ✓ Full proposed method combining controller-guided unfolded SPSA, best-unfold selection, and proximal drift regularization. Although DUQFL-drift and DUQFL-Prox both include drift-control mechanisms, they regularize drift differently. DUQFL-drift uses an explicit displacement constraint, whereas DUQFL-Prox uses a proximal penalty that continuously discourages deviation from the broadcast global model. Similarly, FedProx-QFL and DUQFL-Prox both use proximal regularization, but at different levels: FedProx-QFL regularizes a conventional local optimization routine, whereas DUQFL-Prox embeds proximal regularization within a controller-guided unfolded QNN optimization trajectory. E-B Effect of Best-Unfold Selection and Proximal Drift Control Figure 9 presents the ablation study on the Genome non-IID setting. The comparison isolates the contribution of the main DUQFL-Prox components: controller-guided unfolding, validation-based best-unfold selection, and drift control. DUQFL-last uploads the final unfolded local state, DUQFL-best uploads the validation-preferred unfolded checkpoint, DUQFL-drift adds explicit displacement control, and DUQFL-Prox incorporates proximal drift regularization. The results show that DUQFL-last improves over Default-QFL in some rounds, but its client-level performance remains less stable than DUQFL-best and DUQFL-Prox. This indicates that the final unfolded state is not necessarily the most generalizable checkpoint under heterogeneous client data. Validation-based best-unfold selection improves this behaviour by preventing over-aggressive or poorly generalizing later unfold states from being uploaded to the server. Adding drift control further improves stability. DUQFL-Prox achieves the highest final mean client test accuracy, the lowest train–test gap, and the smallest client fairness gap among the DUQFL variants. These trends suggest that proximal regularization improves aggregation compatibility by limiting excessive client movement, while best-unfold selection improves the quality of the checkpoint selected for aggregation. The comparison with FedProx-QFL is also important. Although both FedProx-QFL and DUQFL-Prox include proximal regularization, FedProx-QFL applies it to a standard local optimization process. In contrast, DUQFL-Prox applies proximal regularization inside an adaptive unfolded QNN optimization trajectory whose SPSA learning rate and perturbation scale are generated by a controller. Therefore, the improvement of DUQFL-Prox over FedProx-QFL indicates that proximal regularization alone is insufficient; stable heterogeneous QFL also requires adaptive control of the local QNN optimization process before aggregation. (a) Generated learning rate η (b) Generated perturbation δ Figure 10: Client- and unfold-level controller behaviour in DUQFL-Prox. The heatmaps show that the generated SPSA learning rate and perturbation scale vary across clients and unfold steps, supporting the claim that DUQFL-Prox performs controller-guided local optimization rather than using a fixed SPSA schedule. (a) param shift over rounds (b) scatter lr pert loss delta Figure 11: Trace-level diagnostics of local optimization in DUQFL-Prox. The parameter-shift curve monitors local movement during unfolded optimization, while the scatter plot shows that controller-generated SPSA hyperparameters remain within a stable numerical range and produce positive local loss reduction for most client-unfold updates. (a) Normalized η, δ, and loss reduction (b) selected unfold distribution Figure 12: Unfold-level behaviour of DUQFL-Prox. The normalized plot shows that the largest local loss reduction often occurs during early unfold steps, while the selected-unfold distribution confirms that the final unfolded state is not always the checkpoint uploaded for aggregation. These diagnostics support the use of validation-based best-unfold selection. Appendix F Additional Controller Trace Analysis This appendix provides trace-level diagnostics for DUQFL-Prox. While the main ablation study evaluates global and client-level performance, the following figures examine the internal behaviour of the controller during unfolded local QNN optimization. These diagnostics support the claim that DUQFL-Prox does not use a fixed SPSA schedule, but instead generates client- and unfold-dependent optimization behaviour. Figure 10 reports the learning rate η and perturbation scale δ generated by the controller for each client and unfold step. The heatmaps show that the generated SPSA hyperparameters vary across clients and unfold steps, indicating that the controller adapts local optimizer behaviour according to the optimization state rather than applying a single fixed schedule. Figure 11 provides two complementary diagnostics. The parameter-shift curve monitors local movement across communication rounds, whereas the loss-reduction scatter plot relates controller-generated hyperparameters to local loss improvement. Most client-unfold updates yield positive loss reduction while the generated learning rates and perturbation scales remain within a stable numerical range. This supports the interpretation that the controller performs fine-grained adaptation without inducing unstable hyperparameter excursions. Figure 12 further supports the validation-based checkpointing mechanism. The normalized unfold-step analysis shows that the largest average loss reduction often occurs during early unfold steps, while the selected-unfold distribution confirms that the final unfolded state is not always selected for upload. Therefore, best-unfold selection is necessary to prevent later, less generalizable unfold states from dominating the server aggregation. F-A Trace Logging and Reproducibility To support reproducibility and detailed ablation analysis, each experimental run saved round-level, client-level, and classification-level outputs. Round-level CSV files stored global accuracy, validation loss, client accuracy statistics, meta-loss, fairness information, and timing. Client-level trace files stored the unfold index, generated learning rate, perturbation scale, local losses, local train/test accuracy, validation loss, parameter displacement, clipping indicator, selected unfold index, client size, and heterogeneity value. These trace logs make it possible to verify whether the controller changes the local SPSA behaviour across clients, unfold steps, and communication rounds. The main saved outputs were: • global_accuracies.csv: global and client train/test accuracies across communication rounds; • validation.csv: validation loss across rounds; • client_trace.csv: unfold-level client traces, including generated η, generated δ, validation loss, parameter shift, and selected unfold index; • outer_meta.csv: outer SPSA controller-update information, including perturbed meta-loss values and controller-gradient norm; • classification_metrics.csv: precision, recall, F1-score, ROC-AUC, PR-AUC, MCC, specificity, and confusion-matrix values; • global_params.npz: saved global QNN parameter vectors across communication rounds. The result files were organized by method, dataset, split type, random seed, and timestamp. This structure allows each run to be regenerated from the stored configuration and output files. Appendix G IBM Quantum Hardware Workload Evidence Figure 13: IBM Quantum Platform workload evidence for real-hardware execution of trained QFL checkpoints. The workload history shows completed executions on the ibm_fez backend using finite-shot measurement. This screenshot is included as execution evidence only; the quantitative simulator–hardware accuracy comparison is reported separately. To assess practical deployability, selected trained global QNN checkpoints were executed on real IBM quantum hardware. This experiment was performed as post-training hardware validation rather than full hardware-in-the-loop federated training. The saved global checkpoints were assigned to the trained QNN circuits, transpiled for the selected IBM backend, and executed using finite-shot measurement. Figure 13 provides IBM Quantum Platform workload evidence for the hardware execution. The workload history shows completed executions on the ibm_fez backend using the IBM Quantum open instance. This screenshot is included only as execution evidence; the quantitative simulator–hardware accuracy comparison is reported separately in the main experimental results.