Paper deep dive
FedCritic-MIMO: Communication-Efficient Serverless Federated Critic Learning for Massive-MIMO Resource Control in Open and Disaggregated 6G RANs
Amin Farajzadeh, Melike Erol-Kantarci
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/5/2026, 3:24:36 AM
Summary
The paper introduces FedCritic-MIMO, a communication-efficient, serverless federated multi-agent reinforcement learning framework designed for resource control in open and disaggregated 6G Radio Access Networks (RANs). It addresses the challenge of coordinating independently deployable cell-level controllers in massive-MIMO OFDMA deployments without a central trainer. The framework allows controllers to retain local actors and personalized critic components while exchanging only compatible shared critic parameters peer-to-peer. Key innovations include wireless-aware event triggering, adaptive layer-wise top-k sparse critic exchange with error feedback, and balanced interference-aware fusion. Simulations demonstrate that FedCritic-MIMO achieves superior performance-communication tradeoffs, reducing critic-communication overhead by 76% compared to uncompressed distributed methods, while improving throughput, SINR, and QoS satisfaction.
Entities (12)
Relation Signals (8)
Amin Farajzadeh → affiliatedwith → University of Ottawa
confidence 95% · A. Farajzadeh and M. Erol-Kantarci are with the NETCORE Lab... University of Ottawa
Melike Erol-Kantarci → affiliatedwith → University of Ottawa
confidence 95% · A. Farajzadeh and M. Erol-Kantarci are with the NETCORE Lab... University of Ottawa
FedCritic-MIMO → targets → 6G RANs
confidence 95% · FedCritic-MIMO targets reuse-1 multi-cell massive-MIMO OFDMA deployments in open and disaggregated 6G RANs.
FedCritic-MIMO → uses → Serverless Learning
confidence 92% · FedCritic-MIMO is a communication-efficient serverless federated multi-agent reinforcement learning framework.
FedCritic-MIMO → employs → Top-k sparse exchange
confidence 90% · FedCritic-MIMO combines... adaptive layer-wise top-k sparse critic exchange with error feedback.
FedCritic-MIMO → reducesoverheadby → 76%
confidence 90% · It reduces critic-communication overhead by 76% relative to uncompressed distributed critic exchange.
FedCritic-MIMO → optimizes → QoS
confidence 88% · FedCritic-MIMO targets... long-term QoS with limited inter-controller signaling.
FedCritic-MIMO → improves → SINR
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:This paper proposes FedCritic-MIMO, a communication-efficient serverless federated multi-agent reinforcement learning framework for AI-native resource control across independently deployable cell-level controllers in open and disaggregated 6G RANs. Controllers share no trainer, retain local actors and personalized critic components, and exchange only compatible shared critic parameters. FedCritic-MIMO targets reuse-$1$ multi-cell massive-MIMO OFDMA deployments, where RAN controllers jointly manage user scheduling, per-stream power allocation, beamforming, interference, and long-term QoS with limited inter-controller signaling. Each base station locally executes its actor without centralized training or actor federation, while critic knowledge is exchanged peer-to-peer over an interference-aware graph. It enables this collaboration through wireless-aware event triggering, adaptive layer-wise top-$k$ sparse critic exchange with error feedback, and balanced interference-aware fusion. We establish conditional finite-time stationarity and consensus guarantees for the balanced, compressed peer-to-peer critic recursion under a fixed-policy, frozen-target critic-regression model. In strongly interference-coupled reuse-$1$ simulations, FedCritic-MIMO achieves the best performance-communication tradeoff among heuristic, independent-learning, centralized-training, and communication-ablation baselines. It achieves the highest held-out throughput, improves user-rate distribution and mean SINR, increases QoS satisfaction, and attains the lowest interference cost per delivered bit among learning baselines. It reduces critic-communication overhead by $76\%$ relative to uncompressed distributed critic exchange. These results demonstrate that serverless exchange of compatible shared critic parameters can coordinate RAN controllers without centralized trajectory collection or parameter-server aggregation.
Tags
Links
- Source: https://arxiv.org/abs/2608.03852v1
- Canonical: https://arxiv.org/abs/2608.03852v1
Trouble viewing inline? Open PDF directly →
Full Text
93,470 characters extracted from source content.
Expand or collapse full text
FedCritic-MIMO: Communication-Efficient Serverless Federated Critic Learning for Massive-MIMO Resource Control in Open and Disaggregated 6G RANs Amin Farajzadeh Member IEEE Melike Erol-Kantarci Fellow IEEE A. Farajzadeh and M. Erol-Kantarci are with the NETCORE Lab, School of Electrical Engineering and Computer Science, University of Ottawa, Ottawa, ON K1N 6N5, Canada (e-mails: amin.farajzadeh, melike.erolkantarci@uottawa.ca). Abstract This paper proposes FedCritic-MIMO, a communication-efficient serverless federated multi-agent reinforcement learning framework for AI-native resource control across independently deployable cell-level controllers in open and disaggregated 6G radio access networks (RANs). We consider a setting in which neighboring controllers do not rely on a common trainer, retain local actors and personalized critic components, and exchange only compatible shared critic parameters. FedCritic-MIMO targets reuse-11 multi-cell massive multiple-input multiple-output (massive-MIMO) orthogonal frequency-division multiple access (OFDMA) deployments where distributed RAN controllers jointly manage user scheduling, per-stream power allocation, beamforming, interference, and long-term quality-of-service (QoS) under limited inter-controller signaling. Rather than centralizing training or federating actors, each base station (BS) executes its actor locally, while critic knowledge is exchanged through peer-to-peer coordination over an interference-aware graph. To make such collaboration practical, FedCritic-MIMO combines wireless-aware event triggering, adaptive layer-wise top-k sparse critic exchange with error feedback, and balanced interference-aware fusion. We further establish conditional finite-time stationarity and consensus guarantees for the proposed balanced, compressed peer-to-peer critic recursion under a fixed-policy, frozen-target critic-regression model. In strongly interference-coupled reuse-11 simulations, FedCritic-MIMO achieves the best performance–communication tradeoff among heuristic, independent-learning, centralized-training, and communication-ablation baselines. In particular, it achieves the highest held-out network throughput, improves user-rate distribution and mean signal-to-interference-plus-noise ratio (SINR), increases QoS satisfaction, and attains the lowest interference cost per delivered bit among learning baselines. It also reduces critic-communication overhead by about 76%76\% relative to uncompressed distributed critic exchange. These results demonstrate that serverless exchange of compatible shared critic parameters can coordinate open and disaggregated RAN controllers without centralized trajectory collection, parameter-server aggregation, or actor homogenization. Index Terms: AI-native 6G RAN, Open RAN, federated reinforcement learning, multi-agent reinforcement learning, serverless learning, massive MIMO, resource allocation, communication-efficient learning. I Introduction Ultra-dense sixth-generation (6G) radio access networks (RANs), formally associated with IMT-2030, are expected to support substantially higher capacity, reliability, connectivity, and intelligence than current cellular systems [1, 2]. These requirements must be met not only through new air-interface capabilities, but also through open, disaggregated, and programmable RAN architectures in which radio-control functions are no longer confined to a monolithic BS implementation [4, 3]. In such architectures, neighboring cell-level controllers may be independently deployed, operate with local observations and implementation-specific model components, and lack a common centralized trainer. The architectural challenge is therefore clear: interference-coupled controllers must learn coordinated radio policies without centralized trajectory collection, unrestricted model sharing, or periodic aggregation through a parameter server. This challenge is particularly important for ultra-dense massive multiple-input multiple-output (massive-MIMO) orthogonal frequency-division multiple access (OFDMA) systems. Massive MIMO improves spatial multiplexing, coverage, and interference suppression, while OFDMA provides flexible time–frequency resource allocation and compatibility with multiuser scheduling [5, 6]. This paper considers a reuse-11 multi-cell massive-MIMO OFDMA downlink in which each BS serves multiple single-antenna user equipments (UEs) over orthogonal subcarriers and spatially multiplexes several UEs per subcarrier through space-division multiple access (SDMA). In this setting, aggressive frequency reuse improves spectral utilization but creates strong inter-cell interference. At the same time, multiuser spatial multiplexing introduces intra-cell SDMA interference whenever the co-scheduled UEs are not perfectly separated by the selected beamformers. The network performance is therefore jointly determined by user scheduling, stream activation, per-stream power allocation, beamforming, inter-cell interference, intra-cell interference, and the evolving quality-of-service (QoS) states of the UEs. These decisions affect throughput, cell-edge rates, fairness, and long-term QoS satisfaction, yielding a high-dimensional, mixed discrete–continuous, and time-varying control problem [7]. Classical optimization methods provide useful benchmarks, but they commonly require accurate global channel and interference information, explicit system models, and repeated solution of mixed-integer nonconvex programs [8]. These requirements are difficult to satisfy in ultra-dense RANs, where channels, traffic demands, queue states, and interference patterns vary rapidly. They are also difficult to reconcile with open and disaggregated control deployments, where local measurements, experience, actors, and implementation-specific critic components remain within independently operated cell-level controllers. Multi-agent reinforcement learning (MARL) is a natural candidate because each BS can learn from its local channel, queue, and interference observations [9, 10]. However, reuse-11 cells are not independent learners: the action of one BS changes the interference, rates, queues, and future scheduling priorities of neighboring cells. Purely local MARL can therefore learn unstable or poorly coordinated policies. Centralized reinforcement learning and centralized-training decentralized-execution (CTDE) can improve coordination, but typically require joint observations, coordinated trajectory collection, centralized critics, or common training infrastructure [11]. In the considered open and disaggregated RAN setting, such requirements reintroduce a common trainer across controllers that are intended to remain independently deployable. To address this problem, we develop FedCritic-MIMO, a communication-efficient serverless federated critic framework for MARL-based joint scheduling, power allocation, and beamforming in interference-coupled multi-cell massive-MIMO OFDMA networks. Each cell-level controller executes its actor locally and retains its local experience, actor parameters, and personalized critic components. Collaboration is restricted to a predefined compatible shared critic subnetwork, which is exchanged directly with neighboring controllers over the interference graph without a central trainer or parameter server. FedCritic-MIMO makes critic collaboration both learning-aware and radio-aware. Critic exchange is triggered according to critic innovation, queue urgency, and interference intensity. Adaptive layer-wise top-k sparsification with error feedback reduces inter-controller model traffic, while symmetric interference-aware fusion assigns greater weight to strongly coupled neighbors. The resulting design coordinates distributed controllers without centralized trajectory aggregation, periodic full-model synchronization, or homogenization of cell-specific actors and personalized critic heads. I-A Related Work I-A1 AI-Native Control in Open and Disaggregated RANs Open and disaggregated RAN architectures expose programmable control functions and interfaces for intelligent radio-resource management [4]. Recent work has studied learning-enabled control for O-RAN resource management, including joint scheduling, O-RU association, and power allocation [12], as well as federated learning for O-RAN slicing and resource management [13]. These works demonstrate the value of AI-enabled control in programmable RAN architectures. However, they do not address serverless critic collaboration among independently deployable cell-level controllers when exchange is restricted to a compatible shared model component. They also do not jointly consider massive-MIMO multiuser scheduling, per-stream power allocation, beamforming, dynamic inter-cell interference, and long-term per-user QoS. I-A2 Learning-Based Wireless Resource Allocation and Beamforming Deep reinforcement learning (DRL) and MARL have been widely studied for dynamic spectrum access, user scheduling, power control, and interference management [14, 15]. These methods enable wireless controllers to adapt to time-varying channels and traffic without repeatedly solving nonconvex optimization problems online. More recent studies have considered higher-dimensional wireless actions, including joint resource allocation, beamforming, and beam combining [16, 17], as well as joint beamforming and subcarrier allocation under queue-aware delay objectives [18]. These works are closely related to the physical-layer control aspect of this paper, but they do not address communication-efficient serverless critic learning over the physical interference graph. I-A3 Centralized and Decentralized MARL Centralized and CTDE-based MARL methods use centralized critics, joint observations, value decomposition, or global training signals to improve coordination among interacting agents [19, 20]. However, collecting channel, queue, interference, scheduling, and beamforming information from multiple BSs creates substantial signaling and scalability requirements [21]. More importantly for the considered architecture, centralized training assumes a common learning function across cell-level controllers that are otherwise independently deployable. Fully decentralized MARL removes this dependency by allowing each BS to learn from local observations [14, 15, 25]. Decentralized actor–critic methods can further exchange critic information through graph-based consensus [22, 23]. Nevertheless, unconditional consensus incurs persistent model traffic, treats neighboring updates without considering their time-varying radio relevance, and can dilute locally useful value information under heterogeneous observations, rewards, and traffic dynamics [24]. These limitations motivate selective critic exchange that preserves local actors and personalized critic components. I-A4 Federated and Communication-Efficient MARL Federated learning supports collaborative model training without exchanging raw local data, while FedAvg and FedProx provide foundational aggregation mechanisms under statistical and systems heterogeneity [26, 27]. Federated MARL has been applied to wireless and edge control [28], including channel assignment and power control [29], decentralized policy collaboration [30], and communication–computation co-optimization [31]. Most existing methods nevertheless rely on a central parameter server, periodic aggregation, or actor sharing. These mechanisms can create communication bottlenecks, homogenize policies across heterogeneous cells, and assume broader model compatibility than is available in the considered open and disaggregated control setting. Gossip-based decentralized optimization provides a basis for serverless model exchange [32], while error feedback compensates for the bias introduced by sparse compressors such as top-k [33]. However, generic gossip and compression mechanisms do not determine when a critic update is useful for wireless control. In an interference-coupled RAN, the value of an update depends not only on parameter innovation, but also on queue urgency and physical interference coupling. I-A5 Positioning of This Work General decentralized actor–critic and compressed-consensus methods provide the algorithmic foundations for peer-to-peer learning. Our earlier FedCritic framework introduced serverless critic collaboration for scheduling and power control in multi-cell OFDMA networks [34]. FedCritic-MIMO addresses the additional architectural and radio-control problem considered here: collaborative learning among independently deployable Open RAN cell-level controllers that lack a common trainer and exchange only a predefined shared critic subnetwork. Within this constraint, FedCritic-MIMO jointly addresses multiuser massive-MIMO scheduling, per-stream power allocation, structured beamforming, dynamic inter-cell interference, and long-term per-user QoS. To the best of our knowledge, it is the first framework to combine communication-efficient serverless critic collaboration, compatibility- restricted shared-model exchange, interference-aware model fusion, multiuser massive-MIMO beamforming, and long-term QoS management. FedCritic-MIMO provides this functionality through: i) local actors and personalized critic heads, with peer exchange restricted to the shared critic subnetwork; i) utility-aware triggering based on critic innovation, queue urgency, and interference intensity; i) adaptive layer-wise top-k exchange with error feedback; and iv) symmetric interference-aware balanced fusion. I-B Contributions The main contributions of this paper are summarized as follows: • We address the joint improvement of network throughput, long-term QoS satisfaction, and interference efficiency across independently deployable cell-level controllers in a reuse-11 open and disaggregated massive-MIMO OFDMA RAN. The scheduling, power-allocation, and beamforming problem is formulated as an interference-coupled decentralized partially observable Markov decision process (Dec–POMDP) that captures multiuser spatial multiplexing, intra-cell and inter-cell interference, budget-safe power control, structured beamforming, and virtual-queue-based QoS constraints. • We develop a personalized serverless federated critic architecture without a common trainer or parameter server. Each controller retains its local actor, experience, and personalized critic head, while only a predefined compatible shared critic subnetwork is exchanged with neighboring controllers. • We introduce a communication-efficient collaboration mechanism that combines utility-aware event triggering, adaptive layer-wise top-k sparsification, and error feedback. The triggering utility jointly captures critic innovation, queue urgency, and interference intensity, while symmetric interference-aware balanced fusion prioritizes strongly coupled neighbors and preserves the network-average shared critic update. • We establish conditional finite-time stationarity and consensus bounds for the balanced shared-critic recursion under a fixed policy and frozen-head, frozen-target critic-regression objective. Under the stated regularity and tracking assumptions, the analysis yields an (T−1/2)+(logT/T)O(T^-1/2)+O( T/T) randomized-iterate stationarity rate. • We evaluate FedCritic-MIMO against heuristic, independent-learning, centralized-training, and communication-ablation baselines. FedCritic-MIMO achieves the strongest held-out throughput, user-rate, signal-to-interference-plus-noise ratio (SINR), QoS, and interference-efficiency performance among the considered learning methods, while reducing critic-communication overhead by approximately 76%76\% relative to uncompressed distributed critic exchange. I System Model I-A Network Setting and Channel Modeling We consider the downlink of an open and disaggregated multi-cell massive-MIMO OFDMA RAN with reuse-11 frequency allocation. Each cell is managed by an independently deployable programmable RAN controller indexed by the corresponding BS index. The controller performs cell-level radio-resource control using locally available channel measurements, queue states, interference observations, and QoS information. The controllers do not rely on a common trainer or parameter server; local actors, local experience, and personalized critic components remain private, while only compatible shared critic parameters are eligible for peer-to-peer exchange during training. This architectural constraint defines the information and model-sharing structure considered in this work, while the physical-layer signal model follows the standard multi-cell massive-MIMO OFDMA downlink. The network consists of N BSs sharing a total bandwidth B, partitioned into K orthogonal subcarriers of equal width Δf=B/K f=B/K. Each BS n∈≜1,…,Nn \1,…,N\ is equipped with LnL_n transmit antennas and serves a set of associated single-antenna UEs ℳnM_n. Time is slotted with index t=0,1,2,…t=0,1,2,…, and in each slot, the controller associated with BS n makes online decisions on user scheduling, linear precoding/beamforming, and transmit-power allocation. For notational simplicity, “BS n” refers to BS n together with its associated local RAN controller unless the distinction is needed explicitly. The slot duration is normalized to one. If an explicit slot duration TsT_s is used, the rate expressions below are in bit/s and the corresponding per-slot service terms used in the virtual queues should be TsRn,m(t)T_sR_n,m(t) and TsRn,mminT_sR _n,m. Equivalently, one may interpret all rates below as normalized per-slot service units. On each subcarrier k, BS n may simultaneously multiplex multiple UEs through linear precoding. Let xn,k,m(t)∈0,1x_n,k,m(t)∈\0,1\ indicate whether UE m∈ℳnm _n is scheduled by BS n on subcarrier k in slot t. The intra-cell spatial multiplexing constraint is ∑m∈ℳnxn,k,m(t)≤Sn,∀n,k, _m _nx_n,k,m(t)≤ S_n, ∀ n,k, (1) where SnS_n denotes the maximum number of simultaneously served streams on each subcarrier at BS n, with 1≤Sn≤minLn,|ℳn|1≤ S_n≤ \L_n,|M_n|\. For each scheduled UE m on subcarrier k, BS n allocates transmit power pn,k,m(t)≥0p_n,k,m(t)≥ 0 and a unit-norm beamforming vector n,k,m(t)∈ℂLn×1v_n,k,m(t) ^L_n× 1 satisfying ‖n,k,m(t)‖2=1,∀n,k,m.\|v_n,k,m(t)\|^2=1, ∀ n,k,m. (2) For unscheduled streams, pn,k,m(t)p_n,k,m(t) is forced to zero by the scheduling–power coupling constraint introduced below, and the corresponding beamforming vector is irrelevant to the transmitted signal. Let sn,k,m(t)∼(0,1)s_n,k,m(t) (0,1) denote the information symbol intended for UE m. The transmitted signal vector from BS n on subcarrier k is n,k(t)=∑m∈ℳnxn,k,m(t)pn,k,m(t)n,k,m(t)sn,k,m(t).u_n,k(t)= _m _nx_n,k,m(t) p_n,k,m(t)\,v_n,k,m(t)\,s_n,k,m(t). (3) Since ‖n,k,m(t)‖2=1\|v_n,k,m(t)\|^2=1 and [|sn,k,m(t)|2]=1E[|s_n,k,m(t)|^2]=1, the average transmit power used by BS n on subcarrier k is ∑m∈ℳnxn,k,m(t)pn,k,m(t) _m _nx_n,k,m(t)p_n,k,m(t). The per-BS power budget is therefore ∑k=1K∑m∈ℳnxn,k,m(t)pn,k,m(t)≤Pn,∀n. _k=1^K _m _nx_n,k,m(t)\,p_n,k,m(t)≤ P_n, ∀ n. (4) For any BS b and UE m∈ℳnm _n, let b,k,m(n)(t)∈ℂLb×1h_b,k,m^(n)(t) ^L_b× 1 denote the downlink channel vector from BS b to UE m associated with BS n on subcarrier k in slot t. For the desired link (b=nb=n), we use the shorthand n,k,m(t)≜n,k,m(n)(t)h_n,k,m(t) _n,k,m^(n)(t). We model each BS–UE channel vector as b,k,m(n)(t)=βb,m(n)(¯b,m(n))1/2~b,k,m(n)(t),h_b,k,m^(n)(t)= _b,m^(n)\, ( R_b,m^(n) )^1/2 h_b,k,m^(n)(t), (5) where βb,m(n)>0 _b,m^(n)>0 denotes the linear-scale channel-power gain including large-scale fading, ¯b,m(n)⪰ R_b,m^(n) 0 is the normalized spatial correlation matrix satisfying 1Lbtr(¯b,m(n))=1 1L_btr( R_b,m^(n))=1, and ~b,k,m(n)(t)∼(,Lb) h_b,k,m^(n)(t) (0,I_L_b) captures the small-scale fading. With this normalization, [‖b,k,m(n)(t)‖2]=Lbβb,m(n)E[\|h_b,k,m^(n)(t)\|^2]=L_b _b,m^(n). The large-scale fading coefficient is modeled in dB as βb,m(n)[dB]=β0−10αlog10(db,m(n)d0)+Fb,m(n), _b,m^(n)[dB]= _0-10α _10\! ( d_b,m^(n)d_0 )+F_b,m^(n), (6) where db,m(n)d_b,m^(n) is the distance between BS b and UE m, α is the path-loss exponent, and Fb,m(n)∼(0,σsh2)F_b,m^(n) (0, _sh^2) is the shadowing term in dB. The coefficient used in (5) is therefore βb,m(n)=10βb,m(n)[dB]/10 _b,m^(n)=10 _b,m^(n)[dB]/10. The temporal evolution of the small-scale fading is modeled by a first-order complex Gauss–Markov process [35] as ~b,k,m(n)(t)=ρ~b,k,m(n)(t−1)+1−ρ2b,k,m(n)(t), h_b,k,m^(n)(t)=ρ\, h_b,k,m^(n)(t-1)+ 1-ρ^2\,w_b,k,m^(n)(t), (7) where b,k,m(n)(t)∼(,Lb)w_b,k,m^(n)(t) (0,I_L_b) is i.i.d. across all indices, and 0≤ρ<10≤ρ<1 is the temporal correlation coefficient. The process is assumed to be initialized with ~b,k,m(n)(0)∼(,Lb) h_b,k,m^(n)(0) (0,I_L_b), which preserves the marginal distribution over time. The channel is assumed constant within each slot and evolves across slots according to (7). The received signal at UE m∈ℳnm _n on subcarrier k in slot t is yn,k,m(t)=n,k,mH(t)xn,k,m(t)pn,k,m(t)n,k,m(t)sn,k,m(t)⏟desired signal y_n,k,m(t)= h_n,k,m^H(t)\,x_n,k,m(t) p_n,k,m(t)\,v_n,k,m(t)\,s_n,k,m(t)_desired signal +∑j∈ℳnj≠mn,k,mH(t)xn,k,j(t)pn,k,j(t)n,k,j(t)sn,k,j(t)⏟intra-cell interference -2.84526pt+ _ subarraycj _n\\ j≠ m subarrayh_n,k,m^H(t)\,x_n,k,j(t) p_n,k,j(t)\,v_n,k,j(t)\,s_n,k,j(t)_intra-cell interference +∑b≠n∑j∈ℳb(b,k,m(n)(t))Hxb,k,j(t)pb,k,j(t)b,k,j(t)sb,k,j(t)⏟inter-cell interference+zn,k,m(t), -2.84526pt+ _b≠ n _j _b (h_b,k,m^(n)(t) )^Hx_b,k,j(t) p_b,k,j(t)\,v_b,k,j(t)\,s_b,k,j(t)_inter-cell interference+z_n,k,m(t), (8) where zn,k,m(t)∼(0,N0Δf)z_n,k,m(t) (0,N_0 f) is additive noise and N0N_0 denotes the noise PSD under the adopted complex-baseband convention. The intra-cell and inter-cell interference powers experienced by UE m∈ℳnm _n on subcarrier k are respectively defined as In,k,mintra(t) I^intra_n,k,m(t) ≜∑j∈ℳnj≠mxn,k,j(t)pn,k,j(t)|n,k,mH(t)n,k,j(t)|2, _ subarraycj _n\\ j≠ m subarrayx_n,k,j(t)p_n,k,j(t) |h_n,k,m^H(t)v_n,k,j(t) |^2, (9) In,k,minter(t) I^inter_n,k,m(t) ≜∑b∈∖n∑j∈ℳbxb,k,j(t)pb,k,j(t)|(b,k,m(n)(t))Hb,k,j(t)|2. _b \n\ _j _bx_b,k,j(t)p_b,k,j(t) | (h_b,k,m^(n)(t) )^Hv_b,k,j(t) |^2. (10) Assuming single-user decoding and treating residual intra-cell and inter-cell interference as noise, the resulting downlink SINR is SINRn,k,m(t)=xn,k,m(t)pn,k,m(t)|n,k,mH(t)n,k,m(t)|2In,k,mintra(t)+In,k,minter(t)+N0Δf.SINR_n,k,m(t)= x_n,k,m(t)p_n,k,m(t) |h_n,k,m^H(t)v_n,k,m(t) |^2I^intra_n,k,m(t)+I^inter_n,k,m(t)+N_0 f. (11) The instantaneous rate of UE m∈ℳnm _n on subcarrier k is Rn,k,msc(t)=Δflog2(1+SINRn,k,m(t)).R_n,k,m^sc(t)= f _2\! (1+SINR_n,k,m(t) ). (12) The per-UE, per-subcarrier-per-BS, and per-BS rates are, respectively, Rn,m(t) R_n,m(t) =∑k=1KRn,k,msc(t), = _k=1^KR_n,k,m^sc(t), (13) cn,k(t) c_n,k(t) =∑m∈ℳnRn,k,msc(t), = _m _nR_n,k,m^sc(t), (14) cn(t) c_n(t) =∑k=1Kcn,k(t). = _k=1^Kc_n,k(t). (15) The minimum-rate parameters Rn,mminR _n,m are assumed to have the same normalized units as Rn,m(t)R_n,m(t) in the virtual-queue update. I-B Optimization Problem Formulation The long-term QoS requirement for UE m∈ℳnm _n is lim infT→∞1T∑t=0T−1[Rn,m(t)]≥Rn,mmin,∀n,m∈ℳn. _T→∞ 1T _t=0^T-1E [R_n,m(t) ]≥ R _n,m, ∀ n,\;m _n. (16) We assume that the vector of minimum-rate requirements is feasible under the available bandwidth, power, scheduling, and beamforming constraints; otherwise, the corresponding virtual queues will diverge and indicate persistent QoS violation. In each slot t, the local cell controllers jointly selects the user scheduling variables (t)=xn,k,m(t)X(t)=\x_n,k,m(t)\, the per-stream powers (t)=pn,k,m(t)P(t)=\p_n,k,m(t)\, and the beamforming vectors (t)=n,k,m(t)V(t)=\v_n,k,m(t)\ to maximize the instantaneous network-wide downlink sum-rate while accounting for the long-term minimum-rate requirements. The long-term QoS constraints are handled through virtual queues, which yields the following queue-weighted per-slot surrogate problem: max(t),(t),(t) -8.53581pt _X(t),\,P(t),\,V(t)\; ∑n=1N∑k=1Kcn,k(t)−∑n=1N∑m∈ℳnQn,m(t)(Rn,mmin−Rn,m(t)) _n=1^N _k=1^Kc_n,k(t)- _n=1^N _m _nQ_n,m(t) (R _n,m-R_n,m(t) ) (17a) s.t. ∑m∈ℳnxn,k,m(t)≤Sn,∀n,k, _m _nx_n,k,m(t)≤ S_n, 00000000000∀ n,k, (17b) ∑k=1K∑m∈ℳnxn,k,m(t)pn,k,m(t)≤Pn,∀n, _k=1^K _m _nx_n,k,m(t)\,p_n,k,m(t)≤ P_n, 00∀ n, (17c) 0≤pn,k,m(t)≤Pnxn,k,m(t),∀n,k,m∈ℳn, 0≤ p_n,k,m(t)≤ P_n\,x_n,k,m(t), 00∀ n,k,\;m _n, (17d) ‖n,k,m(t)‖2=1,∀n,k,m∈ℳn, \|v_n,k,m(t)\|^2=1, 0000000∀ n,k,\;m _n, (17e) xn,k,m(t)∈0,1,∀n,k,m∈ℳn. x_n,k,m(t)∈\0,1\, 000000000∀ n,k,\;m _n. (17f) After normalizing the rate units, (17) is the unit-weight form of the standard drift-plus-penalty surrogate. Equivalently, one may multiply the sum-rate term by a control parameter Vsr>0V_sr>0; we use Vsr=1V_sr=1 throughout. The term −∑n,mQn,m(t)Rn,mmin- _n,mQ_n,m(t)R _n,m is action-independent within slot t, but it is retained to show the deficit-penalty interpretation. The virtual queues evolve as Qn,m(t+1)=[Qn,m(t)+Rn,mmin−Rn,m(t)]+,∀n,m∈ℳn, -5.69054ptQ_n,m(t+1)= [\,Q_n,m(t)+R _n,m-R_n,m(t)\, ]^+,\ \ ∀ n,\;m _n, -5.69054pt (18) where [⋅]+≜max⋅,0[·]^+ \·,0\. Hence, repeatedly solving (17) online steers the system toward satisfying long-term average minimum-rate constraints while still prioritizing instantaneous sum-rate. More precisely, stability of the virtual queues implies satisfaction of the corresponding long-term average-rate constraints in (16). Problem (17) is a mixed-integer nonconvex program because of the binary scheduling variables, coupled power and beamforming decisions, and reuse-11 intra-cell and inter-cell interference. Solving it exactly in every slot would require timely access to network-wide CSI, queue states, QoS states, and interference-coupling information, followed by repeated centralized optimization. This is impractical in large-scale dense deployments and conflicts with the considered open and disaggregated RAN setting, where independently deployable cell-level controllers retain local observations and control components. Purely isolated local control, however, cannot adequately capture the cross-cell effect of each BS’s actions on neighboring rates, queues, and interference. These properties motivate the proposed AI-native learning architecture, in which resource-control execution remains local while compatible shared critic parameters are exchanged selectively among interference-neighbor controllers, without centralized trajectory collection or parameter-server aggregation. I Reformulation as an Interference–Coupled Dec–POMDP The online resource-control problem in (17) can be reformulated as a cooperative decentralized partially observable Markov decision process (Dec–POMDP), in which each BS acts as an autonomous agent and jointly learns with the other BSs to maximize a common long-term network utility. The resulting Dec–POMDP is interference-coupled, since the rate achieved by each BS depends not only on its own scheduling, power-allocation, and beamforming decisions, but also on the simultaneous decisions made by neighboring BSs over the same subcarriers. I-A Dec–POMDP Definition We model the system as the tuple =⟨,,nn∈,nn∈,ℙ,,r,μ0⟩,G= ,S,\A_n\_n ,\O_n\_n ,P,O,r, _0 , (19) where μ0 _0 is the initial-state distribution and the remaining components are defined below. (a) Agents: The agent set is the set of BSs, =1,…,NN=\1,…,N\. Each agent n∈n independently selects its local scheduling, power-allocation, and beamforming decisions at every slot. (b) Global state: The global state at slot t is defined as s(t)=((t),(t))∈,s(t)= (H(t),Q(t) ) , (20) where (t)=b,k,m(n)(t)H(t)=\h_b,k,m^(n)(t)\ collects all direct-link and cross-link channel vectors over all BSs, UEs, and subcarriers, and (t)=Qn,m(t)Q(t)=\Q_n,m(t)\ contains the virtual queues. Under (7) and (18), s(t)\s(t)\ is Markov. If lagged measurements are used as policy inputs, they are either included in an augmented state or handled through the agent’s local action–observation history. (c) Local observations: Since BS n does not have access to the full network state, it observes only a local observation on(t)∈no_n(t) _n. The joint observation is generated according to the observation kernel ((t)∣s(t))O(o(t) s(t)), with (t)=(o1(t),…,oN(t))o(t)=(o_1(t),…,o_N(t)). A deterministic observation map on(t)=Ωn(s(t))o_n(t)= _n(s(t)) is a special case. A natural observation for BS n is on(t)=(nloc(t),n(t),n(t))∈n,o_n(t)= (H^loc_n(t),Q_n(t),I_n(t) ) _n, (21) where nloc(t)=n,k,m(t)H^loc_n(t)=\h_n,k,m(t)\ contains the direct-link CSI available at BS n, n(t)=Qn,m(t)Q_n(t)=\Q_n,m(t)\, contains its local virtual queues, and n(t)I_n(t) contains only locally available interference-side information before the current action, such as measured interference powers, estimated cross-link CSI from dominant neighboring BSs, or compact neighbor summaries exchanged over the coordination graph. Interference created by the simultaneous actions in current slot t is not assumed known before those actions are selected. (d) Local action: At each slot t, BS n chooses an action an(t)=(n(t),n(t),n(t))∈n,a_n(t)= (X_n(t),P_n(t),V_n(t) ) _n, (22) where the feasible set is defined by (17b)–(17f) for BS n. Thus, the local action jointly determines: i) which UEs are scheduled on each subcarrier, i) how much power is allocated to each active stream, and i) which beamforming/precoding vector is used for each scheduled UE. When a structured rule such as RZF is adopted, n(t)V_n(t) is computed from the scheduled-user CSI and a lower-dimensional learned regularization action. The policy is masked and normalized as specified in Section IV so that sampled actions satisfy the feasibility constraints. (e) Joint action: The joint action of all BSs is (t)=(a1(t),…,aN(t))∈≜∏n=1Nn.a(t)= (a_1(t),…,a_N(t) ) _n=1^NA_n. (23) (f) State transition kernel: The transition kernel ℙ(s(t+1)∣s(t),(t))P (s(t+1) s(t),a(t) ) is induced jointly by the stochastic channel evolution and the deterministic queue update. Specifically, the channel component evolves according to (7), while the virtual queues evolve as (18). Since the achieved rate Rn,m(t)R_n,m(t) depends on the SINR in (11), the queue evolution at BS n is coupled not only to its own action an(t)a_n(t), but also to the simultaneous actions of the interfering BSs. (g) Rewards: The common team reward combines the instantaneous network sum-rate with a virtual-queue-weighted QoS term that penalizes rate deficits relative to the minimum-rate requirements and prioritizes users with larger accumulated deficits. Specifically, r(s(t),(t))= -5.69054ptr (s(t),a(t) )= ∑n=1N∑k=1Kcn,k(t)−∑n=1N∑m∈ℳnQn,m(t)(Rn,mmin−Rn,m(t)). -8.53581pt _n=1^N _k=1^Kc_n,k(t)- _n=1^N _m _nQ_n,m(t) (R_n,m -R_n,m(t) ). (24) For decentralized training, BS n uses the corresponding local contribution to the team reward, rn(t)=∑k=1Kcn,k(t)−∑m∈ℳnQn,m(t)(Rn,mmin−Rn,m(t)). -8.53581ptr_n(t)= _k=1^Kc_n,k(t)- _m _nQ_n,m(t) (R_n,m -R_n,m(t) ). (25) The local utilities satisfy r(s(t),(t))=∑n=1Nrn(t)r(s(t),a(t))= _n=1^Nr_n(t), but each rn(t)r_n(t) remains action-coupled because its rates depend on neighboring BS actions through inter-cell interference and on co-scheduled streams through intra-cell interference. I-B Interference-Coupled Structure Let ℬnint⊆∖nB^int_n \n\ denote the set of dominant interferers of BS n. Then, the rate of UE m∈ℳnm _n can be approximated as Rn,m(t)≈ℛn,mloc(s(t),an(t),ab(t)b∈ℬnint),R_n,m(t) _n,m^loc (s(t),a_n(t),\a_b(t)\_b ^int_n ), (26) where ℛn,mloc(⋅)R_n,m^loc(·) denotes the rate mapping obtained from (11) after retaining only the dominant inter-cell interference terms. This local-interference structure motivates graph-based coordination and neighbor-limited information exchange in the proposed decentralized learning framework. The Dec–POMDP is therefore interference-coupled in two senses. First, the immediate reward of each BS depends on neighboring actions through the inter-cell interference terms in (11). Second, the queue dynamics are coupled across BSs through the achieved rates, since stronger interference from neighboring cells increases the virtual-queue backlog of the affected UEs and thus changes future scheduling priorities. I-C Decentralized Control Objective Let πn(an∣on) _n(a_n o_n) denote the stochastic policy of BS n, and let the joint policy factorize as ((t)∣(t))=∏n=1Nπn(an(t)∣on(t)), π(a(t) (t))= _n=1^N _n(a_n(t) o_n(t)), (27) where (t)=(o1(t),…,oN(t))o(t)=(o_1(t),…,o_N(t)). The decentralized control objective is to learn a set of policies πnn=1N\ _n\_n=1^N that maximizes the long-term expected team utility maxπnn=1Nlim infT→∞1T[∑t=0T−1r(s(t),(t))]. _\ _n\_n=1^N\; _T→∞ 1TE_ π [ _t=0^T-1r (s(t),a(t) ) ]. (28) Because the virtual queues are part of the system state and the reward in (24) is derived from the Lyapunov-drift reformulation, optimizing (28) encourages policies that jointly improve network throughput and stabilize the virtual queues, which in turn promotes satisfaction of the long-term minimum-rate requirements. IV Proposed Communication-Efficient Serverless Federated Critic Learning This section presents the proposed communication-efficient serverless federated critic learning framework for the interference-coupled Dec–POMDP in Section I. The key idea is to retain the fully decentralized graph-based collaboration structure of FedCritic while replacing periodic full-parameter gossip with a utility-aware and communication-efficient critic-exchange mechanism. In particular, critic communication is driven not only by local critic drift, but also by wireless-control relevance, namely local queue urgency and interference coupling. Each BS maintains its own local actor and critic, performs local actor–critic updates from locally collected trajectories, and exchanges compressed critic-side information only with neighboring BSs over the coordination graph when the local critic update is deemed sufficiently important. This reduces communication overhead while preserving the benefits of serverless federated value learning in the interference-coupled multi-cell massive-MIMO OFDMA setting. For clarity, we consider synchronous training rounds and reliable peer-to-peer communication over a fixed undirected coordination graph. Critic collaboration is used only during training; after training, each BS executes its local actor using only its local observation. Superscript t denotes a federated training round, whereas τi _i denotes the iith wireless slot in the rollout collected at round t. IV-A Design Rationale Compared with the single-antenna OFDMA setting [34], the massive-MIMO extension substantially enlarges the local state and action spaces due to per-subcarrier multiuser scheduling, per-stream power allocation, and beamforming design. As a result, critic models become larger and more sensitive to local environmental heterogeneity. In such a setting, periodic full-parameter gossip is inefficient, since neighboring BSs may repeatedly exchange highly redundant critic information even when their local critic states have changed only marginally. More importantly, in interference-coupled wireless control, not all critic updates are equally valuable. A critic change occurring at a BS with low queue pressure and weak interference coupling may have little impact on network performance, whereas even a moderate critic change at a BS operating under high QoS pressure or strong interference coupling may be highly relevant to its neighbors. This motivates a wireless-aware communication policy. To address these issues, we propose a communication-efficient serverless critic-learning mechanism based on three principles: 1. critic-only federation: only critic-side information is exchanged across BSs, while actors remain local; 2. utility-aware event-triggered communication: a BS communicates only when its critic update is sufficiently important, as determined jointly by critic innovation, local queue urgency, and interference intensity; 3. compressed balanced interference-aware fusion: when communication is triggered, only a compressed shared-critic increment is exchanged and fused over the interference graph using symmetric interference-relevance weights. IV-B Local Actor–Critic Parameterization and Objectives For BS n∈n , let θnt _n^t denote the local actor parameters. To allow value-function personalization while preserving parameter-space compatibility, write the critic as Vϑnt(o)=hωnt(fψnt(o)),ϑnt=(ψnt,ωnt),V_ _n^t(o)=h_ _n^t\! (f_ _n^t(o) ), _n^t=( _n^t, _n^t), (29) where ψnt∈ℝdc _n^t ^d_c denotes the shared critic parameters and ωnt _n^t denotes a BS-specific value head. All BSs employ compatible shared critic architectures and use the common initialization ψ10=⋯=ψN0=ψ0. _1^0=·s= _N^0=ψ^0. (30) Only ψnt _n^t is exchanged; θnt _n^t and ωnt _n^t remain local. BS n uses the local utility in (25), denoted by r¯n(τ)≜rn(τ). r_n(τ) r_n(τ). Its critic approximates the discounted local value Vϑn(on(τ))≈[∑ℓ=0∞γℓr¯n(τ+ℓ)|on(τ)],V_ _n(o_n(τ)) _ π\! [ _ =0^∞γ r_n(τ+ )\, |\,o_n(τ) ], (31) where γ∈(0,1)γ∈(0,1). Since on(τ)o_n(τ) is generally a partial observation, (31) is a local value approximation rather than the full-state Markov value. The discounted objective is used as a tractable training surrogate for the average-return control objective in (28); no exact equivalence between the two criteria is assumed. At round t, BS n collects the on-policy batch nt=(ont(τi),ant(τi),r¯nt(τi),ont(τi+1))i=0Hnt−1.D_n^t= \ (o_n^t( _i),a_n^t( _i), r_n^t( _i),o_n^t( _i+1) ) \_i=0^H_n^t-1. -5.69054pt (32) Using the current critic, define δn,it _n,i^t =r¯nt(τi)+γVϑnt(ont(τi+1))−Vϑnt(ont(τi)), = r_n^t( _i)+γ V_ _n^t\! (o_n^t( _i+1) )-V_ _n^t\! (o_n^t( _i) ), (33) A^n,it A_n,i^t =∑ℓ=0Hnt−i−1(γλ)ℓδn,i+ℓt,λ∈[0,1], = _ =0^H_n^t-i-1(γλ) _n,i+ ^t, λ∈[0,1], (34) R^n,it R_n,i^t =A^n,it+Vϑnt(ont(τi)). = A_n,i^t+V_ _n^t\! (o_n^t( _i) ). (35) The targets R^n,it\ R_n,i^t\ are treated as constants during critic optimization. The critic regression loss is ℒnc(ϑn;nt)=1Hnt∑i=0Hnt−1(Vϑn(ont(τi))−R^n,it)2.L_n^c( _n;D_n^t)= 1H_n^t _i=0^H_n^t-1 (V_ _n\! (o_n^t( _i) )- R_n,i^t )^2. (36) For the actor, define ϱn,it(θn)=πθn(ant(τi)∣ont(τi))πθnt(ant(τi)∣ont(τi)). _n,i^t( _n)= _ _n\! (a_n^t( _i) o_n^t( _i) ) _ _n^t\! (a_n^t( _i) o_n^t( _i) ). (37) Then the local PPO loss is calculated as ℒna(θn;nt)= _n^a( _n;D_n^t)= −1Hnt∑i=0Hnt−1minϱn,it(θn)A^n,it, - 1H_n^t _i=0^H_n^t-1 \! \ _n,i^t( _n) A_n,i^t, . clip(ϱn,it(θn),1−ϵπ,1+ϵπ)A^n,it .clip\! ( _n,i^t( _n),1- _π,1+ _π ) A_n,i^t \ −βHHnt∑i=0Hnt−1ℋ(πθn(⋅∣ont(τi))). - _HH_n^t _i=0^H_n^t-1H\! ( _ _n(· o_n^t( _i)) ). (38) After local critic optimization, BS n obtains (ψ~nt+1,ω~nt+1)( ψ_n^t+1, ω_n^t+1). For the convergence analysis, one effective update of the shared critic parameters is written as ψ~nt+1=ψnt−ηcgnt,gnt≜∇ψℒnc(ϑnt;nt), ψ_n^t+1= _n^t- _cg_n^t, g_n^t _ψL_n^c( _n^t;D_n^t), (39) where ηc>0 _c>0. Multiple local critic steps are covered only when their effective direction satisfies the stochastic-gradient conditions stated in Section IV-G. The actor and the personalized critic head remain local. IV-C Coordination Graph and Wireless-Relevance Scores Let c=(,ℰ)G_c=(N,E) be the undirected coordination graph and ℬn=b∈∖n:(n,b)∈ℰB_n= \b \n\:(n,b) \ (40) be the neighbor set of BS n. Critic collaboration is peer-to-peer; no parameter server is used. For b∈ℬnb _n, define the interference contributed by BS b to stream (k,m)(k,m) served by BS n during rollout slot τi _i as Inb,k,mt(τi)=∑j∈ℳbxb,k,jt(τi)pb,k,jt(τi)|(b,k,m(n),t(τi))Hb,k,jt(τi)|2. -8.53581ptI_nb,k,m^t( _i)= _j _bx_b,k,j^t( _i)p_b,k,j^t( _i) | (h_b,k,m^(n),t( _i) )^Hv_b,k,j^t( _i) |^2.\ (41) Let nt(τi)=(k,m):xn,k,mt(τi)=1S_n^t( _i)= \(k,m):x_n,k,m^t( _i)=1 \. The normalized queue urgency, total interference intensity, and directional neighbor relevance are Q¯nt=1Hnt|ℳn|∑i=0Hnt−1∑m∈ℳnQn,mt(τi)Qn,mt(τi)+q0,q0>0, Q_n^t= 1H_n^t|M_n| _i=0^H_n^t-1 _m _n Q_n,m^t( _i)Q_n,m^t( _i)+q_0, 18.49988ptq_0>0, (42) I¯nt= I_n^t= 1Hnt∑i=0Hnt−11max1,|nt(τi)|∑(k,m)∈nt(τi)∑b∈ℬnInb,k,mt(τi)∑b∈ℬnInb,k,mt(τi)+N0Δf, 1H_n^t _i=0^H_n^t-1 1 \1,|S_n^t( _i)|\ _(k,m) _n^t( _i) _b _nI_nb,k,m^t( _i) _b _nI_nb,k,m^t( _i)+N_0 f, (43) κ¯nbt= κ_nb^t= 1Hnt∑i=0Hnt−11max1,|nt(τi)|∑(k,m)∈nt(τi)Inb,k,mt(τi)∑b′∈ℬnInb′,k,mt(τi)+N0Δf. 1H_n^t _i=0^H_n^t-1 1 \1,|S_n^t( _i)|\ _(k,m) _n^t( _i) I_nb,k,m^t( _i) _b _nI_nb ,k,m^t( _i)+N_0 f. (44) An empty scheduled-stream set contributes zero. Hence, 0≤Q¯nt<10≤ Q_n^t<1 and 0≤I¯nt,κ¯nbt≤1.0≤ I_n^t, κ_nb^t≤ 1. IV-D Utility-Aware Event Trigger and Compressed Critic Exchange Let ψ^nt ψ_n^t denote the public reconstruction of BS n’s shared critic parameters maintained consistently by BS n and its neighbors. Importantly, ψ^nt ψ_n^t is a communication-side reference and need not equal the current local critic parameters after neighbor fusion. The utility-aware trigger score is Γnt=‖ψ~nt+1−ψ^nt‖2‖ψ^nt‖2+ϵtr(1+αQQ¯nt+αII¯nt), _n^t= \| ψ_n^t+1- ψ_n^t \|_2 \| ψ_n^t \|_2+ _tr (1+ _Q Q_n^t+ _I I_n^t ), (45) where ϵtr>0 _tr>0 and αQ,αI≥0. _Q, _I≥ 0. The communication decision is ξnt=Γnt≥τtht, _n^t=1 \ _n^t≥ _th^t \, (46) where τtht>0 _th^t>0 is a prescribed threshold. A decreasing threshold promotes asymptotically tighter tracking, whereas a positive threshold floor trades residual disagreement for persistent communication savings, as commonly observed in event-triggered learning and federated optimization schemes [36]. Partition the shared critic parameters into LcL_c layers with dimensions d1,…,dLcd_1,…,d_L_c and ∑ℓ=1Lcdℓ=dc. _ =1^L_cd_ =d_c. At round t, layer ℓ retains kn,ℓtk_n, ^t coordinates, where 1≤kn,ℓt≤dℓ.1≤ k_n, ^t≤ d_ . Let kn,ℓtC_k_n, ^t retain the kn,ℓtk_n, ^t largest-magnitude entries of its argument. Define the layer-wise compressor nt()=colℓ=1Lckn,ℓt(ℓ).C_n^t(x)=col_ =1^L_c \C_k_n, ^t(x_ ) \. (47) It satisfies ‖nt()−‖22≤(1−δc)‖22,δc≜infn,t,ℓkn,ℓtdℓ>0. \|C_n^t(x)-x \|_2^2≤(1- _c)\|x\|_2^2, _c _n,t, k_n, ^td_ >0. (48) Initialize ψ^n0=ψ~n0=ψ0 ψ_n^0= ψ_n^0=ψ^0 and n0=.e_n^0=0. Then the variables satisfy nt=ψ~nt−ψ^nt.e_n^t= ψ_n^t- ψ_n^t. (49) Accordingly, the error-compensated discrepancy and transmitted sparse message are nt _n^t =ψ~nt+1−ψ~nt+nt=ψ~nt+1−ψ^nt, = ψ_n^t+1- ψ_n^t+e_n^t= ψ_n^t+1- ψ_n^t, (50) nt _n^t =ξntnt(nt). = _n^tC_n^t(d_n^t). (51) The public reconstruction and residual memory are updated as ψ^nt+1 ψ_n^t+1 =ψ^nt+nt, = ψ_n^t+ _n^t, (52) nt+1 _n^t+1 =nt−nt=ψ~nt+1−ψ^nt+1. =d_n^t- _n^t= ψ_n^t+1- ψ_n^t+1. (53) Thus, all unsent coordinates remain in the public-reconstruction error and are reconsidered in later rounds. Lemma 1 (One-step tracking-error control). For every BS n and round t, ‖nt+1‖22≤(τtht)2(‖ψ^nt‖2+ϵtr)2,ξnt=0,(1−δc)‖nt‖22,ξnt=1.\|e_n^t+1\|_2^2≤ cases ( _th^t)^2 (\| ψ_n^t\|_2+ _tr )^2,& _n^t=0,\\[5.69054pt] (1- _c)\|d_n^t\|_2^2,& _n^t=1. cases (54) Proof: See Appendix A. ∎ IV-E Balanced Interference-Aware Serverless Fusion Because κ¯nbt κ_nb^t and κ¯bnt κ_bn^t are generally different, we define the symmetric edge score as snbt=ϵw+κ¯nbt+κ¯bnt,(n,b)∈ℰ,s_nb^t= _w+ κ_nb^t+ κ_bn^t, (n,b) , (55) where ϵw>0 _w>0, and let snt=∑j∈ℬnsnjts_n^t= _j _ns_nj^t. For b∈ℬnb _n, define wnbt=(1−ωself)snbtmaxsnt,sbt,0<ωself<1,w_nb^t=(1- _self) s_nb^t \s_n^t,s_b^t\, 0< _self<1, (56) and wnnt=1−∑b∈ℬnwnbt.w_n^t=1- _b _nw_nb^t. (57) All other entries are zero. Since snbt=sbnt,s_nb^t=s_bn^t, the matrix t=[wnbt]W^t=[w_nb^t] is symmetric and row stochastic, hence doubly stochastic. Moreover, wnnt≥ωself.w_n^t≥ _self. The shared critic parameters are fused according to ψnt+1=ψ~nt+1+∑b∈ℬnwnbt(ψ^bt+1−ψ^nt+1), _n^t+1= ψ_n^t+1+ _b _nw_nb^t ( ψ_b^t+1- ψ_n^t+1 ), -5.69054pt (58) while the personalized head remains local as ωnt+1=ω~nt+1 _n^t+1= ω_n^t+1. IV-F Integrated Critic Recursion Define t=[ψ1t,…,ψNt] ^t=[ _1^t,…, _N^t], ~t+1=[ψ~1t+1,…,ψ~Nt+1] ^t+1=[ ψ_1^t+1,…, ψ_N^t+1], ^t+1=[ψ^1t+1,…,ψ^Nt+1] ^t+1=[ ψ_1^t+1,…, ψ_N^t+1], and t=[g1t,…,gNt]G^t=[g_1^t,…,g_N^t]. Then (58) is equivalently the following equation: t+1=~t+1+^t+1((t)T−N). ^t+1= ^t+1+ ^t+1 ((W^t)^T-I_N ). (59) Using ~t+1=t−ηct, ^t+1= ^t- _cG^t, we obtain t+1=(t−ηct)(t)T+t, ^t+1= ( ^t- _cG^t )(W^t)^T+R^t, (60) where t=(~t+1−^t+1)(N−(t)T). ^t= ( ^t+1- ^t+1 ) (I_N-(W^t)^T ). (61) Lemma 2 (Average preservation and perturbation control). Let εt2≜1N[‖~t+1−^t+1‖F2]=1N∑n=1N‖nt+1‖22. _t^2 1NE [ \| ^t+1- ^t+1 \|_F^2 ]= 1N _n=1^NE\|e_n^t+1\|_2^2. (62) Then t ^t1 =, =0, (63) ψ¯t+1 ψ^t+1 =ψ¯t−ηcg¯t,g¯t≜1N∑n=1Ngnt, = ψ^t- _c g^t, g^t 1N _n=1^Ng_n^t, (64) 1N‖t‖F2 1NE\|R^t\|_F^2 ≤4εt2. ≤ 4 _t^2. (65) Proof: See Appendix B. ∎ IV-G Conditional Convergence of the Shared Critic Recursion Convergence analysis considers a fixed joint policy π whose induced trajectory process admits a stationary distribution. During the analyzed optimization window, the personalized critic heads and stop-gradient regression targets are held fixed. For BS n, define Fn(ψ)=n∼dn[ℒn,frc(ψ;n)],F(ψ)=1N∑n=1NFn(ψ),F_n(ψ)=E_D_n d_n π [L_n,fr^c(ψ;D_n) ], F(ψ)= 1N _n=1^NF_n(ψ), (66) where ℒn,frcL_n,fr^c denotes the frozen-head, frozen-target loss. Let ψ¯t=1N∑n=1Nψnt,ℰψt=1N∑n=1N‖ψnt−ψ¯t‖22. ψ^t= 1N _n=1^N _n^t, _ψ^t= 1N _n=1^N\| _n^t- ψ^t\|_2^2. (67) Assumption 1. (A1) The matrices tW^t are symmetric and doubly stochastic, and satisfy the following uniform mixing condition: there exists λW∈[0,1) _W∈[0,1) such that ‖t−1NT‖2≤λW,∀t. \|W^t- 1N11^T \|_2≤ _W, ∀ t. -5.69054pt (68) (A2) Each FnF_n is lower bounded and L-smooth, i.e., ‖∇Fn(ψ)−∇Fn(ψ′)‖2≤L‖ψ−ψ′‖2.\|∇ F_n(ψ)-∇ F_n(ψ )\|_2≤ L\|ψ-ψ \|_2. (69) (A3) Let ℱtF_t contain the iterates and all randomness available before the round-t trajectory batches are sampled. The directions satisfy [gnt∣ℱt]=∇Fn(ψnt).E[g_n^t _t]=∇ F_n( _n^t). (70) For finite constants σg2 _g^2 and σav2 _av^2, 1N∑n=1N[∥gnt−∇Fn(ψnt)∥22|ℱt] 1N _n=1^NE\! [ \|g_n^t-∇ F_n( _n^t) \|_2^2 |F_t ] ≤σg2, ≤ _g^2, (71) [∥1N∑n=1N(gnt−∇Fn(ψnt))∥22|ℱt] \! [ \| 1N _n=1^N (g_n^t-∇ F_n( _n^t) ) \|_2^2 |F_t ] ≤σav2. ≤ _av^2. (72) This formulation permits cross-BS gradient-noise correlation. Under conditional independence, σav2≤σg2/N. _av^2≤ _g^2/N. (A4) The local-objective heterogeneity is bounded as 1N∑n=1N‖∇Fn(ψ)−∇F(ψ)‖22≤ζ2,∀ψ. 1N _n=1^N \|∇ F_n(ψ)-∇ F(ψ) \|_2^2≤ζ^2, ∀ψ. -5.69054pt (73) (A5) The shared critic parameters use the common initialization (30), and 0<ηc≤ηmax,0< _c≤ _ , where ηmax _ is sufficiently small relative to L and the uniform spectral gap 1−λW1- _W. Theorem 1. Under Assumption 1, there exist constants C1,…,C6>0C_1,…,C_6>0 independent of T such that 1T∑t=0T−1[‖∇F(ψ¯t)‖22]≤ 1T _t=0^T-1E [\|∇ F( ψ^t)\|_2^2 ]≤ C1(F(ψ¯0)−F⋆)ηcT+C2ηcσav2 C_1 (F( ψ^0)-F ) _cT+C_2 _c _av^2 +C3ηc2(σg2+ζ2)(1−λW)2+C4(1−λW)21T∑t=0T−1εt2, +C_3 _c^2( _g^2+ζ^2)(1- _W)^2+ C_4(1- _W)^2 1T _t=0^T-1 _t^2, -5.69054pt (74) where F⋆F is a lower bound of F. Moreover, 1T∑t=0T−1[ℰψt]≤C5ηc2(σg2+ζ2)(1−λW)2+C6(1−λW)21T∑t=0T−1εt2. 1T _t=0^T-1E[E_ψ^t]≤ C_5 _c^2( _g^2+ζ^2)(1- _W)^2+ C_6(1- _W)^2 1T _t=0^T-1 _t^2. -5.69054pt (75) Proof: See Appendix C. ∎ In this bound, the tracking-error term T−1∑t=0T−1εt2T^-1 _t=0^T-1 _t^2 plays the role of a communication-induced optimization-error term, consistent with wireless federated-learning analyses in which unreliable or resource-constrained communication appears explicitly in the convergence behavior [37]. Corollary 1. Let JTJ_T be uniformly distributed over 0,…,T−1\0,…,T-1\ and independent of the training randomness. If ηc=cηT _c= c_η T -5.69054pt (76) for sufficiently small cη>0c_η>0 and 1T∑t=0T−1εt2=(logT), 1T _t=0^T-1 _t^2=O\! ( TT ), -5.69054pt (77) then [‖∇F(ψ¯JT)‖22]=(T−1/2)+(logT).E [\|∇ F( ψ^J_T)\|_2^2 ]=O(T^-1/2)+O\! ( TT ). (78) For a fixed learning rate and lim supT→∞T−1∑t=0T−1εt2≤ε¯2, _T→∞T^-1 _t=0^T-1 _t^2≤ ^2, the method converges to a stationarity neighborhood whose size is governed by ηcσav2 _c _av^2, ηc2(σg2+ζ2)/(1−λW)2 _c^2( _g^2+ζ^2)/(1- _W)^2, and ε¯2/(1−λW)2. ^2/(1- _W)^2. Proof: See Appendix D. ∎ Remark 1. Lemma 1 gives the exact one-step effect of event triggering and top-k compression on the public-reference error. Corollary 1 deliberately states the additional decay condition on the accumulated tracking error rather than assuming that a particular threshold schedule automatically guarantees it. Establishing (77) requires corresponding control of the shared-critic drift and communication schedule. Remark 2. Theorem 1 concerns first-order stationarity of the fixed-policy, frozen-head, frozen-target aggregate critic objective. It does not establish convergence or global optimality of the evolving PPO actors, the complete actor–critic process, the Dec–POMDP, or the original mixed-integer wireless resource-allocation problem. IV-H Communication Overhead Let knt=∑ℓ=1Lckn,ℓt.k_n^t= _ =1^L_ck_n, ^t. If each transmitted value uses bvb_v bits and a coordinate index uses bi=⌈log2dc⌉b_i= _2d_c bits, the directed logical critic payload at round t is ℬcomm(t)=∑n=1N|ℬn|ξntknt(bv+bi)bits.B_comm(t)= _n=1^N|B_n| _n^tk_n^t(b_v+b_i) . (79) If one physical broadcast is received by all neighbors, the factor |ℬn||B_n| is omitted for that transmitting BS. Equation (79) excludes protocol headers and the scalar relevance summaries required to construct (55). Because the current balanced weights require both κ¯nbt κ_nb^t and κ¯bnt κ_bn^t, these low-rate summaries are exchanged independently of whether the critic trigger fires, or their most recently received values must be used. IV-I Training Procedure Algorithm 1 summarizes the training procedure. Critic collaboration affects training only; decentralized execution uses the local actor at each BS. Input: Graph cG_c; local actors θn0\ _n^0\; common shared-critic initialization ψ0ψ^0; local heads ωn0\ _n^0\; trigger, compression, and fusion parameters. Initialize ψ^n0=ψ~n0=ψn0=ψ0 ψ_n^0= ψ_n^0= _n^0=ψ^0 and n0=e_n^0=0 for all n for training round t=0,1,…t=0,1,… do foreach BS n∈n in parallel do Collect the local rollout ntD_n^t Compute TD residuals, GAE advantages, and frozen return targets using (33)–(35) Perform local critic optimization to obtain (ψ~nt+1,ω~nt+1)( ψ_n^t+1, ω_n^t+1) Update the local actor using (38) Compute Q¯nt Q_n^t, I¯nt I_n^t, and κ¯nbtb∈ℬn\ κ_nb^t\_b _n and exchange the scalar directional-relevance summaries with neighboring BSs Evaluate the trigger, form nt _n^t, and update ψ^nt+1 ψ_n^t+1 and nt+1e_n^t+1 if ξnt=1 _n^t=1 then Transmit the nonzero critic coordinates and their indices to ℬnB_n foreach BS n∈n in parallel do Construct tW^t from (55)–(57) Update the shared critic parameters using (58) and retain ωnt+1=ω~nt+1 _n^t+1= ω_n^t+1 Algorithm 1 Communication-Efficient FedCritic-MIMO V Simulation Results V-A Simulation Setup We evaluate FedCritic-MIMO in a strongly interference-coupled reuse-11 multi-cell massive-MIMO OFDMA downlink consistent with Section I. The network has 77 BSs, 88 UEs per BS, 1616 subcarriers, 3232 BS antennas, and up to 33 simultaneous spatial streams per subcarrier. Consistent with the open and disaggregated RAN setting, each BS is treated as an independently deployable cell-level controller with local observations, local actor execution, and private experience. FedCritic-MIMO does not use a central trainer or parameter server; collaboration is limited to peer-to-peer exchange of shared critic parameters. Large-scale gains follow the lognormal model in Table I, and small-scale channels follow temporally correlated, spatially i.i.d. Rayleigh fading with Gauss–Markov coefficient ρ=0.55ρ=0.55. The radius-22 graph is used as the logical inter-controller critic-exchange graph for critic exchange and fusion; SINR computation includes interference from all co-channel BSs. One environment step is one network-wide wireless slot in which all BSs observe local states, select actions, receive rewards, and update channel and virtual-queue states. Each policy update uses a rollout of H=128H=128 joint steps; thus, U updates correspond to 128U128U network-wide interactions, or 32,00032,000 interactions for U=250U=250. Interactions are counted per wireless slot, not per BS. All learning-based methods use the same local actor architecture, action space, PPO hyperparameters, rollout budget, and teacher-guided warm start. Behavior-cloned actors and pretrained critics are copied identically to all compatible methods before method-specific training, so differences arise only from information scope, critic parameterization, and critic coordination. For FedCritic-MIMO, the actor and personalized critic components remain local during training, while only the shared critic component participates in peer-to-peer collaboration. Each method is trained with six independent random seeds. During training, validation on a fixed held-out channel set is used for checkpoint selection; selected checkpoints are then evaluated without exploration or gradient updates on unseen channel realizations. Unless otherwise stated, learning curves show run-level means with 95%95\% confidence intervals, and final bars and distributions are computed from the selected checkpoints on the held-out evaluation set. Reported communication overhead refers to training-side critic-model traffic among distributed controllers. It does not represent user-plane payload traffic, fronthaul traffic, or a specific standardized Open RAN interface. The principal simulation and learning parameters are summarized in Table I. TABLE I: Main simulation and learning parameters. Parameter Value Parameter Value Network and channel Number of BSs, N 77 UEs per BS, |ℳn||M_n| 88 BS antennas, LnL_n 3232 Maximum spatial streams, SnS_n 33 Subcarriers, K 1616 Frequency reuse 11 Per-BS power budget, PnP_n 11 normalized unit Noise PSD, N0N_0 10−310^-3 Temporal channel correlation, ρ 0.550.55 Small-scale fading Temporally correlated, spatially i.i.d. Rayleigh Direct-link large-scale gain lnβdir∼(−2.3,1.102) β^dir (-2.3,1.10^2) Cross-link gain multiplier 3.03.0 Minimum-rate target, Rn,mminR_n,m 1.91.9 normalized units Coordination graph Radius-22 ring; all BSs contribute physical interference Power-utilization levels, υn _n 0.2,0.4,0.6,0.8,1.00.2,0.4,0.6,0.8,1.0 Normalized RZF levels, α¯ α 10−3,10−2,5×10−2,10−1,5×10−110^-3,10^-2,5×10^-2,10^-1,5×10^-1 Actor–critic training Discount factor, γ 0.990.99 GAE parameter, λ 0.950.95 Actor learning rate, ηa _a 3×10−43× 10^-4 Critic learning rate, ηc _c 5×10−45× 10^-4 PPO clipping, ϵπ _π 0.200.20 Entropy coefficient, βH _H 0.020.02 Actor/critic epochs 4/44/4 Gradient-norm limit 0.500.50 Rollout horizon/minibatch size 128/64128/64 Maximum training updates 250250 Independent training runs 66 Evaluation 66 validation and 3030 held-out channel seeds/run Queue normalization, q0q_0 1010 Reward scaling 0.010.01 Federated communication and fusion Shared critic dimension, dcd_c 50,04850,048 Value/index precision 32/1632/16 bits Overall compression budget, ρc _c 0.200.20 Per-layer ratio bounds [0.10,0.25][0.10,0.25] Initial/floor trigger thresholds τth,0=0.020 _th,0=0.020, τth,min=0.001 _th, =0.001 Queue/interference coefficients αQ=1.0 _Q=1.0, αI=1.5 _I=1.5 Local fusion weight, ωself _self 0.550.55 Exchanged parameters Shared critic only V-B Baselines and Comparison Protocol We compare FedCritic-MIMO with representative heuristic, independent-learning, centralized-training, and communication-ablation baselines. • Random: feasible scheduling, stream activation, power, and RZF-regularization actions are selected randomly. • Greedy-MaxGain: each BS schedules UEs with the largest instantaneous direct-channel gains and uses the maximum available power-fraction level. • Greedy-Queue: each BS applies a queue-weighted channel-gain rule to account for long-term rate deficits without learning. • Greedy-IA-Queue: the Greedy-Queue metric is further normalized by incoming and outgoing interference-risk terms, forming the strongest non-learning heuristic. • Strict-Independent-PPO: each BS trains a purely local actor–critic using only own-cell channel, queue, occupancy, and measured-SINR information, without explicit neighbor features or outgoing-interference pricing. • No-Federation-IA-PPO: each BS uses the same interference-aware observations, reward, actor, and critic architectures as FedCritic-MIMO, but without critic exchange. This isolates the benefit of serverless critic collaboration. • CTDE-MAPPO: decentralized actors are trained with a centralized critic using concatenated BS-level critic features, while execution remains local. • Periodic-Full: every BS exchanges its complete shared critic parameters with coordination neighbors at every training update. • Event-Uncompressed: the proposed event rule determines communication times, but transmitted critic increments are uncompressed. • Proposed: the complete FedCritic-MIMO framework with utility-aware event-triggered shared-critic exchange, adaptive layer-wise top-k compression with error feedback, and balanced interference-aware serverless fusion. ((a)) ((b)) ((c)) Figure 1: Validation learning dynamics over training updates: (a) reward, (b) QoS satisfaction, and (c) interference cost per unit sum rate. V-C Learning Dynamics Fig. 1 reports the evolution of the validation reward, the QoS-satisfaction metric, and the interference cost per unit sum rate. All learning-based methods improve substantially beyond the common warm-start reference, confirming that the subsequent policy updates provide material gains over the teacher-guided initialization. As Fig. 1 (a) shows, the proposed method separates from the other methods after approximately the first one hundred updates and reaches a validation reward of about 56.356.3 at update 250250, compared with approximately 52.552.5, 52.052.0, 51.551.5, 50.350.3, and 48.848.8 for CTDE-MAPPO, Event-Uncompressed, No-Federation-IA-PPO, Periodic-Full, and Strict-Independent-PPO, respectively. The QoS trajectories in Fig. 1 (b), exhibit a similar, although less separated, ordering. At the final displayed update, the proposed method reaches approximately 0.780.78, while CTDE-MAPPO, No-Federation-IA-PPO, Event-Uncompressed, Periodic-Full, and Strict-Independent-PPO reach approximately 0.760.76, 0.740.74, 0.730.73, 0.710.71, and 0.670.67, respectively. The confidence bands of the proposed method and CTDE-MAPPO overlap near the end of training; hence, the figure supports the conclusion that the proposed decentralized method matches the strongest centralized-training baseline in QoS, rather than establishing a statistically significant QoS advantage from the learning curves alone. As shown in Fig. 1 (c), the proposed method also produces the most consistent reduction in the interference-per-rate metric, for which lower values are preferable. It decreases from approximately 2.8×10−42.8× 10^-4 at initialization to about 1.8×10−41.8× 10^-4, whereas the strongest competing learning methods finish between approximately 2.0×10−42.0× 10^-4 and 2.1×10−42.1× 10^-4. Thus, the reward improvement is accompanied by more interference-efficient operation, rather than being obtained solely by increasing the delivered rate without controlling the resulting inter-cell interference. Because the proposed reward curve remains mildly increasing at the last displayed update, these results should be interpreted as finite-budget performance rather than evidence of empirical convergence. Figure 2: Held-out episodic reward distribution. Figure 3: Cumulative training-side communication overhead. V-D Held-Out Reward Performance Fig. 2 evaluates the selected policies on channel realizations not used for training or checkpoint selection. The proposed method attains a mean episodic reward of approximately 57.057.0. Event-Uncompressed and CTDE-MAPPO achieve approximately 53.353.3 and 53.153.1, respectively, while No-Federation-IA-PPO reaches approximately 52.252.2. Strict-Independent-PPO and Periodic-Full attain approximately 50.450.4 and 50.350.3, and Greedy-Queue reaches approximately 46.446.4. Accordingly, the proposed method improves the mean held-out reward by about 6.8%6.8\% relative to Event-Uncompressed, 7.2%7.2\% relative to CTDE-MAPPO, and 9.1%9.1\% relative to No-Federation-IA-PPO. The proposed distribution is shifted to the right and has a comparatively compact interquartile range. This indicates that the gain is present across the held-out realizations rather than being generated by a small number of exceptionally favorable episodes. Nevertheless, episode-level samples generated by the same trained policy are not independent training replicates. Therefore, claims of statistical significance should be based on paired comparisons across the independent training runs, while the boxplot is used as a descriptive representation of held-out variability. V-E Communication Efficiency Fig. 3 reports cumulative training-side communication under the adopted accounting rule. At update 250250, Periodic-Full and Event-Uncompressed require approximately 11.211.2 and 11.011.0 Gbits, respectively, whereas CTDE-MAPPO and the proposed method require approximately 3.653.65 and 2.652.65 Gbits. The proposed method therefore reduces the reported communication by approximately 76%76\% relative to both uncompressed distributed critic-exchange baselines and by approximately 27%27\% relative to CTDE-MAPPO. The small difference between Periodic-Full and Event-Uncompressed shows that the selected event trigger remains active during most training rounds. Consequently, the measured reduction relative to the uncompressed distributed baselines is attributable primarily to layer-wise sparse critic exchange and error feedback, rather than to a large decrease in the number of communication rounds. The experiment therefore demonstrates the effectiveness of compressed exchange at the selected performance-oriented operating point, but does not by itself isolate the gain attributable to event triggering. Such isolation would require a periodic-compressed baseline with the same sparsification budget. Methods without critic coordination have zero federated-model communication, but their lower reward, QoS, and interference efficiency show the performance cost of eliminating collaboration entirely. V-F Long-Term QoS and Mean SINR Fig. 4 compares the held-out QoS metric and mean SINR. Fig. 4 (a) shows that the proposed method obtains a QoS-satisfaction ratio of approximately 0.780.78, closely followed by CTDE-MAPPO at approximately 0.770.77. No-Federation-IA-PPO and Event-Uncompressed achieve approximately 0.740.74 and 0.730.73, while Periodic-Full and Strict-Independent-PPO attain approximately 0.700.70 and 0.690.69. The overlapping error bars of the proposed method and CTDE-MAPPO again indicate comparable QoS performance; the more pronounced benefit of the proposed method appears in the interference-related metrics. In Fig. 4 (b), the proposed method reaches a mean SINR of approximately −2.3-2.3 dB, compared with approximately −4.0-4.0 dB for Event-Uncompressed, −4.8-4.8 dB for CTDE-MAPPO, −5.1-5.1 dB for No-Federation-IA-PPO, −5.6-5.6 dB for Periodic-Full, and −5.9-5.9 dB for Strict-Independent-PPO. Hence, the proposed method improves the mean SINR by approximately 1.71.7 dB relative to Event-Uncompressed and by approximately 2.52.5 dB relative to CTDE-MAPPO. The negative absolute SINR values are consistent with the deliberately interference-limited reuse-11 operating regime; the relevant observation is the substantial relative improvement achieved through coordinated scheduling, power control, and beamforming. V-G User-Rate Distribution Fig. 5 presents the empirical CDF of the per-UE time-average held-out rate. The proposed curve is shifted to the right of those of the learning and heuristic baselines over most of the distribution. The separation is particularly visible in the lower and median portions of the CDF, indicating that the proposed method improves not only the average network outcome but also the rates experienced by comparatively weak UEs. The curves approach one another in the upper tail, suggesting that the principal benefit is improved service to disadvantaged and typical users rather than an isolated increase in the largest UE rates. The vertical line at Rn,mmin=1.9R_n,m =1.9 identifies the long-term target. The rate CDF and the reported QoS bar must be interpreted using explicitly distinct definitions if the latter is computed as a time average of instantaneous UE-slot satisfaction indicators. If the QoS bar is intended instead to denote the fraction of UEs whose time-average rate satisfies the long-term constraint, it must equal 1−F^R¯(Rn,mmin)1- F_ R(R_n,m ) and should be recomputed directly from the same samples used for Fig. 5. This distinction is required to avoid conflating instantaneous service reliability with satisfaction of the long-term average-rate constraint. ((a)) ((b)) Figure 4: Held-out QoS and SINR performance of the selected checkpoints: (a) long-term QoS-satisfaction ratio and (b) mean SINR. Figure 5: Empirical CDF of per-UE time-average held-out rates. Figure 6: QoS–interference operating points of the evaluated methods. V-H Interference Management The interference-per-rate results provide a normalized view of how effectively each policy converts the interference it creates into useful throughput. The proposed method attains approximately 1.78×10−41.78× 10^-4, compared with 1.93×10−41.93× 10^-4 for CTDE-MAPPO, 1.98×10−41.98× 10^-4 for No-Federation-IA-PPO, 2.03×10−42.03× 10^-4 for Event-Uncompressed, and 2.40×10−42.40× 10^-4 for Periodic-Full. These values correspond to reductions of approximately 8%8\%, 10%10\%, 12%12\%, and 26%26\%, respectively. Fig. 6 summarizes the joint QoS–interference operating points. The preferred region is the upper-left corner, representing larger QoS satisfaction and smaller interference cost. The proposed method occupies the most favorable displayed point. CTDE-MAPPO provides nearly the same QoS ratio but at a visibly larger interference cost, whereas No-Federation-IA-PPO and Event-Uncompressed sacrifice both QoS and interference efficiency. The comparison with Strict-Independent-PPO and the greedy baselines further shows that explicit interference awareness alone is insufficient to attain the proposed operating point; collaborative critic learning contributes additional value. V-I Overall Operating Point Taken together, the figures indicate that the proposed method provides the strongest performance–communication operating point among the coordinated learning methods considered. It achieves the largest validation and held-out reward, the highest mean SINR, the smallest interference cost per unit sum rate, and a QoS level comparable to the best centralized-training baseline. Relative to uncompressed distributed critic exchange, these benefits are obtained with an approximately 76%76\% reduction in reported training-side communication. Relative to the non-federated learning baselines, the proposed method uses additional training communication but obtains materially better reward, user-rate, QoS, and interference outcomes. The result should therefore be interpreted as a favorable joint operating point, not as a claim that the proposed method minimizes communication in isolation. VI Conclusion This paper proposed FedCritic-MIMO, a communication-efficient serverless federated critic learning framework for massive-MIMO resource control in open and disaggregated 6G RANs. The framework enables independently deployable cell-level controllers to coordinate scheduling, power allocation, and beamforming without centralized trajectory collection, actor sharing, or parameter-server aggregation. By keeping local actors and personalized critic components private while exchanging only compatible shared critic parameters, FedCritic-MIMO preserves local control autonomy and supports collaboration over the physical interference graph. FedCritic-MIMO combines utility-aware triggering, adaptive layer-wise top-k compression with error feedback, and balanced interference-aware fusion to make peer-to-peer critic collaboration both communication- efficient and radio-aware. We also established conditional finite-time stationarity and consensus guarantees for the balanced shared-critic recursion under a fixed-policy, frozen-target critic-regression model. Simulation results showed that FedCritic-MIMO achieves the most favorable performance–communication tradeoff among the considered baselines, improving throughput, user-rate distribution, SINR, QoS satisfaction, and interference efficiency while substantially reducing critic-model traffic. Future work will consider asynchronous peer coordination, mobility-aware interference graphs, imperfect CSI, and larger heterogeneous deployments. Appendix A Proof of Lemma 1 Proof: If ξnt=0 _n^t=0, then Γnt<τth _n^t< _th. Since 1+αQQ¯nt+αII¯nt≥11+ _Q Q_n^t+ _I I_n^t≥ 1 and nt+1=nte_n^t+1=d_n^t, (45) gives ‖nt+1‖2<τth(‖ψ^nt‖2+ϵtr)\|e_n^t+1\|_2< _th (\| ψ_n^t\|_2+ _tr ). If ξnt=1 _n^t=1, then nt+1=nt−nt(nt)e_n^t+1=d_n^t-C_n^t(d_n^t), and (48) yields ‖nt+1‖22≤(1−δc)‖nt‖22\|e_n^t+1\|_2^2≤(1- _c)\|d_n^t\|_2^2. These two cases prove (54). ∎ Appendix B Proof of Lemma 2 Proof: Because tW^t is doubly stochastic, (N−(t)T)=(I_N-(W^t)^T)1=0. Equation (61) therefore gives t=R^t1=0. Right-multiplying (60) by /N1/N then gives (64). Finally, ‖N−(t)T‖2≤2\|I_N-(W^t)^T\|_2≤ 2 for a symmetric stochastic matrix; hence, 1N‖t‖F2≤4N‖~t+1−^t+1‖F2=4εt2 1NE\|R^t\|_F^2≤ 4NE\| ^t+1- ^t+1\|_F^2=4 _t^2, which proves (65). ∎ Appendix C Proof of Theorem 1 Proof: Let ≜1NTJ 1N11^T, ≜N− _N-J, and t≜tZ^t ^t . Then N−1‖t‖F2=ℰψtN^-1\|Z^t\|_F^2=E_ψ^t. With t≜t−A^t ^t-J, double stochasticity gives t=t=tW^t = W^t=A^t and ‖t‖2≤λW\|A^t\|_2≤ _W. Using t=tR^t =R^t, the disagreement recursion is t+1=(t−ηct)(t)T+t.Z^t+1= (Z^t- _cG^t )(A^t)^T+R^t. (C.1) Let t=[∇F1(ψ1t),…,∇FN(ψNt)]U^t=[∇ F_1( _1^t),…,∇ F_N( _N^t)]. By smoothness, bounded heterogeneity, and (71), 1N‖t‖F2≤c0(L2[ℰψt]+σg2+ζ2) 1NE\|G^t \|_F^2≤ c_0\! (L^2E[E_ψ^t]+ _g^2+ζ^2 ) (C.2) for a constant c0>0c_0>0. Applying Young’s inequality to (C.1), using (65), and choosing ηc≤ηmax _c≤ _ sufficiently small relative to L and 1−λW1- _W, yields constants c1,c2>0c_1,c_2>0 and ρdis∈(0,1) _dis∈(0,1) with 1−ρdis≥c1(1−λW)1- _dis≥ c_1(1- _W) such that [ℰψt+1]≤ρdis[ℰψt]+c2ηc21−λW(σg2+ζ2)+c21−λWεt2.E[E_ψ^t+1]≤ _dis\,E[E_ψ^t]+ c_2 _c^21- _W( _g^2+ζ^2)+ c_21- _W _t^2. (C.3) Since ℰψ0=0E_ψ^0=0, summing (C.3) and using ∑j≥0ρdisj=(1−ρdis)−1 _j≥ 0 _dis^j=(1- _dis)^-1 gives 1T∑t=0T−1[ℰψt]≤C5ηc2(σg2+ζ2)(1−λW)2+C6(1−λW)21T∑t=0T−1εt2, 1T _t=0^T-1E[E_ψ^t]≤ C_5 _c^2( _g^2+ζ^2)(1- _W)^2+ C_6(1- _W)^2 1T _t=0^T-1 _t^2, (C.4) which proves (75). It remains to bound the aggregate objective. By (64), ψ¯t+1=ψ¯t−ηcg¯t ψ^t+1= ψ^t- _c g^t. Define h¯t≜1N∑n=1N∇Fn(ψnt) h^t 1N _n=1^N∇ F_n( _n^t) and t≜h¯t−∇F(ψ¯t)b^t h^t-∇ F( ψ^t). Smoothness and Jensen’s inequality imply ‖t‖22≤L2ℰψt\|b^t\|_2^2≤ L^2E_ψ^t. Moreover, Assumption 1 gives [g¯t∣ℱt]=h¯tE[ g^t _t]= h^t and [‖g¯t−h¯t‖22∣ℱt]≤σav2E[\| g^t- h^t\|_2^2 _t]≤ _av^2. The descent lemma, followed by Young’s inequality, therefore gives, for a sufficiently small ηc _c, [F(ψ¯t+1)]≤ [F( ψ^t+1)]≤ [F(ψ¯t)]−ηc2‖∇F(ψ¯t)‖22 [F( ψ^t)]- _c2E\|∇ F( ψ^t)\|_2^2 +c3ηcL2[ℰψt]+c4Lηc2σav2. +c_3 _cL^2E[E_ψ^t]+c_4L _c^2 _av^2. (C.5) Summing (C.5), using the lower bound F⋆F , dividing by ηcT _cT, and substituting (C.4) yields (74) after absorbing numerical and smoothness constants into C1,…,C4C_1,…,C_4. Together with (C.4), this proves the theorem. ∎ Appendix D Proof of Corollary 1 Proof: Uniformity and independence of JTJ_T imply ‖∇F(ψ¯JT)‖22=1T∑t=0T−1‖∇F(ψ¯t)‖22E\|∇ F( ψ^J_T)\|_2^2= 1T _t=0^T-1E\|∇ F( ψ^t)\|_2^2. Substituting ηc=cη/T _c=c_η/ T and (77) into (74) gives, respectively, (T−1/2)O(T^-1/2), (T−1/2)O(T^-1/2), (T−1)O(T^-1), and (logT/T)O( T/T) for its four right-hand-side terms, proving (78). For fixed ηc _c, taking the limit superior in (74) and using lim supT→∞T−1∑t<Tεt2≤ε¯2 _T→∞T^-1 _t<T _t^2≤ ^2 gives the stated stationarity neighborhood. ∎ References [1] E. Hossain and A. I. Vera-Rivera, “6G cellular networks: Mapping the landscape for the IMT-2030 framework,” IEEE Trans. Technol. Soc., vol. 6, no. 4, p. 377–392, Dec. 2025. [2] M. Na et al., “Operator’s perspective on 6G: 6G services, vision, and spectrum,” IEEE Commun. Mag., vol. 62, no. 8, p. 178–184, Aug. 2024. [3] B. Agarwal et al., “Open RAN for 6G networks: Architecture, use cases and open issues,” IEEE Commun. Surveys Tuts., vol. 28, p. 2881–2924, Secondquarter 2026. [4] M. Polese et al., “Empowering the 6G cellular architecture with open RAN,” IEEE J. Sel. Areas Commun., vol. 42, no. 2, p. 245–262, Feb. 2024. [5] F. Mazzenga and A. Vizzarri, “Time synchronous OFDMA for dense wireless access in open-RAN,” IEEE Commun. Lett., vol. 30, p. 66–70, 2026. [6] W. Chen et al., “Toward standardization of 6G and NextG: Key technologies to enable fundamental enhancements,” IEEE J. Sel. Areas Commun., vol. 44, p. 4333–4365, Mar. 2026. [7] M. Elyasi and A. Vosoughi, “Massive MIMO over correlated fading channels: Multi-cell MMSE processing, pilot assignment and power control,” IEEE Trans. Wirel. Commun., vol. 25, p. 30–46, Jun. 2026. [8] A. Tusha and H. Arslan, “Interference burden in wireless communications: A comprehensive survey from PHY layer perspective,” IEEE Commun. Surveys Tuts., vol. 27, no. 4, p. 2204–2246, Fourthquarter 2025. [9] T. Li et al., “Applications of multi-agent reinforcement learning in future internet: A comprehensive survey,” IEEE Commun. Surveys Tuts., vol. 24, no. 2, p. 1240–1279, Secondquarter 2022. [10] M. Chafii et al., “Emergent communication in multi-agent reinforcement learning for future wireless networks,” IEEE Internet Things Mag., vol. 6, no. 4, p. 18–24, Dec. 2023. [11] A. Kopic, E. Perenda, and H. Gacanin, “A collaborative multi-agent deep reinforcement learning-based wireless power allocation with centralized training and decentralized execution,” IEEE Trans. Commun., vol. 72, no. 11, p. 7006–7016, Nov. 2024. [12] A. E. Matemu, M. Kim, and K. Lee, “Joint scheduling, O-RU association, and power allocation in O-RAN via model-based optimization and PPO-based reinforcement learning,” IEEE Trans. Wirel. Commun., vol. 25, p. 20102–20117, Jul. 2026. [13] H. Zhang, H. Zhou, and M. Erol-Kantarci, “Federated deep reinforcement learning for resource allocation in O-RAN slicing,” in Proc. IEEE Glob. Commun. Conf. (GLOBECOM), Rio de Janeiro, Brazil, 2022, p. 958–963. [14] Y. S. Nasir and D. Guo, “Multi-agent deep reinforcement learning for dynamic power allocation in wireless networks,” IEEE J. Sel. Areas Commun., vol. 37, no. 10, p. 2239–2250, Oct. 2019. [15] N. Naderializadeh et al., “Resource management in wireless networks via multi-agent deep reinforcement learning,” IEEE Trans. Wirel. Commun., vol. 20, no. 6, p. 3507–3523, Jun. 2021. [16] A. Al-Habashna et al., “Decentralized and joint resource allocation, beamforming, and beamcombining for 5G networks with heterogeneous MARL,” IEEE Access, vol. 13, p. 101491–101506, 2025. [17] Y. Chen et al., “Proportional fair resource scheduling for dynamic beyond 5G networks: A distributed hierarchical DRL approach,” IEEE Trans. Mobile Comput., vol. 25, no. 7, p. 10893–10909, Jul. 2026. [18] Y. Ma et al., “Joint beamforming and resource allocation for delay optimization in RIS-assisted OFDM systems: A DRL approach,” arXiv preprint arXiv:2506.03586, 2025. [19] C. Amato, “An introduction to centralized training for decentralized execution in cooperative multi-agent reinforcement learning,” arXiv preprint arXiv:2409.03052, 2024. [20] J. Zhang, “Multi-agent reinforcement learning in wireless distributed networks for 6G,” arXiv preprint arXiv:2502.05812, 2025. [21] E. Eldeeb and H. Alves, “An offline multi-agent reinforcement learning framework for radio resource management,” IEEE Trans. Mobile Comput., vol. 25, no. 1, p. 1137–1150, Jan. 2026. [22] K. Zhang et al., “Fully decentralized multi-agent reinforcement learning with networked agents,” in Proc. 35th Int. Conf. Mach. Learn. (ICML), Stockholm, Sweden, 2018, p. 5872–5881. [23] Z. Chen et al., “Sample and communication-efficient decentralized actor–critic algorithms with finite-time analysis,” in Proc. 39th Int. Conf. Mach. Learn. (ICML), Baltimore, MD, USA, 2022, p. 3794–3834. [24] J. Jiang, K. Su, and Z. Lu, “Fully decentralized cooperative multi-agent reinforcement learning: a survey,” arXiv preprint arXiv:2401.04934, 2024. [25] Y. Tao, J.-C. He, Z.-J. Liu, and S. Yang, “Learn-to-share: A decentralized multi-agent spectrum sharing framework for heterogeneous networks in the 6G era,” IEEE J. Sel. Areas Commun., vol. 44, p. 3490–3506, Jan. 2026. [26] H. B. McMahan et al., “Communication-efficient learning of deep networks from decentralized data,” in Proc. 20th Int. Conf. Artif. Intell. Statist. (AISTATS), Fort Lauderdale, FL, USA, 2017, p. 1273–1282. [27] T. Li et al., “Federated optimization in heterogeneous networks,” in Proc. Mach. Learn. Sys. (MLSys), vol. 2, Austin, TX, USA, 2020, p. 429–450. [28] P. Parhizgar et al., “Federated reinforcement learning for energy-efficient D2D-IoT networks with AoI awareness,” IEEE Open J. Veh. Technol., vol. 6, p. 2828–2841, Sep. 2025. [29] Z. Ji, Z. Qin, and X. Tao, “Meta federated reinforcement learning for distributed resource allocation,” IEEE Trans. Wirel. Commun., vol. 23, no. 7, p. 7865–7876, Jul. 2024. [30] X. Yu et al., “Communication-efficient soft actor–critic policy collaboration via regulated segment mixture,” IEEE Internet Things J., vol. 12, no. 4, p. 3929–3947, Feb. 2025. [31] R. Tan et al., “Pareto actor–critic for communication and computation co-optimization in non-cooperative federated learning services,” IEEE Trans. Mobile Comput., vol. 25, no. 2, p. 1628–1643, Feb. 2026. [32] A. Koloskova, S. U. Stich, and M. Jaggi, “Decentralized stochastic optimization and gossip algorithms with compressed communication,” in Proc. 36th Int. Conf. Mach. Learn. (ICML), Long Beach, CA, USA, 2019, p. 3478–3487. [33] S. P. Karimireddy et al., “Error feedback fixes SignSGD and other gradient compression schemes,” in Proc. 36th Int. Conf. Mach. Learn. (ICML), Long Beach, CA, USA, 2019, p. 3252–3261. [34] A. Farajzadeh and M. Erol-Kantarci, “FedCritic: Serverless federated critic learning-based resource allocation for multi-cell OFDMA in 6G,” arXiv preprint arXiv:2605.21418, 2026. [35] Z. Jiang et al., “Partitioned edge learning over fast fading channels,” IEEE Trans. Veh. Technol., vol. 74, no. 6, p. 8561–8576, Jun. 2025. [36] X. He et al., “Asymptotic analysis of federated learning under event-triggered communication,” IEEE Trans. Signal Process., vol. 71, p. 2654–2667, 2023. [37] Z. Chen et al., “Robust federated learning for unreliable and resource-limited wireless networks,” IEEE Trans. Wireless Commun., vol. 23, no. 8, p. 9793–9809, Aug. 2024.