Paper deep dive
Spatio-Temporal Scheduling Prediction Under Backhaul Delay for Resilient Coordinated Beamforming
Prashant Kumar Singh, Shubham Vaishnav, Ahmet Hasim Gökceoglu, Li Wang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 7/10/2026, 5:57:42 AM
Summary
This paper addresses backhaul latency in distributed 5G networks by proposing a two-stage predictive framework that uses a Spectral Temporal Graph Neural Network (StemGNN) to forecast user equipment scheduling states. By replacing stale inter-cell scheduling information with predicted states, the framework integrates with a Signal-to-Leakage-plus-Noise Ratio (SLNR) coordinated beamformer to mitigate performance degradation. Evaluated on a three-cell massive MIMO downlink, StemGNN achieves 87.57% prediction accuracy, outperforming recurrent and Markov baselines, and recovers 57-73% of sum rate loss caused by backhaul delay while improving fairness for cell-edge users.
Entities (10)
Relation Signals (8)
Backhaul Delay → degrades → Coordinated Beamforming
confidence 95% · backhaul latency makes this information stale. Even a single transmission time interval (TTI) of delay can reduce CBF-SLNR performance
StemGNN → predicts → User Equipment Scheduling
confidence 95% · StemGNN predicts future user equipment (UE) scheduling states from delayed historical observations
StemGNN → outperforms → GRU
confidence 90% · outperforming LSTM, GRU, Simple RNN, and Markov chain baselines at all evaluated horizons
StemGNN → outperforms → Markov Chain
confidence 90% · outperforming LSTM, GRU, Simple RNN, and Markov chain baselines at all evaluated horizons
StemGNN → outperforms → LSTM
confidence 90% · outperforming LSTM, GRU, Simple RNN, and Markov chain baselines at all evaluated horizons
StemGNN → recovers → Sum Rate
confidence 90% · recover 57–73% of the sum rate loss caused by one TTI of backhaul delay
Quadriga Urban Micro → usedin → Simulation
confidence 90% · Evaluated on a three-cell massive MIMO downlink... under Quadriga Urban Micro (UMi) channels
SLNR → integratedwith → Coordinated Beamforming
confidence 85% · feeds these predictions into a Signal-to-Leakage-plus-Noise Ratio (SLNR) coordinated beamformer
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Coordinated beamforming in distributed 5G networks relies on the timely exchange of inter-cell scheduling information, but backhaul latency makes this information stale. Even a single transmission time interval (TTI) of delay can reduce CBF-SLNR performance below the uncoordinated baseline, because the precoder suppresses interference toward users that are no longer active. Coordination on stale information is therefore worse than no coordination at all. To address this, we propose a two-stage predictive framework in which a Spectral Temporal Graph Neural Network (StemGNN) predicts future user equipment (UE) scheduling states from delayed historical observations, and the predictions replace stale inputs to the CBF-SLNR precoder. Evaluated on a three-cell massive MIMO downlink with 60 UEs and 64 antennas per base station under Quadriga Urban Micro (UMi) channels and a proportional fair scheduler, StemGNN achieves a mean scheduling prediction accuracy of 87.57%, outperforming LSTM, GRU, Simple RNN, and Markov chain baselines at all evaluated horizons, with gains of up to 7.71% over LSTM at longer horizons where inter-UE structural dependencies dominate over temporal autocorrelation. When integrated into coordinated beamforming, the predictions recover 57-73% of the sum rate loss caused by one TTI of backhaul delay, improving sum rate by 9.58-14.35% over the no-prediction baseline and recovering up to 83% of the Lag-1 fairness loss for cell-edge users, with fairness gains persisting at higher lag values where throughput gains diminish. These results show that treating backhaul latency as a spatio-temporal forecasting problem is an effective approach for robust inter-cell coordination in delay-constrained networks.
Tags
Links
- Source: https://arxiv.org/abs/2607.08454v1
- Canonical: https://arxiv.org/abs/2607.08454v1
Trouble viewing inline? Open PDF directly →
Full Text
47,763 characters extracted from source content.
Expand or collapse full text
11institutetext: Stockholm University, Stockholm, Sweden 11email: prku7110, shubham.vaishnav@dsv.su.se 22institutetext: Huawei R&D, Stockholm, Sweden 22email: ahmet.hasim.gokceoglu1, leo.li.wang@huawei.com Spatio-Temporal Scheduling Prediction Under Backhaul Delay for Resilient Coordinated Beamforming Prashant Kumar Singh Shubham Vaishnav Ahmet Hasim Gökceoglu Li Wang Abstract Coordinated beamforming in distributed 5G networks relies on the timely exchange of inter-cell scheduling information, but backhaul latency makes this information stale. Even a single transmission time interval (TTI) of delay can reduce CBF-SLNR performance below the uncoordinated baseline, because the precoder suppresses interference toward users that are no longer active. Coordination on stale information is therefore worse than no coordination at all. To address this, we propose a two-stage predictive framework in which a Spectral Temporal Graph Neural Network (StemGNN) predicts future user equipment (UE) scheduling states from delayed historical observations, and the predictions replace stale inputs to the CBF-SLNR precoder. Evaluated on a three-cell massive MIMO downlink with 60 UEs and 64 antennas per base station under Quadriga Urban Micro (UMi) channels and a proportional fair scheduler, StemGNN achieves a mean scheduling prediction accuracy of 87.57%, outperforming LSTM, GRU, Simple RNN, and Markov chain baselines at all evaluated horizons, with gains of up to 7.71% over LSTM at longer horizons where inter-UE structural dependencies dominate over temporal autocorrelation. When integrated into coordinated beamforming, the predictions recover 57–73% of the sum rate loss caused by one TTI of backhaul delay, improving sum rate by 9.58–14.35% over the no-prediction baseline and recovering up to 83% of the Lag-1 fairness loss for cell-edge users, with fairness gains persisting at higher lag values where throughput gains diminish. These results show that treating backhaul latency as a spatio-temporal forecasting problem is an effective approach for robust inter-cell coordination in delay-constrained networks. 1 Introduction Modern 5G networks, and emerging 6G ones, use coordinated beamforming (CBF) across multiple base stations (BSs) to reduce inter-cell interference and keep spectral efficiency high [10, 27, 5, 17]. In distributed radio access networks, neighboring BSs exchange channel state and scheduling information over backhaul links. These links have limited bandwidth, are often asynchronous, and frequently add delay [25, 3]. As a result, the view each BS has of the wider network is almost always out of date, while user activity changes much faster: users enter, leave, or change their scheduling state continuously [29, 15]. Standard CBF algorithms such as Weighted Minimum Mean Square Error (WMMSE) optimization [27, 26] assume that each BS has fresh, global channel and scheduling information. Once this assumption breaks, the sum rate drops and service quality becomes uneven across cells. From an edge AI point of view, each BS is an autonomous edge node that has to make real-time, latency-critical decisions using incomplete and partially outdated information about its peers. Backhaul latency, jitter, and asynchronous updates act as a kind of partial fault on the links between nodes: the information a BS needs is not lost outright, it just arrives late and stale. A beamforming controller at the edge therefore needs to stay robust against this kind of information-staleness fault to keep service available and quality of experience consistent across cells. The trust side of this problem matters too. Each BS feeds its beamformer with scheduling and channel reports it receives from neighbors, and there is no built-in way to tell whether an unusual or late-arriving report reflects genuine network behavior, a backhaul fault, or a tampered message injected somewhere along the path. The more directly a BS depends on raw, last-received neighbor data, the larger this trust surface becomes. A learning-based predictor trained on the legitimate temporal patterns of neighbor scheduling reduces this dependence: the BS can rely on its own predicted view of the neighborhood, and treat late-arriving reports as a correction rather than as ground truth. We do not aim to detect adversarial inputs in this work, but the same prediction step that handles delay also shrinks the attack surface exposed to an attacker on the backhaul. Despite a lot of progress on CBF and on machine-learning-assisted interference management, most existing work either assumes ideal information exchange, or treats the spatial and temporal sides of the problem separately. Graph Neural Networks (GNNs) capture the spatial structure of wireless networks well by representing BSs and users as nodes in a graph [35, 16, 11], and recent GNN-based beamforming methods [13, 23, 8, 9] perform strongly but usually assume static user configurations and synchronous information access. On the other side, spectral-temporal forecasting models such as StemGNN [2, 31] capture multivariate temporal dependencies but are not connected to physical-layer beamforming decisions. As a result, no existing framework jointly addresses (i) delayed inter-BS information, (i) a dynamic user population, and (i) coordinated beamforming in a form that fits an edge deployment. In this paper, we recast coordinated beamforming under backhaul delay as a prediction-assisted decision problem that runs locally at each BS. Each edge node uses a spectral-temporal GNN, based on StemGNN [2], to predict the current user-equipment (UE) scheduling state of neighboring cells from past, delayed observations, and feeds these predictions into a Signal-to-Leakage-plus-Noise Ratio (SLNR) coordinated beamformer (see Fig. 1 for an overview). To handle the dynamic user populations seen in real networks, we extend the architecture to support a variable number of UEs and a permutation-invariant representation, so the same model still works as users join, leave, or are reordered. The resulting controller suppresses inter-cell interference even when the backhaul is delayed and user activity is stochastic, and gives a concrete example of a fault-tolerant edge AI pipeline for cooperative wireless inference. Figure 1: Proposed prediction-assisted coordinated beamforming framework. Each base station uses a spectral-temporal GNN to predict the current scheduling state of neighboring cells from delayed backhaul observations, and feeds the predictions into an SLNR-based coordinated beamformer. Contributions. The main contributions of this paper are as follows: • We formulate inter-BS UE scheduling prediction under backhaul delay as a multivariate binary time-series problem, and benchmark a spectral-temporal GNN against recurrent (RNN, LSTM, GRU) and Markov baselines. • We extend StemGNN with a permutation-invariant input and output representation that supports a dynamic, variable number of UEs without retraining, making it suitable for realistic edge deployments. • We integrate the predicted scheduling states into an SLNR-based CBF pipeline and measure the resulting sum-rate gain over conventional beamforming that uses stale information directly, under a range of backhaul delays and dynamic user-activity conditions. 2 Related Work 2.1 Coordinated beamforming and the cost of stale information Massive MIMO and coordinated beamforming have become the standard tools for managing interference in dense multi-cell deployments [22, 17, 20, 1, 10]. The classical algorithms in this area, mainly WMMSE [27] and fractional programming [26], reach near-optimal sum rate when each base station has accurate and timely channel state and scheduling information from all of its neighbors. In practice this is rarely the case. In cloud and distributed RAN architectures, backhaul links are bandwidth-limited and asynchronous, and the inter-cell information used to drive these algorithms is delayed, partial, or both [3, 25]. The repeated matrix inversions and iterative steps in WMMSE and FP also push their computational cost beyond what is reasonable for real-time execution on an edge node such as a base station [20, 22]. Decentralized variants reduce some of this burden, but still require frequent inter-cell signaling, which is affected by the same backhaul delays [10, 19]. From an edge-deployment standpoint, this points to a clear failure mode: when the data feeding the coordination algorithm is late or incomplete, the algorithm produces a confidently wrong decision, and the system has no fallback. 2.2 Machine learning approaches and their edge implications To work around these limitations, a large body of recent work uses machine learning for wireless resource allocation and interference management [7, 33, 12]. The general idea is to replace iterative optimization with a learned mapping from network state to control actions, which lowers inference latency and fits the timing budget of edge execution. Reinforcement learning has been applied to beamforming, power allocation, and inter-cell coordination by framing them as sequential decision problems [21, 24], though stability under fast-changing conditions remains a concern. Supervised deep-learning models that map channel observations directly to beamforming vectors [28] cut complexity further, but tend to ignore the structural dependencies between cells. From a deployment standpoint, these models also raise questions about trust and robustness: a base station that runs a learned model at the edge inherits the integrity of its training data and the inputs it consumes at runtime, and most of these works do not discuss how the model behaves when those inputs are stale, missing, or tampered with. 2.3 Graph and spatio-temporal models Wireless networks have a natural graph structure, with base stations and users as nodes and interference relationships as edges. Graph Neural Networks (GNNs) [16, 11, 30, 35] are a good fit for this structure because their message passing is permutation invariant and respects the connectivity of the network. GNN-based beamforming methods, including heterogeneous variants such as HSTGNN [13, 34] and the approaches in [23, 8, 9], achieve near-optimal sum rate with millisecond-level inference, which is compatible with edge execution. The main limitation across this line of work is the assumption of static, or slowly varying, user configurations. Real cellular traffic does not behave this way: users enter, leave, and change scheduling state on much shorter timescales [5, 29, 15]. On the temporal side, spectral-temporal models such as StemGNN [2] forecast multivariate time series by jointly modeling temporal dependencies and inter-series correlations in the frequency domain. Graph WaveNet [31] combines graph convolution with dilated temporal convolutions for a similar purpose, and DCRNN and STGCN follow related ideas [18, 32]. These models are designed as forecasters, and are not typically connected to a physical-layer decision such as beamforming. 2.4 Positioning of this paper Pulling these threads together, three observations stand out for an edge AI deployment in a distributed RAN. Classical CBF assumes information conditions that the edge does not actually have. ML-based beamforming captures spatial structure but treats user activity as static. Spatio-temporal forecasters capture the temporal side but stop short of the beamforming pipeline. To the best of our knowledge, no existing framework jointly addresses (i) the spatial structure of the network, (i) the temporal evolution of user activity, and (i) the staleness of inter-base-station information, in a form deployable at the edge. This paper targets exactly that intersection by combining a spectral-temporal GNN predictor for delayed neighbor scheduling with an SLNR-based coordinated beamformer, so that each base station can take a beamforming decision without trusting a possibly faulty or stale neighbor report at face value. 3 System Model and Problem Formulation 3.1 Multi-Cell Downlink Model We consider a downlink Massive MIMO network with B base stations (BSs) operating on the same time–frequency resource, where each BS b∈1,…,Bb∈\1,…,B\ has M antennas and serves KbK_b single-antenna user equipments (UEs). The total number of UEs in the network is K=∑b=1BKbK= _b=1^BK_b. We denote the channel vector from BS b′b to UE k served by BS b as b′,b,k∈ℂMh_b ,b,k ^M. In our experiments, these channels are generated with the Quadriga Urban Micro propagation model [5], which captures path loss, shadowing, and multipath fading. At time slot (TTI) t, BS b transmits the linearly precoded signal b(t)=∑k=1Kbb,k(t)sb,k(t),x_b(t)= _k=1^K_bw_b,k(t)\,s_b,k(t), (1) where b,k(t)∈ℂMw_b,k(t) ^M is the beamforming vector for UE k in cell b and sb,k(t)∼(0,1)s_b,k(t) (0,1) is the unit-power data symbol. Each BS is subject to a per-cell transmit power constraint ∑k=1Kb‖b,k(t)‖2≤Pmax _k=1^K_b\|w_b,k(t)\|^2≤ P_ . The signal received at UE k in cell b is yb,k(t)=b,b,kHb,ksb,k+∑j≠kb,b,kHb,jsb,j+∑l≠b∑j=1Kll,b,kHl,jsl,j+ub,k,y_b,k(t)=h_b,b,k^Hw_b,ks_b,k+ _j≠ kh_b,b,k^Hw_b,js_b,j+ _l≠ b _j=1^K_lh_l,b,k^Hw_l,js_l,j+u_b,k, (2) where ub,k∼(0,σu2)u_b,k (0, _u^2) is additive Gaussian noise (time index suppressed for readability). The three sums on the right-hand side are, respectively, the desired signal, the intra-cell interference, and the inter-cell interference. The resulting SINR and achievable rate at UE (b,k)(b,k) are γb,k(t)=|b,b,kHb,k|2∑j≠k|b,b,kHb,j|2+∑l≠b∑j|l,b,kHl,j|2+σu2,Rb,k(t)=log2(1+γb,k(t)). _b,k(t)= |h_b,b,k^Hw_b,k|^2 _j≠ k|h_b,b,k^Hw_b,j|^2+ _l≠ b _j|h_l,b,k^Hw_l,j|^2+ _u^2, R_b,k(t)= _2 (1+ _b,k(t) ). (3) 3.2 User Scheduling and Backhaul Delay In each TTI, BS b runs a proportional-fair scheduler [15, 29] that selects a subset of its UEs for transmission. We represent this decision by the binary scheduling vector b(t)∈0,1Kb,sb,k(t)=1⇔UE (b,k) is active at TTI t.s_b(t)∈\0,1\^K_b, s_b,k(t)=1 (b,k) is active at TTI t. (4) Only active UEs contribute to the interference terms in (2), so the global scheduling state (t)=[1(t);…;B(t)]s(t)=[s_1(t);…;s_B(t)] directly shapes the interference pattern at every UE in the network. In a distributed RAN, each BS shares its scheduling decisions and channel state with its neighbors over a backhaul link. We model the backhaul as introducing a fixed delay of τ≥1τ≥ 1 TTIs. At TTI t, each BS therefore observes only b′(t−τ),b′≠b,s_b (t-τ), b ≠ b, (5) about its neighbors. From an edge-deployment perspective, this delay is a structural property of the link between edge nodes, not an occasional event, so any coordination logic running at the BS has to tolerate it by design. 3.3 Coordinated Beamforming with Stale Information Coordinated beamforming aims to maximize the network sum rate under per-BS power constraints: maxb,k∑b=1B∑k=1KbRb,k(t)s.t.∑k=1Kb‖b,k‖2≤Pmax,∀b. _\w_b,k\\;\; _b=1^B _k=1^K_bR_b,k(t) .t. _k=1^K_b\|w_b,k\|^2≤ P_ ,\;\;∀ b. (6) This problem is non-convex; iterative methods such as WMMSE [27] and fractional programming [26] solve it under the assumption that each BS has timely access to the global channel and scheduling state. Coordinated beamforming via the Signal-to-Leakage-plus-Noise Ratio (SLNR) criterion offers a closed-form per-UE alternative that is well suited to distributed execution at the edge [10]. For UE k in cell b, the SLNR is SLNRb,k(t)=|b,b,kHb,k|2b,kHb(t)b,k+σu2‖b,k‖2,SLNR_b,k(t)= |h_b,b,k^Hw_b,k|^2w_b,k^HQ_b(t)\,w_b,k+ _u^2\|w_b,k\|^2, (7) where the leakage covariance b(t)=∑l≠b∑j=1sl,j(t)=1Klb,l,jb,l,jHQ_b(t)\;=\; _l≠ b\; _ subarraycj=1\\ s_l,j(t)=1 subarray^K_lh_b,l,j\,h_b,l,j^H (8) sums only over UEs that are currently active in neighboring cells. The SLNR-maximizing precoder is the dominant generalized eigenvector of the matrix pair (b,b,kb,b,kH,b(t)+σu2)(h_b,b,kh_b,b,k^H,\,Q_b(t)+ _u^2I), and in practice we use a rank-L approximation of b(t)Q_b(t) to keep complexity within the per-TTI budget of an edge node. The dependence on the current scheduling state l(t)s_l(t) in (8) is the central issue. When the BS has only the stale observation l(t−τ)s_l(t-τ), it can either (i) plug the stale observation straight into (8), which builds the wrong leakage covariance and steers nulls toward UEs that are no longer active while ignoring UEs that became active in the meantime, or (i) infer the current scheduling state from the delayed history. We pursue the second option. 3.4 Prediction-Assisted Formulation Let ℋb′(t)=[b′(t−τ),b′(t−τ−1),…,b′(t−τ−Twin+1)]∈0,1Twin×Kb′H_b (t)\;=\; [s_b (t-τ),\,s_b (t-τ-1),\,…,\,s_b (t-τ-T_win+1) ]\;∈\;\0,1\^T_win× K_b (9) denote the most recent TwinT_win delayed scheduling vectors of BS b′b as observed at a neighboring BS. Each BS applies a learned per-cell predictor fθb′f_ _b to produce a current-time estimate ^b′(t)= 1[fθb′(ℋb′(t))> 0.5]∈0,1Kb′, s_b (t)\;=\;1\! [\,f_ _b \! (H_b (t) )\,>\,0.5\, ]\;∈\;\0,1\^K_b , (10) which replaces b′(t−τ)s_b (t-τ) in (8). The overall design problem then becomes a joint learning-and-control task: learn per-BS predictors fθb′\f_ _b \ that, when wired into the SLNR precoder of (7)–(8), recover as much as possible of the sum rate in (6) at delay τ. Two practical constraints follow from the edge setting. First, the predictor must run locally on the BS at TTI granularity, which rules out any centralized or high-latency solution. Second, the predictor should not depend on a fixed UE identity or fixed ordering, since users enter and leave the cell continuously. We therefore require fθf_θ to be permutation-invariant in the UE dimension, so that the same trained model handles a varying Kb′K_b without retraining. 4 Proposed Spatio-Temporal Scheduling Prediction Framework The proposed framework is a two-stage pipeline (Fig. 1). In Stage 1, each BS trains a spectral-temporal GNN predictor on its own observed history of (delayed) neighbor scheduling. In Stage 2, the trained predictor is wired into the SLNR coordinated beamformer so that the leakage covariance (8) is built from ^b′(t) s_b (t) instead of b′(t−τ)s_b (t-τ). Both stages share the same simulator, channel realizations, and proportional-fair scheduler, so the predictor is trained on the same operating regime in which it is later deployed. 4.1 Spectral-Temporal GNN Predictor We use a StemGNN architecture [2], a spectral-temporal graph neural network originally proposed for multivariate time-series forecasting. The fit is direct: the Kb′K_b scheduling sequences of cell b′b are correlated through the proportional-fair scheduler that produced them, and StemGNN jointly models temporal and inter-series dependencies in the spectral domain. The network has three components. Latent correlation layer. The input window ∈0,1Twin×Kb′X∈\0,1\^T_win× K_b is passed through a GRU [4], and a self-attention layer over the hidden states produces a learned adjacency matrix ∈ℝKb′×Kb′A ^K_b × K_b . This A captures which pairs of UEs tend to be co-scheduled and is learned from data rather than specified a priori. StemGNN block. A Graph Fourier Transform (GFT) using the eigenbasis of A projects the multivariate input into a spectral domain where inter-UE correlations become orthogonal. A Spectral Sequential Cell then applies a Discrete Fourier Transform along the temporal axis of each (now decorrelated) component, processes the frequency representation with a 1D convolution and a Gated Linear Unit, and inverts back to the time domain. A learned graph convolution and an inverse GFT close the block. Two such blocks are stacked with a residual connection, where the second block learns the reconstruction residual of the first. Output layer. A Gated Linear Unit followed by fully connected layers produces per-UE activation logits, which are passed through a sigmoid σ(⋅)σ(·) and thresholded at 0.50.5 as in (10) to yield the binary scheduling prediction. Since scheduling prediction is a binary classification task rather than the regression task in [2], we train the network with the binary cross-entropy loss ℒ(θ)=−1||∑(,)∈∑k=1Kb′[yklogy^k+(1−yk)log(1−y^k)],L(θ)\;=\;- 1|D| _(X,y) _k=1^K_b [y_k y_k+(1-y_k) (1- y_k) ], (11) where y^k=σ(fθ())k y_k=σ(f_θ(X))_k and D is the training set of (history window, target scheduling) pairs. 4.2 Permutation Invariance and Variable User Population To meet the second edge constraint above, the input UEs are treated as nodes of the latent graph rather than as fixed positions in a vector. The latent correlation layer infers A from data on each forward pass, the spectral graph convolution is permutation-equivariant by construction, and the per-UE output head is shared across UE indices. As a result, the same trained model handles a varying number of UEs in cell b′b without retraining: UEs that enter the cell are added as new nodes, and UEs that leave are dropped, with no change to the parameters θ. 4.3 Integration with the SLNR Beamformer At inference time, each receiving BS b runs B−1B-1 copies of the trained predictor, one per neighboring cell, on the delayed history ℋb′(t)H_b (t) it has on hand. The predicted neighbor scheduling vectors ^b′(t) s_b (t) are fed into (8) to build the estimated leakage covariance ^b(t) Q_b(t), and the SLNR precoder of (7) is then computed via a rank-L generalized eigenvalue decomposition. No other change is made to the beamforming algorithm, which keeps the predictor and the beamformer cleanly decoupled and lets the predictor be retrained or swapped without touching the physical-layer code path. In operational terms, each BS keeps producing a beamforming decision at every TTI even if neighbor reports arrive late, and the decision is no longer a direct function of the most recent received message. The predictor acts as a learned model of expected neighbor behaviour, so a single missing, late, or anomalous report cannot, on its own, drive the beamformer into a faulty state. 4.4 Algorithm Algorithm 1 summarizes the complete pipeline. Lines 3–5 cover the initial data-gathering phase, during which the BSs run conventional CBF-SLNR with stale scheduling and accumulate a history tensor ℋ∈0,1B×K×TtotalH∈\0,1\^B× K× T_total. At a fixed TTI ttraint_train, each BS trains its own predictor on its accumulated history (lines 16–19). From ttrain+1t_train+1 onward, the trained predictor replaces the stale neighbor scheduling in the SLNR precoder (lines 7–10). Training uses the Adam optimizer with learning rate 10−310^-3, mini-batch size 3232, gradient clipping at maximum norm 1.01.0, early stopping on validation loss with patience 1515, and a strictly chronological 80/2080/20 train–validation split that avoids any temporal leakage from future to past. Input : Channels H; backhaul delay τ; history window TwinT_win; horizon TtotalT_total; training trigger ttraint_train. Output : Beamformers b,k(t)\w_b,k(t)\; sum rate; edge rate. 1exInitialize history ℋ∈0,1B×K×TtotalH∈\0,1\^B× K× T_total, PF rate histories r¯b(t) r_b(t), predictors ←∅← for t←0t← 0 to Ttotal−1T_total-1 do // Scheduling for each BS b do b(t)←PF(,r¯b(t))s_b(t)← PF(H, r_b(t)); store in ℋ[⋅,⋅,t]H[·,·,t] // Neighbor scheduling estimate if predictors=∅predictors= then ^b′(t)←b′(t−τ) s_b (t) _b (t-τ) for all b′b // stale baseline else for each BS b′b do ℋb′(t)←ℋ[b′,⋅,t−τ−Twin+1:t−τ]H_b (t) [b ,\,·,\,t-τ-T_win+1\,:\,t-τ] ^b′(t)←[fθb′(ℋb′(t))>0.5] s_b (t) 1\! [f_ _b (H_b (t))>0.5 ] end for end if // CBF-SLNR Build ^b(t) Q_b(t) from ^b′(t)b′≠b\ s_b (t)\_b ≠ b via (8) Compute SLNR precoder b,k(t)w_b,k(t) via rank-L GEVD of (b,b,kb,b,kH,^b(t)+σu2)(h_b,b,kh_b,b,k^H,\, Q_b(t)+ _u^2I) Compute rates Rb,k(t)R_b,k(t) via (3); update r¯b(t) r_b(t) // Train predictors once if t=ttraint=t_train then for each BS b′b do Build b′D_b from ℋ[b′,⋅, 0:ttrain]H[b ,\,·,\,0:t_train] with lag τ, window TwinT_win; chronological 80/2080/20 split Train fθb′f_ _b by minimizing (11) with Adam, BCE loss, gradient clipping, early stopping end for predictors←fθb′b′=1Bpredictors←\f_ _b \_b =1^B end if end for return b,k(t)\w_b,k(t)\, sum rate, edge rate Algorithm 1 Two-stage prediction-assisted CBF-SLNR pipeline. 5 Experimental Results 5.1 Experimental Setup We evaluate the proposed framework on a three-cell Massive MIMO downlink network with 60 single-antenna UEs (22, 20, and 18 in cells 0, 1, and 2 respectively), 64 antennas per base station, and an identical proportional-fair scheduler running at every BS. Channel matrices are generated with the Quadriga Urban Micro (UMi) propagation model [5], which captures path loss, shadowing, and multipath fading under realistic urban conditions and is widely used in 5G and 6G system-level evaluation. The configuration follows parameters representative of operational Massive MIMO deployments rather than a stylised academic setup. Two subcarrier configurations are evaluated, 1 SC (narrowband) and 48 SC (wideband), which together cover the operating regimes typical of real deployments. Backhaul delays of τ∈1,3,5τ∈\1,3,5\ TTIs are considered. All models use a 12-step delayed history window (Twin=12T_win=12). StemGNN is benchmarked against LSTM [14], GRU [4], Simple RNN [6], third-order Markov chains, and a Simple Moving Average baseline. Two experiments are conducted. Experiment 1 evaluates standalone scheduling prediction over 2000 TTIs with a 70/20/10 chronological train/validation/test split. Experiment 2 integrates the trained predictors into the live CBF-SLNR pipeline over 2000 TTIs, with model training triggered at TTI 1600 on the scheduling history accumulated under coordinated beamforming conditions, and performance measured over the final 400 TTIs. The relative improvement of StemGNN over a baseline is reported as (VStemGNN−Vbaseline)/Vbaseline×100%(V_StemGNN-V_baseline)/V_baseline× 100\%. Results are aggregated across 18 independent operating conditions, namely three cells with different UE populations, two subcarrier configurations, and three backhaul lags, so the trends below reflect consistent behaviour rather than a single favourable run. 5.2 Scheduling Prediction Accuracy Table 1 reports prediction accuracy at Horizon 1 along with StemGNN’s relative improvement over each baseline. Table 1: Scheduling prediction accuracy (%) at Horizon 1, and StemGNN’s relative improvement over each baseline. Model BS 0 BS 1 BS 2 Mean Improv. (%) Simple Moving Average 66.59 59.36 60.47 62.14 ++40.92 3rd Order Markov 86.00 83.53 86.28 85.27 ++2.70 Simple RNN 86.19 85.05 86.21 85.82 ++2.04 GRU 86.34 85.34 86.33 86.00 ++1.83 LSTM 86.24 84.87 86.87 85.99 ++1.84 StemGNN (proposed) 89.71 85.71 87.30 87.57 — StemGNN reaches 87.57% mean accuracy at Horizon 1 and leads on every cell. It is ahead of the strongest recurrent baseline (GRU) by 1.83 percentage points, ahead of the third-order Markov chain by 2.70, and ahead of the Simple Moving Average by a much larger 40.92. The recurrent baselines cluster tightly between 85.82% and 86.00%, which suggests that gating mechanisms add little over plain recurrence for short-range binary scheduling dynamics. At Horizon 1, the lead of the spectral-graph predictor over a well-tuned recurrent baseline is modest but consistent across cells. Table 2 reports accuracy at Horizons 3 and 5. Table 2: Scheduling prediction accuracy (%) at Horizons 3 and 5, and StemGNN’s mean improvement over each recurrent baseline. Horizon 3 Horizon 5 Model BS 0 BS 1 BS 2 Improv. BS 0 BS 1 BS 2 Improv. Simple RNN 77.52 75.29 77.54 ++6.38% 76.54 72.68 76.58 ++5.40% GRU 77.44 75.04 78.33 ++5.72% 75.43 72.57 76.48 ++5.41% LSTM 76.39 73.47 77.41 ++7.71% 74.39 70.91 75.88 ++7.44% StemGNN 83.00 76.00 82.00 — 80.00 75.00 79.00 — StemGNN keeps a 3 to 6 percentage-point lead over every recurrent baseline at both longer horizons, and the margin grows with horizon. The advantage over LSTM widens from 1.84% at Horizon 1 to 7.71% at Horizon 3, and the advantage over GRU widens from 1.83% to 5.72% over the same range. As short-range temporal autocorrelation decays at longer horizons, the dominant predictive signal shifts to inter-UE co-scheduling structure, and the spectral graph component of StemGNN captures that structure jointly across all UEs while the recurrent models still process each sequence in isolation. 5.3 CBF-SLNR Integration Results Table 3 reports CBF-SLNR performance under stale scheduling with no ML compensation. This setting isolates what happens when the backhaul fault is left uncorrected. Table 3: CBF-SLNR performance under backhaul latency without ML prediction. Sum rate and edge rate are in bits/s/Hz. 48 SC 1 SC Configuration Sum Rate Edge Rate Sum Rate Edge Rate REZF (uncoordinated) 149.40 1.13 198.61 1.37 CBF-SLNR Lag 0 (ideal) 156.99 1.30 257.11 1.87 CBF-SLNR Lag 1 138.93 1.12 205.89 1.51 CBF-SLNR Lag 3 146.17 1.19 222.80 1.64 CBF-SLNR Lag 5 145.40 1.17 223.61 1.37 One TTI of backhaul delay reduces the 48-SC sum rate by 11.5% and actually pushes performance below the uncoordinated REZF baseline (138.93 vs 149.40), which means that running coordinated beamforming on stale inputs is worse than not coordinating at all. At Lag 1, the precoder nulls towards UEs that are no longer scheduled and steers interference towards UEs that are actually active. Lag 3 partially recovers because information several slots old loses its specificity and the precoder degrades towards generic, uncoordinated beamforming, which is harmless by comparison. This is the scenario in which a deployed system needs a fault-tolerance mechanism that activates when fresh information is missing. Table 4 and Fig. 2 report end-to-end CBF-SLNR performance with ML-predicted schedules at Lag 1, the operating point where the gain is largest. Table 4: CBF-SLNR performance with ML-predicted schedules at Lag 1. Sum rate and edge rate are in bits/s/Hz, and Improv. is StemGNN’s relative gain in sum rate over the baseline. 48 SC 1 SC Model Sum Rate Edge Rate Improv. Sum Rate Edge Rate Improv. No prediction 138.93 1.117 ++9.58% 205.89 1.507 ++14.35% LSTM 149.90 1.195 ++1.56% 228.79 1.729 ++2.91% GRU 150.74 1.202 ++0.99% 228.39 1.723 ++3.09% StemGNN 152.24 1.283 — 235.44 1.757 — Ideal (Lag 0) 156.99 1.300 — 257.11 1.871 — Figure 2: CBF-SLNR sum rate versus backhaul latency with ML-predicted schedules for 1 SC (left) and 48 SC (right). StemGNN consistently leads at all lag values. The dashed line marks the ideal Lag 0 upper bound. With the StemGNN predictor in the loop at Lag 1, the 48-SC sum rate climbs from 138.93 to 152.24 bits/s/Hz, recovering 73% of the gap to the ideal Lag-0 upper bound and improving on the no-prediction baseline by 9.58%. In the 1-SC setting, StemGNN recovers 57% of the Lag-1 loss with a 14.35% sum-rate improvement over the no-prediction baseline. The largest gains land at the Lag-1 operating point, which is the same point at which the uncompensated system is most unsafe to deploy. StemGNN also stays ahead of the recurrent baselines under integration: it beats GRU by 0.99% (48 SC) and 3.09% (1 SC), and LSTM by 1.56% (48 SC) and 2.91% (1 SC). The advantage narrows at higher lags as all models converge towards similar accuracy plateaus, which Fig. 2 shows clearly. The edge-rate gains are larger in relative terms than the sum-rate gains, and we consider this one of the more important findings for an edge-availability setting. In the 1-SC configuration, StemGNN raises the edge rate from 1.507 to 1.757 bits/s/Hz at Lag 1, a 16.59% improvement, and recovers 83% of the Lag-1 fairness loss. The improvements persist at Lag 3 and Lag 5 even as the sum-rate benefits start to fade, which says that the framework keeps service quality consistent for the cell-edge users who are most exposed to coordination failure [29]. In an edge deployment, this translates directly into more even quality of experience across cells when the backhaul is degraded. Figure 3: Scheduling prediction accuracy during CBF-SLNR integration (48 subcarriers). All models drop sharply from Lag 1 to Lag 3 and plateau thereafter. StemGNN leads at Lag 1 with 83.88%. Fig. 3 reports prediction accuracy during live CBF-SLNR integration across backhaul latency values. All models drop steeply from Lag 1 to Lag 3, a drop of 21 percentage points for StemGNN (from 83.88% to 62.91%), and plateau thereafter. At Lag 1, StemGNN leads GRU by 2.42 p and LSTM by 4.19 p. The collapse in accuracy at Lag 3 directly explains the shrinking sum-rate benefit visible in Fig. 2. The integrated accuracy at Lag 1 (83.88%) is slightly below the standalone accuracy at Horizon 1 (87.57%). This gap reflects a closed-loop distributional shift: the predictor’s own influence on the beamformer alters the interference environment and subtly shifts the scheduling dynamics away from the patterns it was trained on. This is a known property of any closed-loop learned controller rather than a defect of the model itself. 5.4 Discussion Three findings that stand out are as follows: First, the latency-performance relationship is non-monotonic. Lag 1 is worse than Lag 3 or Lag 5 because the precoder nulls towards users active in the previous slot who may no longer be active, so confident misinformation turns out to be worse than no information at all. The proposed framework directly addresses this failure mode and recovers 57 to 73% of the Lag-1 sum-rate loss with 9.58 to 14.35% improvement over the no-prediction baseline. From a fault-tolerance perspective, this is the worst-case window the edge controller must survive, and it is the window in which the predictor delivers its largest gain. Second, StemGNN’s margin over recurrent baselines grows with prediction horizon, reaching up to 7.71% over LSTM at Horizon 3. This confirms that inter-UE co-scheduling structure becomes the dominant predictive signal once short-range autocorrelation decays. The numbers reported here should be read as conservative estimates. The Quadriga UMi simulator with a proportional-fair scheduler produces scheduling patterns that are more regular than those of a real network with bursty traffic, high user mobility, denser deployments, and many more UEs per cell. In those more realistic conditions, the structural inter-UE dependencies that the spectral graph component captures would carry substantially more predictive value, and we expect the gap over purely temporal recurrent models to widen. Third, accuracy is high at Lag 1 (83.88% for StemGNN) but drops by roughly 21 percentage points at Lag 3 and plateaus thereafter. Two effects compound here: the growing uncertainty of multi-step forecasting, and the closed-loop distributional shift noted above. This accuracy collapse is the main bottleneck of the current framework and is what limits sum-rate recovery at Lag 3 and Lag 5. The shape of this failure profile is itself useful information for an edge deployment: the system is most reliable exactly at the operating point where it most needs to be reliable, and degrades gracefully towards the uncoordinated baseline as latency grows further, rather than collapsing into a faulty state. Across all 18 operating conditions evaluated, the ranking of StemGNN over the recurrent and statistical baselines remains consistent, indicating that the gains reflect a structural property of the framework rather than an artefact of a specific configuration. 6 Conclusions This paper addressed the problem of coordinated beamforming under backhaul latency in distributed 5G networks, framed as a fault-tolerance problem for edge AI controllers that must operate on stale information from their peers. We proposed a two-stage predictive framework in which a spectral-temporal GNN predicts future UE scheduling states from delayed historical observations, and the predictions replace stale inputs to the CBF-SLNR precoder without modifying the core beamforming algorithm. Three conclusions stand out. First, the latency-performance relationship is non-monotonic. One TTI of delay is more harmful than three or five, because the precoder actively suppresses interference toward users that are no longer active. Confident misinformation is worse than no information. Prediction is most valuable at precisely this hardest operating point, where the proposed framework recovers 57 to 73% of the Lag-1 sum-rate loss and brings performance close to the ideal upper bound that assumes zero delay. The framework therefore restores service availability exactly when the system is most exposed to backhaul faults. Second, spatio-temporal graph modelling captures inter-UE co-scheduling dependencies that purely temporal recurrent models cannot. The advantage grows with prediction horizon, which confirms that structural inter-UE correlations become the dominant predictive signal once short-range autocorrelation decays. The gains reported here are conservative, and in larger and more bursty deployments with greater user mobility, the structural advantage of the spectral graph component is expected to widen. Third, the performance benefits fall disproportionately on cell-edge users. The framework recovers 83% of the Lag-1 fairness loss in the 1-SC configuration, which makes it a fairness mechanism for the most exposed users as much as a throughput-recovery tool. For an edge deployment, this means more even quality of experience across cells under the same backhaul-fault conditions. Future work will target the sharp accuracy drop from Lag 1 to Lag 3 through uncertainty-aware prediction, in which the model outputs a confidence estimate that the precoder uses to weight neighbour scheduling states by reliability. This also extends the trust argument from the Introduction by shrinking the surface exposed to late or anomalous reports. Further directions include end-to-end joint training under a task-aware loss, online adaptive learning to handle closed-loop distributional shift, and evaluation on larger network topologies. credits 6.0.1 Acknowledgements The authors thank Huawei Sweden R&D for providing the simulation environment and industrial supervision. This work was carried out within the Master’s thesis programme at the Department of Computer and Systems Sciences, Stockholm University. 6.0.2 Use of LLMs. Claude (Anthropic) was used for language assistance and drafting under author direction; all technical content is the authors’ original work. 6.0.3 This work was conducted in collaboration with Huawei Sweden R&D. The beamforming and scheduling simulation framework contains proprietary components and was used solely for academic research purposes. Authors A. H. Gökceoglu and L. Wang are employed by Huawei. The authors have no other competing interests to declare. References [1] F. Boccardi, R. W. Heath, A. Lozano, T. L. Marzetta, and P. Popovski (2014) Five disruptive technology directions for 5g. IEEE Communications Magazine 52 (2), p. 74–80. External Links: Document Cited by: §2.1. [2] D. Cao et al. (2020) Spectral temporal graph neural network for multivariate time-series forecasting. In Advances in Neural Information Processing Systems (NeurIPS), Vol. 33, p. 17766–17778. Cited by: §1, §1, §2.3, §4.1, §4.1. [3] A. Checko, H. L. Christiansen, Y. Yan, L. Scolari, G. Kardaras, M. S. Berger, and L. Dittmann (2015) Cloud ran for mobile networks—a technology overview. IEEE Communications Surveys & Tutorials. Cited by: §1, §2.1. [4] K. Cho et al. (2014) Learning phrase representations using rnn encoder-decoder. In EMNLP, Cited by: §4.1, §5.1. [5] E. Dahlman, S. Parkvall, and J. Sköld (2018) 5G nr: the next generation wireless access technology. 1st edition, Academic Press, Oxford, United Kingdom. External Links: ISBN 978-0-12-814323-0 Cited by: §1, §2.3, §3.1, §5.1. [6] J. Elman (1990) Finding structure in time. Cognitive Science. Cited by: §5.1. [7] N. Farsad and O. Simeone (2018) Deep learning for wireless physical layer: opportunities and challenges. IEEE Communications Magazine. Cited by: §2.2. [8] J. Ge, Y. Liang, J. Joung, and S. Sun (2020) Deep reinforcement learning for distributed dynamic miso downlink-beamforming coordination. IEEE Transactions on Communications 68 (10), p. 6070–6085. External Links: Document Cited by: §1, §2.3. [9] J. Ge, Y. Liang, L. Zhang, R. Long, and S. Sun (2024) Deep reinforcement learning for distributed dynamic coordinated beamforming in massive mimo cellular networks. IEEE Transactions on Wireless Communications 23 (5), p. 4155–4169. External Links: Document Cited by: §1, §2.3. [10] D. Gesbert et al. (2010) Multi-cell mimo cooperative networks: a new look at interference. IEEE Journal on Selected Areas in Communications 28 (9), p. 1380–1408. External Links: Document Cited by: §1, §2.1, §3.3. [11] J. Gilmer et al. (2017) Neural message passing for quantum chemistry. In ICML, Cited by: §1, §2.3. [12] M. Giordani et al. (2020) Toward 6g networks: use cases and technologies. IEEE Communications Magazine. Cited by: §2.2. [13] C. He, Y. Lu, R. Zhang, Y. Li, B. Ai, and D. Niyato (2025) GNN-enabled coordinated beamforming design for high speed railway communication systems. IEEE Transactions on Mobile Computing, p. 1–13. External Links: Document Cited by: §1, §2.3. [14] S. Hochreiter and J. Schmidhuber (1997) Long short-term memory. Neural Computation. Cited by: §5.1. [15] F. Kelly, A. Maulloo, and D. Tan (1997) Charging and rate control for elastic traffic. European Transactions on Telecommunications. Cited by: §1, §2.3, §3.2. [16] T. Kipf and M. Welling (2017) Semi-supervised classification with graph convolutional networks. ICLR. Cited by: §1, §2.3. [17] E. G. Larsson, O. Edfors, F. Tufvesson, and T. L. Marzetta (2014) Massive mimo for next generation wireless systems. IEEE Communications Magazine 52 (2), p. 186–195. External Links: Document Cited by: §1, §2.1. [18] Y. Li et al. (2018) Diffusion convolutional recurrent neural network for traffic forecasting. ICLR. Cited by: §2.3. [19] Z. Li, T. Gamvrelis, H. A. Ammar, and R. Adve (2022) Decentralized user scheduling and beamforming in multi-cell mimo networks. In ICC 2022 - IEEE International Conference on Communications, Vol. , p. 1980–1985. External Links: Document Cited by: §2.1. [20] L. Lu, G. Y. Li, A. L. Swindlehurst, A. Ashikhmin, and R. Zhang (2014) An overview of massive mimo: benefits and challenges. IEEE Journal of Selected Topics in Signal Processing 8 (5), p. 742–758. External Links: Document Cited by: §2.1. [21] N. C. Luong et al. (2019) Deep reinforcement learning in wireless communications. IEEE Wireless Communications. Cited by: §2.2. [22] T. L. Marzetta, E. G. Larsson, H. Yang, and H. Q. Ngo (2016) Fundamentals of massive mimo. Cambridge University Press, Cambridge. Cited by: §2.1. [23] F. B. Mismar, B. L. Evans, and A. Alkhateeb (2020) Deep reinforcement learning for 5g networks: joint beamforming, power control, and interference coordination. IEEE Transactions on Communications 68 (3), p. 1581–1592. External Links: Document Cited by: §1, §2.3. [24] T. Nguyen et al. (2020) Multi-agent reinforcement learning for wireless networks. IEEE Communications Surveys and Tutorials. Cited by: §2.2. [25] S. Park, O. Simeone, O. Sahin, and S. Shamai (2013) Dynamic coordinated beamforming with limited backhaul capacity. IEEE Transactions on Wireless Communications. Cited by: §1, §2.1. [26] K. Shen and W. Yu (2018) Fractional programming for communication systems—part i: power control and beamforming. IEEE Transactions on Signal Processing 66 (10), p. 2616–2630. External Links: Document Cited by: §1, §2.1, §3.3. [27] Q. Shi, M. Razaviyayn, Z. Luo, and C. He (2011) An iteratively weighted mmse approach to distributed sum-utility maximization for a mimo interfering broadcast channel. IEEE Transactions on Signal Processing 59 (9), p. 4331–4340. External Links: Document Cited by: §1, §2.1, §3.3. [28] H. Sun et al. (2018) Learning to optimize for wireless resource management. IEEE Transactions on Signal Processing. Cited by: §2.2. [29] D. Tse and P. Viswanath (2005) Fundamentals of wireless communication. Cambridge University Press. Cited by: §1, §2.3, §3.2, §5.3. [30] Z. Wu et al. (2020) A comprehensive survey on graph neural networks. IEEE Transactions on Neural Networks and Learning Systems. Cited by: §2.3. [31] Z. Wu et al. (2019) Graph wavenet for deep spatial-temporal graph modeling. In IJCAI, Cited by: §1, §2.3. [32] S. Yan, Y. Xiong, and D. Lin (2018) Spatial temporal graph convolutional networks for skeleton-based action recognition. In Proceedings of the AAAI conference on artificial intelligence, Cited by: §2.3. [33] A. Zappone, M. Di Renzo, and M. Debbah (2019) Wireless networks design in the era of deep learning. IEEE Transactions on Communications. Cited by: §2.2. [34] C. Zhang et al. (2022) Graph neural networks for wireless communications: a survey. IEEE Communications Surveys and Tutorials. Cited by: §2.3. [35] J. Zhou et al. (2020) Graph neural networks: a review of methods and applications. AI Open. Cited by: §1, §2.3.