Paper deep dive
Adversarial Attacks in AI-Driven RAN Slicing: SLA Violations and Recovery
Deemah H. Tashman, Soumaya Cherkaoui
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 4/2/2026, 3:32:53 AM
Summary
This paper investigates the impact of adversarial jamming attacks on AI-driven radio access network (RAN) slicing in Next-Generation (NextG) cellular networks. The authors propose a deep reinforcement learning (DRL)-based framework for resource allocation that incorporates service level agreement (SLA) constraints. They demonstrate that a budget-constrained adversary, using a surrogate DRL agent, can selectively jam slice transmissions to induce significant SLA violations. The study quantifies these violations and analyzes the post-attack recovery behavior of the DRL agent, showing that while the agent is resilient and capable of recovery, the recovery period is slice-dependent and non-negligible.
Entities (6)
Relation Signals (3)
Adversarial Attack → induced → SLA Violations
confidence 95% · Our results indicate that budget-constrained adversarial jamming can induce severe and slice-dependent steady-state SLA violations.
DRL Agent → manages → RAN Slicing
confidence 92% · we study the impact of adversarial attacks on AI-driven RAN slicing decisions... bias deep reinforcement learning (DRL)-based resource allocation
Adversary → targets → URLLC
confidence 90% · the learned attack policy avoids eMBB and instead targets URLLC and mMTC
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Next-generation (NextG) cellular networks are designed to support emerging applications with diverse data rate and latency requirements, such as immersive multimedia services and large-scale Internet of Things deployments. A key enabling mechanism is radio access network (RAN) slicing, which dynamically partitions radio resources into virtual resource blocks to efficiently serve heterogeneous traffic classes, including enhanced mobile broadband (eMBB), massive machine-type communications (mMTC), and ultra-reliable low-latency communications (URLLC). In this paper, we study the impact of adversarial attacks on AI-driven RAN slicing decisions, where a budget-constrained adversary selectively jams slice transmissions to bias deep reinforcement learning (DRL)-based resource allocation, and quantify the resulting service level agreement (SLA) violations and post-attack recovery behavior. Our results indicate that budget-constrained adversarial jamming can induce severe and slice-dependent steady-state SLA violations. Moreover, the DRL agent's reward converges toward the clean baseline only after a non-negligible recovery period.
Tags
Links
- Source: https://arxiv.org/abs/2604.01049v1
- Canonical: https://arxiv.org/abs/2604.01049v1
Trouble viewing inline? Open PDF directly →
Full Text
31,344 characters extracted from source content.
Expand or collapse full text
Adversarial Attacks in AI-Driven RAN Slicing: SLA Violations and Recovery Deemah H. Tashman and Soumaya Cherkaoui Abstract Next-generation (NextG) cellular networks are designed to support emerging applications with diverse data rate and latency requirements, such as immersive multimedia services and large-scale Internet of Things deployments. A key enabling mechanism is radio access network (RAN) slicing, which dynamically partitions radio resources into virtual resource blocks to efficiently serve heterogeneous traffic classes, including enhanced mobile broadband (eMBB), massive machine-type communications (mMTC), and ultra-reliable low-latency communications (URLLC). In this paper, we study the impact of adversarial attacks on AI-driven RAN slicing decisions, where a budget-constrained adversary selectively jams slice transmissions to bias deep reinforcement learning (DRL)–based resource allocation, and quantify the resulting service level agreement (SLA) violations and post-attack recovery behavior. Our results indicate that budget-constrained adversarial jamming can induce severe and slice-dependent steady-state SLA violations. Moreover, the DRL agent’s reward converges toward the clean baseline only after a non-negligible recovery period. I Introduction To support applications such as enhanced mobile broadband (eMBB), massive machine-type communications (mMTC), and ultra-reliable low-latency communications (URLLC), next-generation (NextG) networks introduce major architectural enhancements through radio access network (RAN) slicing mechanisms [1, 2, 3, 4, 5, 6, 7]. In this paradigm, physical radio resources are virtualized and partitioned into resource blocks (RBs) that can be dynamically allocated across slices, enabling slice-aware resource management via intelligent RAN controllers [8, 9, 10]. Service level agreements (SLAs) are crucial to RAN slicing since they provide the anticipated performance guarantees between network operators and tenants in terms of key performance indicators (KPIs), such as latency and throughput [11]. Adhering to the SLA is crucial not just for maintaining satisfactory service, but also for mitigating penalties, safeguarding reputation, and minimizing regulatory concerns, particularly for services that are sensitive to latency and reliability. NextG architectures employ intelligent RAN controllers to translate high-level SLA requirements into granular resource allocation decisions [11]. Moreover, to enable intelligent resource allocation, artificial intelligence (AI) has been increasingly integrated into critical RAN functions, facilitating adaptive decision-making through learning from network measurements and dynamically optimizing slice-level RB assignments [12, 13]. Despite the benefits of AI-driven control in RANs, the reliability of learning-based decision mechanisms is increasingly challenged by adversarial threats [14, 15]. By deliberately manipulating uplink traffic patterns, network measurements, or feedback signals, an adversary can distort the state information observed by AI-based resource allocation agents [16, 17, 18]. Such interference can bias the learned policies toward suboptimal or harmful actions, leading to incorrect slice-level resource block assignments, persistent SLA violations, and, in extreme cases, service disruption across multiple network slices. For instance, reinforcement learning (RL) systems continuously interact with the environment and update their policies based on observed rewards [19, 20]. This adaptive behavior gives rise to properties that can be exploited by an adversary. That is, it becomes feasible to manipulate the reward signal to significantly degrade the learning and decision-making performance of the agent. For example, the authors in [21] evaluated data poisoning and backdoor attacks against deep reinforcement learning (DRL)-based xApps in open RAN (O-RAN), showing that poisoned training data can manipulate control decisions and severely degrade network performance. Moreover, [22] studied adversarial attacks against DRL-based resource allocation in the Near-Real-Time RAN Intelligent Controller (RIC) for vehicular networks, showing that manipulated observations can mislead AI agents and severely degrade RB allocation, data rates, latency, and URLLC reliability. In [23], the authors analyzed adversarial attacks against AI-enabled O-RAN control in multi-operator deployments, showing that a malicious cell can spoof KPIs to deceive traffic steering decisions and gain unfair RB allocations. Furthermore, the authors in [24] explored O-RAN adversarial attacks, including policy infiltration, gradient-based attacks, and signal perturbation. Finally, [25] studied adversarial over-the-air attacks against RL-based network slicing in NextG RANs, where an intelligent jammer selectively disrupts RBs to manipulate the RL reward signal and cause long-term performance degradation. While existing works emphasize the development of specific attacks or defense mechanisms, they do not provide a systematic analysis of the impact of adversarial manipulation on SLA outcomes. Hence, to the best of our knowledge, this is the first work to analyze the impact of adversarial manipulation on AI-driven RAN slicing decisions, with a focus on SLA violations, steady-state degradation, and post-attack recovery behavior. Given this, the main contributions of this work are given as follows: • We propose a slice-level NextG RAN slicing framework that captures RB allocation, slice activity, and SLA constraints for heterogeneous services, including eMBB, URLLC, and mMTC. • We investigate an adversarial DRL–based attack that selectively disrupts slice transmissions under a constrained jamming budget, and analyze its impact on the learning dynamics and decision-making behavior of DRL-driven RAN slicing. • We provide a systematic slice-level evaluation of adversarial effects, quantifying steady-state SLA violations and performance degradation across different slice configurations. • We analyze post-attack recovery behavior after attack removal, highlighting slice-dependent resilience and long-term performance impairment in AI-driven RAN slicing. The remainder of the paper is organized as follows. Section I presents the system model. Section I formulates the optimization problem and describes the proposed DRL-based solution. The adversarial attack model against the victim agent is introduced in Section IV. Simulation results and performance evaluation are provided in Section V, while conclusions are drawn in Section VI I System Model I-A Network Architecture and Slicing Model Fig. 1 illustrates a single-cell downlink RAN in which a gNodeB serves multiple network slices, encompassing eMBB, URLLC, and mMTC. We assume that the time is divided into slots, and within each slot, each slice may submit a transmission request based on a stochastic arrival process. When operational, a slice aggregates traffic from multiple user equipment (UEs) with similar quality of equipment (QoE) requirements, including minimum data rate and service reliability, and is assigned to a priority weight that reflects its relative significance. The gNodeB allocates a limited quantity of RBs to a subset of active slices, adhering to an overall RB capacity limitation. The objective is to optimize long-term system utility while adhering to SLAs. In addition, the system faces an external adversary that selectively disrupts the transmissions of specific slices within a constrained jamming budget, therefore impairing their service performance. Figure 1: The system model. Let =1,2,3K=\1,2,3\ denote the set of network slices corresponding to eMBB, URLLC, and mMTC, respectively. At each time slot t, slice k∈k generates a traffic request with probability pkp_k. We define a binary activity indicator as ak(t)=1,if slice k has an active request at time t,0,otherwise.a_k(t)= cases1,&if slice k has an active request at time t,\\ 0,&otherwise. cases (1) The set of active slices at time t is given by (t)=k∈∣ak(t)=1A(t)=\k a_k(t)=1\. Each active slice is associated with a priority weight wk(t)w_k(t), drawn from a bounded interval, reflecting its relative importance. Moreover, each slice k requires a fixed number of RBs, denoted by FkF_k, to be served in a given time slot. I-B Downlink Data Rate and QoE Constraints For each active slice k∈(t)k (t), the achieved downlink data rate at time t is approximated as [25] Dk(t)=cFk(1−BERk(t)),D_k(t)=c\,F_k\, (1-BER_k(t) ), (2) where c is a constant determined by the modulation and bandwidth configuration, and BERk(t)BER_k(t) denotes the bit error rate of slice k, modeled as a function of the signal-to-noise ratio at user k under the AWGN channel assumption with the adopted modulation and coding scheme. Each slice has a minimum rate requirement RkminR_k . Therefore, a slice satisfies its instantaneous rate constraint if Dk(t)≥Rkmin.D_k(t)≥ R_k . (3) I-C Resource Allocation Constraints The gNodeB has a total of F available RBs per time slot. Let xk(t)∈0,1x_k(t)∈\0,1\ be a binary decision variable indicating whether slice k is served at time t. Given this, the resource allocation must satisfy the following constraint [25] ∑k∈(t)Fkxk(t)≤F. _k (t)F_k\,x_k(t)≤ F. (4) This constraint ensures that the total number of RBs allocated to the set of active and scheduled slices at time slot t does not exceed the gNodeB’s available RB capacity. I-D SLA Formulation and Evaluation Metrics I-D1 Served-Ratio SLA The served ratio captures the fraction of slice transmission requests that are successfully served within a finite observation window, and is used as a slice-level SLA metric. We further define a binary success indicator sk(t)∈0,1s_k(t)∈\0,1\, where sk(t)=1s_k(t)=1 if slice k is successfully served at time t, i.e., it is scheduled (xk(t)=1x_k(t)=1), not jammed, and its achieved rate satisfies the minimum rate requirement; otherwise sk(t)=0s_k(t)=0. For slice k, the served ratio over the window ending at time t is defined as ρk(t)=∑τ=t−W+1tsk(τ)=1∧ak(τ)=1∑τ=t−W+1tak(τ)=1, _k(t)= _τ=t-W+1^tI\s_k(τ)=1 a_k(τ)=1\ _τ=t-W+1^tI\a_k(τ)=1\, (5) where ⋅I\·\ denotes the indicator function and W denotes the length of the sliding observation window (in time slots). We also assume that each slice has a minimum served-ratio requirement of ρkmin _k . I-D2 Average Rate SLA The average achieved data rate of slice k over the same window is given by D¯k(t)=1|k|∑τ∈kDk(τ), D_k(t)= 1|T_k| _τ _kD_k(τ), (6) where kT_k denotes the set of time slots within the window in which slice k is active. I-D3 SLA Satisfaction Indicator A slice is considered to satisfy its SLA if both of the following conditions hold within a sliding observation window: • its average achieved data rate exceeds a predefined minimum rate threshold, and • the fraction of time slots in which its requests are successfully served exceeds a target served-ratio. Hence, the SLA indicator is given as SLAk(t)=[D¯k(t)≥Rkmin∧ρk(t)≥ρkmin].SLA_k(t)=I\! [ D_k(t)≥ R_k \; \; _k(t)≥ _k ]. (7) I Problem Formulation and Proposed DRL-based Solution I-A Problem Formulation The gNodeB aims to dynamically allocate radio resources among active network slices to maximize long-term system utility while maintaining slice-level service reliability. Therefore, we formulate a slice-level resource allocation problem that maximizes the weighted sum of successfully served slice requests, while respecting the SLA constraints given in (7). The optimization is subject to per-slot radio RB capacity constraints and binary scheduling decisions for each slice. Given this, the slice-level resource allocation problem can be formulated as maxxk(t)∑t∑k∈(t)wk(t)sk(t), _\x_k(t)\\; _t _k (t)w_k(t)\,s_k(t), (8) subject to ∑k∈(t)Fkxk(t)≤F,∀t, _k (t)F_k\,x_k(t)≤ F, ∀ t, (9) xk(t)∈0,1,∀k,∀t, x_k(t)∈\0,1\, ∀ k,∀ t, (10) ρk(t)≥ρkmin,∀k,∀t≥W, _k(t)≥ _k , ∀ k,\ ∀ t≥ W, (11) D¯k(t)≥Rkmin,∀k,∀t≥W. D_k(t)≥ R_k , ∀ k,\ ∀ t≥ W. (12) I-B Deep Reinforcement Learning-based Solution To tackle the optimization problem in (8), we employ a DRL framework at the gNodeB. Tabular Q-learning has become obsolete due to the expanding state space resulting from dynamic traffic arrivals and varying slice requirements. Hence, we employ a double deep Q-network (DDQN) architecture to approximate the optimal action-value function. At each time slot t, the gNodeB observes the system state st∈s_t , selects an action at∈a_t , receives a scalar reward rtr_t, and transitions to the next state st+1s_t+1. The state vector captures slice-level information and is defined as st=[F(t),a1(t),w1(t),…,aK(t),wK(t)],s_t= [F(t),\;a_1(t),w_1(t),…,a_K(t),w_K(t) ], (13) where F(t)F(t) denotes the number of available RBs, ai(t)∈0,1a_i(t)∈\0,1\ indicates whether slice i has an active request, and wi(t)w_i(t) is the priority weight associated with slice i, for i∈i . The action space consists of all possible subsets of slices that can be served in a given time slot and is represented as a binary bitmask over slices, =0,1,A=\0,1\^K, (14) where a value of 11 indicates that the slice is selected for service. Infeasible actions (violating the RB capacity) result in zero service in that slot. The reward function is designed to jointly capture throughput maximization and SLA satisfaction. Specifically, the instantaneous reward at time t is defined as rt=∑k∈(t)wk(t)sk(t)−λ∑k∈(1−SLAk(t)),r_t= _k (t)w_k(t)\,s_k(t)\;-\;λ _k (1-SLA_k(t) ), (15) where λ is the SLA penalty coefficient, and SLAk(t)∈0,1SLA_k(t)∈\0,1\ indicates whether slice k satisfies its SLA over the sliding window. The proposed reward formulation incentivizes the gNodeB to balance short-term utility maximization with long-term SLA compliance across heterogeneous network slices. The DDQN framework maintains two neural networks: an online network with parameters θ and a target network with parameters θ−θ^-. The action-value function is updated according to Qθ(st,at) Q_θ(s_t,a_t) ← ← Qθ(st,at)+α(rt+γ Q_θ(s_t,a_t)+α (r_t+γ Qθ−(st+1,argmaxa′Qθ(st+1,a′))−Qθ(st,at)), Q_θ^- (s_t+1, _a Q_θ(s_t+1,a ) )-Q_θ(s_t,a_t) ), where α is the learning rate and γ is the discount factor. This DDQN formulation mitigates the overestimation bias commonly observed in standard DQN. IV Deep Reinforcement Learning–Based Slice-Level Attack The adaptive behavior of the DRL agent heightens the risk of various adversarial attacks. In this paper, we investigate the reward manipulation attack on DRL-based RAN slicing and examine its substantial effects on the agent’s learning and decision-making efficiency. Furthermore, we demonstrate that once the attack is removed, the DRL agent gradually recovers through continued interaction with the environment [26, 27, 28, 29], leading to the restoration of long-term SLA satisfaction. Attack Surface As demonstrated in (13), the state observed by the gNodeB encompasses the available RBs, slice activity indicators, and slice priority weights. The gNodeB internally maintains these quantities, preventing a wireless attacker from directly accessing or manipulating the victim agent’s state or learning process. The attacker can, however, alter the gNodeB’s reward by actively jamming specific slice transmissions [25]. By disrupting successful transmissions, jamming directly affects SLA satisfaction indicators, resulting in diminished rewards and misleading feedback to the learning agent. In addition, when a slice transmission is jammed, the associated request fails to satisfy its quality of service (QoS) requirements. This leads to a negative acknowledgment (NACK) to be sent from the UEs. Limited Jamming Capability We assume a practical adversary with a constrained jamming budget, indicating restrictions on either the finite transmission power or the energy availability. Let B be the maximum number of RBs that the adversary is capable to jam within a specified time interval. It is worth mentioning that the adversary cannot disrupt all slices simultaneously, as each slice employs a predefined quantity of RBs. Hence, the adversary must carefully select which slices to target within this budgetary constraint. Surrogate Attacker Model In an optimal situation, the attacker would precisely identify the slices that the gNodeB will serve and would disrupt them accordingly. This type of knowledge is not available in reality, as the adversary is unaware of the slice weights, scheduling choices, or internal gNodeB states. To mitigate this limitation, we employ a surrogate learning methodology as outlined in [25]. The adversary develops an independent DDQN-based surrogate agent to replicate the behavior of the victim gNodeB. The surrogate does not replicate the victim’s precise strategy; nevertheless, it acquires statistical patterns in slice scheduling and resource availability, thereby facilitating informed attack decisions. Attacker State, Action, and Reward The DRL-based attacker is modeled as follows: • State: Similar to [25], we assume that the attacker observes the available RBs at gNodeB at each time interval by passively observing slice-level transmission and scheduling activity on the radio interface. • Action: The attacker selects a binary action over slices, indicating which slices to jam. The selected slices must adhere to the jamming budget constraint, such that the cumulative RB demand of the jammed slices does not exceed B. • Reward: The attacker receives a reward equals the number of victim slice transmissions that fail due to jamming, as inferred from observed NACKs. The adversary does not decode the NACK contents but only detects their presence, which is feasible since NACK signals are shorter and structurally distinguishable from data transmissions. Learning Algorithm The attacker uses a DDQN to mitigate overestimation bias and enhance learning stability. The surrogate attacker is trained against a static victim policy, enabling it to learn jamming techniques without direct access to the gNodeB’s internal parameters. This DRL-based surrogate attack enables the adversary to systematically diminish slice-level QoS and induce SLA violations, even within stringent jamming budget limitations. V Simulation results In this section, we present the simulation setup for the proposed scenario. Unless otherwise stated, all slice-dependent parameters follow the slice ordering eMBB, URLLC, mMTC. The parameters of the simulation are set to the following values: The minimum rate requirement for each slice is set to 80% of its nominal achievable rate, i.e., Rkmin=0.8cFkR_k =0.8cF_k, F=11F=11, B=5B=5, c=12.59×106c=12.59× 10^6, λ=2λ=2, pk=[0.6,0.4,0.3]p_k=[0.6,0.4,0.3], wk=[0.6,0.4,0.3]w_k=[0.6,0.4,0.3], ρkmin=[0.85,0.95,0.80] _k =[0.85,0.95,0.80], Fk=[5,3,1]F_k=[5,3,1], W=20W=20, γ=0.9γ=0.9, and α=0.1α=0.1. Figure 2: Slice-level average SLA violation rate under no attack and DRL-based adversarial attack Fig. 2 illustrates the steady-state slice-level SLA violation rates under both the no-attack scenario and the proposed DRL-based jamming attack. In the absence of attacks, all slices predominantly meet their SLA criteria, with only negligible violations resulting from stochastic traffic arrivals and temporary resource contention. When the NextG network is under attack, significant SLA violations occur, particularly affecting the URLLC and mMTC slices. However, the impact on eMBB remains minimal. Given the limited jamming budget (B=5)(B=5), and the higher RB demand of eMBB (FeMBB=5)(F_eMBB=5), the learned attack policy avoids eMBB and instead targets URLLC and mMTC, where jamming yields more frequent SLA violations per unit of jamming effort. Additionally, we notice that URLLC and mMTC exhibit nearly identical steady-state SLA violation rates as their combined RB demands remain within the attacker’s jamming budget, allowing both slices to be consistently disrupted and driving their SLA metrics to similar saturation levels. This behavior underscores the attacker’s capacity to carefully manipulate the victim agent’s SLA-aware reward framework by consistently degrading slices with service guarantees that are more susceptible to missed transmissions and accumulated SLA penalties. Fig. 3 illustrates the recovery of slice-level SLA violations following the termination of the DRL-based jamming attack. Throughout the attack phase, all slices exhibit persistently elevated SLA violation rates. This indicates that the attacker can prevent the victim from fulfilling long-term QoE needs. Post-attack, the slices exhibit a distinct and visible recovery pattern. The eMBB slice exhibits the most rapid recovery, with its SLA violation rate swiftly declining to zero. On the other hand, URLLC and mMTC, necessitate a longer recovery period due to their elevated SLA thresholds and window-based assessment, which demand successful resource allocation to proceed before SLA violations can be cleared. The consistent decline in SLA breaches following the removal of the attack indicates that the victim’s DRL policy is not permanently harmed by the attack and can readapt through continious learning. Figure 3: Post-attack SLA recovery behavior per network slice under an RL-based adversarial attack. Fig. 4 illustrates the temporal variation in the victim’s average reward per slot during the proposed DRL-based jamming attack and recovery, contrasted with a clean baseline trained without any adversarial interference. The victim’s reward is greatly reduced during the attack phase. This results from both the immediate transmission failures due to jamming and the accumulated SLA penalties within the reward function. We also compare our approach with the method proposed in [25], whose objective formulation does not explicitly account for SLA constraints. We conclude that the benchmark achieves higher absolute reward values, primarily due to its SLA-unaware objective, which does not penalize persistent violations of service reliability and continuity. In contrast, the proposed SLA-aware formulation jointly captures throughput and long-term reliability requirements, providing a stricter and more realistic assessment of network performance and revealing degradation that would otherwise remain hidden under throughput-only reward metrics. Subsequent to the termination of the attack, there is a significant change in the victim’s policy, which rapidly improves through ongoing engagement with the environment. The restored reward trajectory aligns with the clean baseline, signifying that the RL-based slicing agent may reacquire the ability to allocate resources, even after prolonged exposure to attacks. The temporary gap between the recovery and clean curves immediately following the completion of the attack illustrates the extent of harm that adversarial disruption can inflict; however, the ultimate convergence of the two curves indicates that the attack does not result in permanent damage to the policy. Figure 4: Post-attack recovery of the average reward per slot under an RL-based adversarial attack. VI Conclusions In this paper, we investigated adversarial attacks against a DDQN-based RAN slicing framework designed to maximize resource allocation efficiency for heterogeneous services, namely eMBB, URLLC, and mMTC. The adversary was modeled as an intelligent DRL agent that mimics the victim DDQN policy in order to maximize its jamming impact on slice transmissions. By selectively disrupting slice-level communications, the attacker manipulates the resource allocation process, leading to significant violations of slice-level SLAs. Our results demonstrate the substantial effectiveness of the proposed DRL-based attacker in inducing sustained SLA violations by anticipating the victim agent’s allocation decisions. We further show that the eMBB slice recovers more rapidly than URLLC and mMTC after attack removal, as the learned attack policy primarily targets URLLC and mMTC to greedily maximize the number of disrupted slices under a constrained jamming budget, while largely avoiding the higher-cost eMBB slice. Across all slices, the DRL-based resource allocation agent is ultimately able to reduce SLA violation rates once the attack ceases, indicating a degree of resilience in the learning process. In addition, we observe that adversarial jamming significantly degrades the agent’s average reward due to both transmission failures and accumulated SLA penalties; however, following attack removal, the agent progressively relearns effective allocation strategies and converges back toward the clean, no-attack performance baseline. References [1] S. Novanana et al., “Performance of 5G Slicing With Access Technologies and Diversity: A Review and Challenges,” IEEE Access, vol. 12, p. 170 780–170 802, 2024. [2] A. Filali et al., “Open RAN Slicing for MVNOs With Deep Reinforcement Learning,” IEEE Internet of Things Journal, vol. 11, no. 10, p. 18 711–18 725, 2024. [3] K. M. Naguib et al., “DRL-Driven Edge-Aware Utility Optimization for Multi-Slice 6G Networks,” IEEE Networking Letters, vol. 8, p. 14–18, 2026. [4] P. Keyela et al., “Open RAN Slicing with Quantum Optimization,” in 2025 Global Information Infrastructure and Networking Symposium (GIIS), 2025, p. 1–6. [5] A. Filali et al., “Communication and Computation O-RAN Resource Slicing for URLLC Services Using Deep Reinforcement Learning,” IEEE Communications Standards Magazine, vol. 7, no. 1, p. 66–73, 2023. [6] A. Abouaomar et al., “Federated Deep Reinforcement Learning for Open RAN Slicing in 6G Networks,” IEEE Communications Magazine, vol. 61, no. 2, p. 126–132, 2023. [7] Z. Mlika et al., “Network Slicing with MEC and Deep Reinforcement Learning for the Internet of Vehicles,” IEEE Network, vol. 35, no. 3, p. 132–138, 2021. [8] K. Alam et al., “A Comprehensive Tutorial and Survey of O-RAN: Exploring Slicing-Aware Architecture, Deployment Options, Use Cases, and Challenges,” IEEE Communications Surveys & Tutorials, vol. 28, p. 1637–1678, 2026. [9] M. Sarker et al., “Network Slicing Resource Management in Uplink User-Centric Cell-Free Massive MIMO Systems,” arXiv preprint arXiv:2601.12687, 2026. [10] A. Filali et al., “Dynamic SDN-Based Radio Access Network Slicing With Deep Reinforcement Learning for URLLC and eMBB Services,” IEEE Transactions on Network Science and Engineering, vol. 9, no. 4, p. 2174–2187, 2022. [11] N. M. Yungaicela-Naula et al., “RSLAQ–A Robust SLA-driven 6G O-RAN QoS xApp using deep reinforcement learning,” arXiv preprint arXiv:2504.09187, 2025. [12] Y. Azimi et al., “Applications of Machine Learning in Resource Management for RAN-Slicing in 5G and Beyond Networks: A Survey,” IEEE Access, vol. 10, p. 106 581–106 612, 2022. [13] M. C. Kirana et al., “ML-Enabled Open RAN: A Comprehensive Survey of Architectures, Challenges, and Opportunities,” IEEE Communications Surveys & Tutorials, p. 1–1, 2026. [14] D. H. Tashman et al., “Trustworthy AI-Driven Dynamic Hybrid RIS: Joint Optimization and Reward Poisoning-Resilient Control in Cognitive MISO Networks,” IEEE Transactions on Network and Service Management, p. 1–1, 2026. [15] Y. A. Ergu et al., “Unmasking Vulnerabilities: Adversarial Attacks against DRL-based Resource Allocation in O-RAN,” in ICC 2024 - IEEE International Conference on Communications, 2024, p. 2378–2383. [16] H. Moudoud et al., “Enhancing Open RAN Security with Zero Trust and Machine Learning,” in GLOBECOM 2023 - 2023 IEEE Global Communications Conference, 2023, p. 2772–2777. [17] H. Moudoud et al., “Strengthening Open Radio Access Networks: Advancing Safeguards Through ZTA and Deep Learning,” in GLOBECOM 2023 - 2023 IEEE Global Communications Conference, 2023, p. 80–85. [18] M. C. Kirana et al., “Robust Ensemble Model for Attack Detection in Open RAN,” in GLOBECOM 2025 - 2025 IEEE Global Communications Conference, 2025, p. 4000–4005. [19] D. H. Tashman et al., “Optimizing Cognitive Networks: Reinforcement Learning Meets Energy Harvesting Over Cascaded Channels,” IEEE Systems Journal, vol. 18, no. 4, p. 1839–1848, 2024. [20] D. H. Tashman et al., “Performance Optimization of Energy-Harvesting Underlay Cognitive Radio Networks Using Reinforcement Learning,” in 2023 Int. Wireless Commun. Mobile Comput. (IWCMC), 2023, p. 1160–1165. [21] A. Lacava et al., “How to Poison an xApp: Dissecting Backdoor Attacks to Deep Reinforcement Learning in Open Radio Access Networks,” Computer Networks, p. 111727, 2025. [22] Y. A. Ergu et al., “Efficient Adversarial Attacks Against DRL-Based Resource Allocation in Intelligent O-RAN for V2X,” IEEE Transactions on Vehicular Technology, vol. 74, no. 1, p. 1674–1686, 2025. [23] E. Aizikovich et al., “Rogue Cell: Adversarial Attack and Defense in Untrusted O-RAN Setup Exploiting the Traffic Steering xApp,” arXiv preprint arXiv:2505.01816, 2025. [24] Y. Abera Ergu et al., “RADAR: Robust DRL-Based Resource Allocation Against Adversarial Attacks in Intelligent O-RAN,” IEEE Transactions on Green Communications and Networking, vol. 9, no. 4, p. 2305–2318, 2025. [25] Y. Shi et al., “How to Attack and Defend NextG Radio Access Network Slicing With Reinforcement Learning,” IEEE Open Journal of Vehicular Technology, vol. 4, p. 181–192, 2023. [26] D. H. Tashman et al., “Dynamic Synergy: Leveraging RIS and Reinforcement Learning for Secure, Adaptive Underlay Cognitive Radio Networks,” in 2025 Global Information Infrastructure and Networking Symposium (GIIS), 2025, p. 1–6. [27] D. H. Tashman et al., “Securing Cognitive IoT Networks: Reinforcement Learning for Adaptive Physical Layer Defense,” in 2024 6th International Conference on Communications, Signal Processing, and their Applications (ICCSPA), 2024, p. 1–6. [28] D. H. Tashman et al., “Federated Learning-based MARL for Strengthening Physical-Layer Security in B5G Networks,” in ICC 2024 - IEEE International Conference on Communications, 2024, p. 293–298. [29] D. H. Tashman et al., “Securing Next-Generation Networks against Eavesdroppers: FL-Enabled DRL Approach,” in 2024 International Wireless Communications and Mobile Computing (IWCMC), 2024, p. 1643–1648.