Paper deep dive
Rate-Aware Quantum-Inspired Trajectory Learning for Interference-Limited Multi-UAV Networks
Khaoula Khaled, Muhammad Afaq, Ali Arshad Nasir, Zeeshan Kaleem
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 7/9/2026, 7:48:42 AM
Summary
The paper addresses the curse of dimensionality in trajectory optimization for interference-limited multi-UAV networks by proposing the Rate-Aware Quantum-Annealed Graph Condensation (RA-QAGC) framework. RA-QAGC decouples global coordination from local control by using quantum-inspired probabilistic annealing to compress the high-dimensional spatial search space into a discrete set of high-throughput waypoints. Decentralized Independent Q-Learning agents then navigate this condensed connectivity graph to maximize weighted sum-rate throughput while maintaining QoS requirements. Simulations demonstrate significant performance gains over existing baseline schemes.
Entities (10)
Relation Signals (10)
KFUPM → affiliateswith → Zeeshan Kaleem
confidence 95% · Department of Computer Engineering, King Fahd University of Petroleum & Minerals (KFUPM), Dhahran 31261, Saudi Arabia
Zeeshan Kaleem → authored → RA-QAGC
confidence 95% · Corresponding author: Zeeshan Kaleem (email: zeeshankaleem@gmail.com).
RA-QAGC → solves → Trajectory Optimization
confidence 95% · To overcome this challenge, we propose the Rate-Aware Quantum-Annealed Graph Condensation (RA-QAGC) scheme
RA-QAGC → targets → Multi-UAV Networks
confidence 95% · Unmanned aerial vehicle (UAV) can provide on-demand, high-capacity connectivity in disaster and normal situation.
RA-QAGC → achieves → Throughput
confidence 90% · Simulation results demonstrate the proposal outperformed over existing schemes by achieving 59.4 Mbps total throughput
RA-QAGC → employs → Independent Q-Learning
confidence 90% · This novel design enables lightweight, scalable, and high-performance decentralized Independent Q-Learning (IQL) agents to perform online continuous trajectory tracking
RA-QAGC → maintains → Quality-of-Service
confidence 90% · RA-QAGC effectively balances network capacity by maintaining quality-of-service (QoS) requirements.
RA-QAGC → →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Unmanned aerial vehicle (UAV) can provide on-demand, high-capacity connectivity in disaster and normal situation. However, it faces a challenge of curse of dimensionality in trajectory optimization, where interference-limited environments and vast search spaces make real-time coordination computationally expensive. To overcome this challenge, we propose the Rate-Aware Quantum-Annealed Graph Condensation (RA-QAGC) scheme, which combines rate-aware graph abstraction with decentralized reinforcement learning to enable scalable, interference-aware UAV coordination. By identifying high throughput locations and guiding UAV trajectory adaptation toward throughput-optimal regions, RA-QAGC effectively balances network capacity by maintaining quality-of-service (QoS) requirements. Simulation results demonstrate the proposal outperformed over existing schemes by achieving 59.4 Mbps total throughput and 23.9 Mbps priority-user throughput, representing gains of approximately 15% and 34%, respectively, over the baseline schemes.
Tags
Links
- Source: https://arxiv.org/abs/2606.25480v1
- Canonical: https://arxiv.org/abs/2606.25480v1
Trouble viewing inline? Open PDF directly →
Full Text
41,655 characters extracted from source content.
Expand or collapse full text
Received X Month, X; revised X Month, X; accepted X Month, X; Date of publication X Month, X; date of current version April, 2026. Digital Object Identifier 10.1109/OJVT.2026.x Rate-Aware Quantum-Inspired Trajectory Learning for Interference-Limited Multi-UAV Networks Khaoula Khaled ∗ , Muhammad Afaq † ,Ali Arshad Nasir ‡ , Zeeshan Kaleem § , Senior Member, IEEE 1∗ Department of Computer Engineering, King Fahd University of Petroleum & Minerals (KFUPM), Dhahran 31261, Saudi Arabia 2† Computer Engineering Department and IRC for Intelligent Secure Systems, King Fahd University of Petroleum and Minerals (KFUPM), Dhahran 31261, Saudi Arabia 3‡ Interdisciplinary Research Center for Communication Systems and Sensing (IRC-CSS), Department of Electrical Engineering, King Fahd University of Petroleum and Minerals (KFUPM), Dhahran 31261, Saudi Arabia 4§ Computer Engineering Department and Interdisciplinary Research Center for Smart Mobility and Logistics, King Fahd University of Petroleum & Minerals (KFUPM), Dhahran 31261, Saudi Arabia Corresponding author: Zeeshan Kaleem (email: zeeshankaleem@gmail.com). The authors at King Fahd University of Petroleum & Minerals (KFUPM) would like to acknowledge the support provided by the Deanship of Research (DoR). ABSTRACT Unmanned aerial vehicle (UAV) can provide on-demand, high-capacity connectivity in disaster and normal situation. However, it faces a challenge of curse of dimensionality in trajectory optimization, where interference-limited environments and vast search spaces make real-time coordination computationally expensive. To overcome this challenge, we propose the Rate-Aware Quantum-Annealed Graph Condensation (RA-QAGC) scheme, which combines rate-aware graph abstraction with decentralized reinforcement learning to enable scalable, interference-aware UAV coordination. By identifying high throughput locations and guiding UAV trajectory adaptation toward throughput-optimal regions, RA-QAGC effectively balances network capacity by maintaining quality-of-service (QoS) requirements. Simulation results demonstrate the proposal outperformed over existing schemes by achieving 59.4 Mbps total throughput and 23.9 Mbps priority-user throughput, representing gains of approximately 15% and 34%, respectively, over the baseline schemes. INDEX TERMS Unmanned Aerial Vehicles (UAVs), quantum-inspired computing, trajectory optimization, independent Q-learning, reinforcement learning. I. Introduction T HE integration of Unmanned Aerial Vehicles (UAVs) into next-generation wireless networks represents a paradigm shift toward providing resilient, on-demand three- dimensional connectivity during disaster situation to enable emergency communications. Moving beyond traditional two- dimensional terrestrial coverage, the deployment of UAVs serves as a core pillar for next generation networks [1]. Unlike ground-based communication systems, UAV-enabled networks offer flexibility for rapid deployment in dynamic environments [2]. The major limiting factor in UAV deployment for tra- jectory optimization and network resource allocation is the high computational complexity [3]. As the network size increases, the number of possible network configurations grows combinatorially, significantly elevating the problem and leading to a severe curse of dimensionality [4]. To overcome those challenges, in literature Deep Rein- forcement Learning (DRL) frameworks are widely explored for their adaptability to dynamic environments [5]. However, they suffer from slow convergence rates in massive state- action spaces, where agents frequently get trapped exploring high-dimensional suboptimal actions. Conversely, classical This work is licensed under a Creative Commons Attribution 4.0 License. For more information, see https://creativecommons.org/licenses/by/4.0/ VOLUME 00, 20261 arXiv:2606.25480v1 [cs.MA] 24 Jun 2026 Khoula et al.: Quantum-Aware Trajectory Optimization geometric clustering approaches, such as K-means, have been adopted to compress the environmental state space [6]. Nevertheless, their ability to meet QoS requirements is limited, as they neglect the dynamics of real-time wireless channel conditions. Moreover, frameworks that jointly opti- mize trajectory design, power allocation, and user association lead to highly non-convex optimization problems, posing substantial computational challenges. Existing schemes in the literature addressed the energy efficiency maximization problem by decomposing multi-variable formulations into sequential convex subproblems evaluating trajectory, power allocation, and time-slot assignments [7]. Similarly, Suc- cessive convex approximation and block coordinate descent techniques have also been extensively deployed in wireless- powered communication networks [8], however, these ap- proaches require intensive, centralized iterative computations that scale poorly with network size. In literature, clustering based approaches has also been adopted to compress large user spaces into discrete UAV service zones [9]. These schemes still have a challenge even after simplifying the spatial problem, as they remain channel-blind. They calcu- late spatial centroids based strictly on Euclidean distances, completely ignoring vital wireless channel parameters and interferences. Despite of their success, traditional tabular and standard deep multi-agent reinforcement learning (MARL) frame- works encounter a severe curse of dimensionality as the network expands, demanding decentralized architectures to maintain real-time decision-making capabilities [2]. They also have limitation as they struggle with memory constraints and fail to generalize across unseen states when facing continuous high-dimensional actions. To break the scalability barriers of classical optimization and deep MARL search spaces, quantum computing paradigms, specifically Quantum Annealing (QA) and Quantum-Inspired Optimization (QIO) have emerged as highly efficient alternatives for contin- uous and combinatorial problems. For instance, QA has been successfully applied to satellite communication systems to resolve beam placement and frequency assignments by formulating hybrid quantum-classical Ising pipelines that outperform standard commercial optimization solvers [10]. In aerial networks, quantum annealing (QA)-based ap- proaches have been investigated for sum-rate maximization through the joint optimization of user clustering, subchan- nel assignment, and power allocation [11]. Similarly, [12] addressed joint UAV trajectory design and radio resource allocation by formulating the problem as a Markov decision process and employing a non-iterative cooperative opti- mization strategy to obtain high-quality solutions. Although these studies apply QIO or integrate quantum concepts into reinforcement learning frameworks, their focus remains on trajectory optimization and resource allocation. The use of QIO for reducing the state space prior to learning has received limited attention. In particular, existing works do not exploit quantum-inspired probabilistic annealing together with a rate-aware condensation objective to construct a com- pact representation of the search space, which can improve the scalability of multi-agent reinforcement learning for trajectory optimization. In [13], the UAV navigation problem in cellular-connected networks was formulated as an MDP to jointly min- imize flight delay and communication outage duration. A DRL-based framework was proposed to optimize the UAV trajectory in complex urban environments, while a quantum-inspired experience replay mechanism improved learning efficiency through prioritized sampling. Simula- tion results demonstrated superior performance compared with conventional optimization and existing DRL-based ap- proaches. Moreover, the authors in [14] proposed a Lay- erwise Quantum-Based Deep Reinforcement Learning (LQ- DRL) framework to address large-scale continuous optimiza- tion problems by integrating quantum embedding with deep reinforcement learning. The method jointly optimizes UAV trajectory, user grouping, and power allocation to maximize energy efficiency while satisfying QoS requirements. Results showed that LQ-DRL achieved higher rewards and lower training loss than conventional DRL approaches, with perfor- mance improving as the number of quantum layers increased. Similarly, in our previous work, we proposed quantum- driven state reduction for optimizing UAV trajectory that significantly reduced the outage probability [15]. The aforementioned literature successfully adopted QIO for trajectory optimization targeting various QoS metric but none of these existing frameworks utilize the powerful global search capabilities of quantum annealing to solve the environmental state-space reduction problem. Moreover, the existing schemes proposes channel-blind geometric cluster- ing that ignores co-channel interference to reduce the high- dimensional search space, resulting in slow convergence. To overcome these limitations, we propose Rate- AwareQuantum-AnnealedGraphCondensation(RA- QAGC) scheme, where a rate-aware condensation cost func- tion J is proposed that explicitly characterizes the signal- to-interference-plus-noise ratio (SINR) dynamics. Moreover, we introduce a global, quantum-inspired probabilistic an- nealing mechanism guided by J to prune and compress the high-dimensional deployment space into high-reward discrete waypoint candidate setC. This discrete set explicitly maps the optimal spatial coordinates corresponding to high- rate, low-interference corridors. Finally, to reduce environ- mental complexity, we decouple global coordination from local control. This novel design enables lightweight, scalable, and high-performance decentralized Independent Q-Learning (IQL) agents to perform online continuous trajectory tracking efficiently while mitigating the curse of dimensionality. I. System Model We consider the uplink of a multi-UAV-assisted wireless network spanning a continuous terrestrial region A ⊂ R 2 . The network comprises a set N = 1,...,N of N 2VOLUME 00, 2026 Unmanned Aerial Vehicles (UAVs) acting as Aerial Base Stations (ABSs) to service a set K = 1,...,K of K stationary ground users (GUs). The system operates under a full-frequency reuse scheme across a total bandwidth of B Hz, resulting in an interference-limited network environment. Each ABS n ∈ N flies at a constant altitude h n . At any discrete time slot t ∈ 1,...,T, where T denotes the finite mission horizon, the 3D coordinate vector of ABS n is defined as u n [t] = [x n [t],y n [t],h n ] ⊤ , where q n [t] = [x n [t],y n [t]] ⊤ ∈A denotes its time-varying horizon- tal position. The stationary GUs possess fixed coordinates w k = [x k ,y k , 0] ⊤ , ∀k ∈ K. To implement QoS, K is partitioned into two mutually exclusive subsets: priority users (K pr ) requiring strict data rate guarantees, and normal users (K nr = K pr ) processing standard best-effort traffic. To mitigate the curse of dimensionality inherent in continuous trajectory optimization, the spatial search domain is mapped to a finite set of M discrete candidate waypoints (centroids), precomputed via the RA-QAGC framework as C = c m ∈ R 2 M m=1 . Consequently, the horizontal positioning of each agent must satisfy q n [t] ∈ C, ∀n ∈ N,∀t. Spatial state transitions between successive time slots q n [t] → q n [t + 1] are governed by a global connectivity graph G = (C,E ), where a directional transition edge (c m , c m ′ ) ∈ E exists if and only if: ∥c m − c m ′ ∥ 2 ≤ v max ∆t, where v max is the maximum horizontal velocity of the UAV, and ∆t is the discrete slot duration. The time-varying Euclidean distance d k,n [t] and the cor- responding elevation angle θ k,n [t] between terrestrial GU k and flying ABS n are expressed respectively as: d k,n [t] = ∥u n [t] − w k ∥ 2 , θ k,n [t] = arcsin h n d k,n [t] . Let θ ◦ k,n [t] = 180 π θ k,n [t] represent the elevation angle in degrees. The probability of establishing a Line-of-Sight (LoS) link follows a standard sigmoidal logistic distribution as P LoS,k,n [t] = 1 1+a exp ( −b [ θ ◦ k,n [t]−a ]) , where a and b are environmental con- stant parameters dependent on the urban topology, and the corresponding Non-LoS (NLoS) probability is P NLoS,k,n [t] = 1−P LoS,k,n [t]. The aggregate large-scale effective path loss (in dB) is represented as L eff,k,n [t] =K 0 + 10α log 10 (d k,n [t])+ χ LoS P LoS,k,n [t] + χ NLoS P NLoS,k,n [t], (1) where K 0 =20 log 10 (4πf c /c) represents the free- space path loss at a one meter reference distance for carrier frequency f c . Here, α denotes the environment- specific path loss exponent, and χ LoS ,χ NLoS are state- dependentexcesspathlossvariablesassignedto LoS and NLoS conditions. Compounding large-scale attenuationwithexponentialsmall-scaleRayleigh fading variance |f k,n [t]| 2 ∼ exp(1) yields the total linear channel gain: g k,n [t] = 10 − L eff,k,n [t] 10 · |f k,n [t]| 2 . Tomaintainenergyefficiency,GUsimplementa fractional open-loop power control mechanism. The uplinktransmissionpowerofgrounduser kis dynamically bounded in accordance with: P tx,k [t]= min P max , P 0 + α OL L eff,k,ˆn k [t] [t] + 10 log 10 (N RB ) , where P max is the maximum hardware transmission power threshold, P 0 is the target received power spectral density, α OL ∈ [0, 1] is the path loss compensation factor, and N RB denotes the number of allocated resource blocks. Terrestrial users associate with ABS that provides the maximum average received signal power computed as ˆn k [t] = arg max n∈N (p tx,k [t]· g k,n [t]), where p tx,k [t] is the linear equivalent wattage of the logarithmic value P tx,k [t]. Under a full frequency reuse, the cumulative co-channel inter-cell interference power I n,k [t] = P j∈K,j̸=k p tx,j [t] · g j,n [t] experienced at ABS n, and the resulting received SINR γ k [t] for GU k are γ k [t] = p tx,k [t]·g k,ˆn k [t] [t] σ 2 +I ˆn k [t],k [t] , where σ 2 represents the additive white Gaussian noise (AWGN) ther- mal power. we assume an equal-sharing intra-cell bandwidth allocation policy among associated users, the achievable uplink data rate for GU k associated with its target ABS is defined as R k [t] = B |K ˆn k [t] [t]| · log 2 (1 + γ k [t]), where |K n [t]| = P j∈K ⊮ˆn j [t] = n represents the instantaneous user association cardinality of ABS n, evaluated using the indicator function ⊮·. To enforce priority, the aggregate network utility T [t] is formulated as a weighted sum-rate as T [t]= w pr P k∈K pr R k [t] + w nr P k∈K nr R k [t], where w pr and w nr denote the priority and normal tier weights respectively, satisfying w pr ≫ w nr . The primary network objective is to maximize this utility over the global mission horizon. I. Problem Formulation In this section, we formulated the weighted sum-rate maxi- mization problem by jointly optimizing the UAV trajectory and resource allocation over a finite operational mission horizon T , represented as max q n [t], p tx,k [t] T X t=1 w pr X i∈K pr R i [t] + w nr X j∈K nr R j [t] (2) s.t. X min ≤ x n [t]≤ X max , Y min ≤ y n [t]≤ Y max , (3) ∥q n [t + 1]− q n [t]∥ 2 ≤ v max ∆t, ∀n∈N, t < T, (4) q n [t]∈C, ∀n∈N, ∀t,(5) (q n [t], q n [t + 1])∈E, ∀n∈N, t < T,(6) p min ≤ p tx,k [t]≤ p max , ∀k ∈K, ∀t,(7) X i∈K n [t] B i [t]≤ B, ∀n∈N, ∀t,(8) R i [t]≥ R i,min [t](9) The optimization problem in (2) is supported by various constraints. Here, Constraint (3) specifies that the ABS horizontal trajectories must remain strictly bounded within the designated terrestrial deployment region A. Kinematic step bounds are enforced by (4). Constraints (5) and (6) enforce discrete waypoint restrictions and graph topology VOLUME 00, 20263 Khoula et al.: Quantum-Aware Trajectory Optimization adherence during the trajectory tracking phase. Here, the ABS positions are restricted to the pre-screened condensed waypoint matrix C, and their spatial state transitions must map across connected geometric edges defined within the valid physical edge domain E of the graph G C . The uplink transmit power should be constrained to the limits as defined in (7). IV. Proposed Two-Stage RA-QAGC Optimization Framework To reduce the computational intractability and exponential state-space explosion (O(A N )) of optimizing joint multi- UAV trajectories, this paper proposed RA-QAGC frame- work. The framework decoupled the problem into a se- quential two-stage pipeline: (i) offline global state-space compression via quantum-inspired optimization, and (i) online localized trajectory tracking handled by decentralized MARL. To evaluate the topological fitness of any multi-UAV spa- tial deployment layout, we first construct a communication- centric, rate-aware cost objective function J (q n ) that maps physical wireless channel constraints and co-channel interference characteristics into an energy minimization land- scape: J (q n ) =− (w pr R pri + w nr R nr ) + λ· ̄ I dB (10) where R pri = P i∈K pr R i and R nr = P j∈K nr R j represent the aggregate capacity realizations for priority and standard user tiers, respectively. The term ̄ I dB = 1 K P K i=1 10 log 10 (I ˆn i ,i + ε) quantifies the time-averaged geometric system interfer- ence expressed in decibels, balanced by an insulation con- stant ε = 10 −12 against evaluation faults, and scaled by a multiplier λ = 0.5. Unlike distance-based geometric clustering techniques like K-means that minimize Euclidean centers, inverting the network throughput as a negative energy penalty prevents independent agents from clustering tightly over high-density user areas, thereby suppressing severe co-channel self-interference loops. To minimize this non-convex cost function without be- coming trapped in sub-optimal configurations, a configura- tion state s = [s 1 ,s 2 ,s 3 ] ⊤ tracks coordinate indices pointing to a pre-screened candidate grid N cand = 200. A candidate state s ′ is proposed by selecting n swap = 2 random agents and modifying their positions: s ′ n = ( randi(N cand ), n∈R swap s n ,otherwise. (11) This state transition exploration is governed by an exponen- tial annealing cooling trajectory: T κ+1 = ρ· T κ , T 0 = 1.0(12) where the cooling index ρ = 0.9985 sustains the exploration behavior over a terminal iteration boundary of I max = 3000, reaching a terminal temperature floor of T I max ≈ 0.0119. The state transition couples classical Metropolis metrics with a quantum-inspired probabilistic floor: P accept = max exp − ∆J T , p tunnel · T T 0 (13) where ∆J = J (s ′ ) − J (s). By embedding a baseline tunneling constant p tunnel = 0.05, the system preserves non- zero exploration mechanics through high cost-energy barriers as T → 0. This global search routine ultimately isolates a refined matrix of discrete spatial centroids, constituting the condensed high-reward waypoint candidate set C = c m ∈ R 2 M m=1 . The discrete waypoint matrix C obtained will be used as an input for the trajectory optimization. By using these optimized coordinates to restrict the physical positions of the ABS (q n [t] ∈ C), the continuous search domain is transformed into a lightweight, discrete state-action space governed by a global connectivity graph G = (C,E ). This allows decentralized Independent Q-Learning (IQL) agents to operate with high efficiency. Each agent updates its localized tabular action-value space Q n (s,a), where the local state maps directly to the active waypoint index in C, and the action space is confined to valid edges in E , via the temporal difference model: Q n (s,a)←Q n (s,a) + α r[t]+ γ max a ′ ∈N (s ′ ) Q n (s ′ ,a ′ )− Q n (s,a) (14) where α = 0.25 represents the learning rate, γ = 0.95 is the discount variable. To solve (2) via RL, the multi-UAV trajectory execution is modeled as a discrete-time multi-agent MDP characterized by the standard tuple (S,A,P,R,γ). State space S vector s[t] = [s 1 [t],...,s N [t]] ⊤ ∈ S defines the joint spatial configurations, where index s n [t] maps directly to candidate waypoint c s n [t] ∈ C. Each agent selects its next target waypoint destination restricted by localized spatial connec- tivity graphs as a n [t] ∈ A n (s n [t]) = m ′ ∈ 1,...,M : (c s n [t] , c m ′ )∈E. The joint multi-agent action vector is de- fined as a[t] = [a 1 [t],...,a N [t]] ⊤ . Assuming deterministic kinematic execution of physical actions, the next state is an identity mapping of the chosen action vector: s[t+1] = a[t]. To maximize global spectral performance, the scalar reward r[t] directly mirrors the instantaneous network utility as r[t] =T [t] = w pr R pri [t] + w nr (R total [t]− R pri [t]).(15) We adopt a decentralized IQL framework where each agent n ∈ N maintains an independent action-value function Q n (s n ,a n ) over localized features. Here, using a unified global reward r[t] ensures mutual cooperation despite the non-Markovian environmental shifts typical of concurrent agent updates. This decentralized formulation preserves a lean computational profile ofO(M 2 ) per agent, successfully circumventing the O(M N ) dimensional bottleneck of cen- tralized joint-action configurations. 4VOLUME 00, 2026 TABLE 1: Network Simulation and Channel Parameters ParameterSimulation Configuration Terrain area A1000× 1000 m 2 ABS count N3 ABS altitude h n 100 m GU count K100 Priority ratio K pr 30% Initial candidates N 0 1000 random points Target received power P 0 -90 dBm Number of RB N RB 100 Waypoint set size M100 centroids Carrier freq f c / BW B2.0 GHz / 20 MHz Environment params (a,b)Urban (b 1 = 0.1,b 2 = 1) Path loss exponent α2.5 Excess attenuationχ LoS = 0, χ NLoS = 20 dB Small-scale fadingRayleigh: |f k,n | 2 ∼ exp(1) Noise floor σ 2 −90 dBm SINR threshold γ th −15 dB GU Tx power P tx 20 dBm Q-learning episodes / exploration 5000 / ε-greedy decaying The proposed RA-QAGC framework addresses the com- plexity of multi-UAV trajectory optimization by first reduc- ing the size of the search space through rate-aware graph condensation. A large set of candidate UAV locations is eval- uated using a cost function that balances network through- put and interference levels. A quantum-inspired annealing procedure then explores different deployment configurations and identifies a set of high-quality waypoint locations while avoiding poor local solutions. These selected waypoints are used to construct a connectivity graph that satisfies UAV mobility constraints. Based on this condensed graph, each UAV independently learns its movement policy using Independent Q-Learning (IQL). During operation, UAVs select neighboring waypoints through an ε-greedy strategy and update their Q-values according to a global reward that reflects the weighted throughput of both priority and regular users. By combining intelligent state-space reduction with decentralized learning, RA-QAGC enables efficient trajectory planning, improves network throughput, and main- tains quality-of-service requirements in interference-limited UAV networks. Key steps of the proposal is summarized in Algorithm 1. V. Simulation Results and Discussion To evaluate the performance of the proposed RA-QAGC framework, multi-UAV communication network is simulated using a discrete-event execution model within the MATLAB environment over a 1000 × 1000 m 2 deployment region. Physical network parameters and environmental variables are summarized in Table 1. Figure1 presents the total system throughput and priority- users rate achieved by the proposed RA-QAGC framework against four state-of-the-art baseline schemes random, K- means, Graphically condensed (GC)-K-means, GC-SNR positioning (GC-SNRP). The performance improvement of the RA-QAGC methodology is because of quality-aware network conditioning. Unlike the traditional GC-SNRP ap- proach that evaluates isolated signal power ratios, RA- Algorithm 1 RA-QAGC: Rate-Aware Quantum-Inspired Graph Condensation and Multi-UAV Trajectory Optimiza- tion Require: UAV set N , user set K, initial candidate points N 0 = 1000, pre-screened candidates N cand = 200, condensed waypoints M = 100, cooling factor ρ = 0.9985, initial temperature T 0 = 1.0, maximum itera- tions I max = 3000, tunneling constant p tunnel = 0.05, learning rate α = 0.25, discount factor γ = 0.95 Ensure: Condensed waypoint setC, condensed graphG C = (C,E ), learned Q-tables Q n n∈N 1: Generate N 0 random candidate locations and pre-screen the top N cand points using the rate-aware criterion 2: Initialize a configuration state s and evaluate J (s) =− w pr R pri + w nr R nr + λ ̄ I dB 3: for k ← 0 to I max − 1 do 4:Select n swap = 2 UAV agents and propose a new state s ′ by assigning random positions from the con- densed candidate set 5:Compute ∆J ← J (s ′ )− J (s) 6:Compute the acceptance probability P accept ← max exp − ∆J T k , p tunnel T k T 0 7:Accept s ′ → s with probability P accept 8:Update temperature: T k+1 ← ρT k 9: end for 10: Extract the condensed waypoint set C =c m M m=1 from the optimized state 11: Build the condensed connectivity graph G C = (C,E ) (c m ,c m ′ )∈E ⇐⇒ ∥c m − c m ′ ∥ 2 ≤ v max ∆t 12: for each UAV n∈N do 13:Initialize an independent Q-table Q n (s n ,a n ) over waypoint indices and valid neighbor actions 14: end for 15: for t← 1 to T do 16:for each UAV n∈N do 17:Select action a n [t] using ε-greedy exploration over A n (s n [t]) =m ′ : (c s n [t] ,c m ′ )∈E 18:end for 19:Execute the selected actions and observe the reward r[t] = w pr R pri [t] + w nr R total [t]− R pri [t] 20:for each UAV n∈N do 21:Update Q-table: Q n (s n ,a n ) ← Q n (s n ,a n ) + α h r[t] + γ max a ′ ∈A n (s ′ n ) Q n (s ′ n ,a ′ )− Q n (s n ,a n ) i 22:end for 23: end for 24: return C, G C , and Q n n∈N VOLUME 00, 20265 Khoula et al.: Quantum-Aware Trajectory Optimization FIGURE 1: Proposed RA-QAGC throughput performance comparison with the baseline positioning methods. QAGC explicitly handles user tier prioritization, optimizing both total and priority throughput concurrently. Moreover, RA-QAGC is environment-aware spatial clustering schemes unlike the conventional K-means heuristics by mapping complex, physics-driven wireless channel behaviors directly into the clustering routine, bypassing channel-blind geomet- ric limitations. Also, the integration of a quantum-inspired probabilistic annealing trajectory actively prevents optimiza- tion routines from getting trapped in high-energy suboptimal local cost configurations. GC-SNRP baseline achieved the lowest total throughput performance around 37.8 Mbps as compared with other schemes. However, the proposed RA-QAGC framework proves that strict QoS requirements can be achieved without sacrificing aggregate system throughput. Figure2 show the optimal spatial deployment positions generated by the RA- QAGC. The deployment reveals that the ABS are centered directly over high-density user hotspots while simultaneously positioning themselves near priority nodes. The condensed graph setC acts as an information-dense representation of the terrestrial network topology, filtering out low-reward spatial positions to accelerate execution. The cumulative distribution function (CDF) of the indi- vidual per-user data rates is illustrated in Figure 3. The empirical distribution profiles demonstrate that the priority- user achieved high data rate as compared to the normal user. Specifically, priority users attain a median data rate of approximately 0.85 Mbps, whereas normal users achieve around 0.55 Mbps. It can be clearly noticed that for edge- user (5% CDF) the data rate is above zero, indicating that the framework successfully prevents cell-edge users while maintaining QoS. To validate the scalability and convergence of the proposed scheme under dynamic conditions, performance is bench- marked directly against state-of-the-art graph vision and communication (GVis&Comm) frameworks in [16], which FIGURE 2: Optimal multi-UAV spatial deployment layout generated by RA-QAGC. FIGURE 3: CDF of per-user data rates for the optimized RA-QAGC. couples multi-agent actor-critic loops with continuous over- the-air graph neural network, and with the Joint UAV Tra- jectory and Association Planning (JUTAP) framework [5], which leverages centralized Deep Q-Networks to manage macro-association actions. The aggregate network sum-rate of the proposed scheme with different schemes are shown in Fig. 4, which clearly shows that the proposed schemes outperformed all the schemes. The numerical result demonstrate that the RA-QAGC architecture maintains a consistent capacity advantage over the existing schemes across the entire time-step window. Al- though the scheme in [16] attempts to adjust trajectories via graph convolutions, its convergence behavior degrades heav- ily in deep-fading environments where spatial proximity does not map linearly to real-time SINR changes. Conversely, while the scheme in [5] minimizes handover disconnectivity via macro planning, it lacks a fine-grained, interference- aware trajectory tracking mechanism. This causes significant throughput degradation in high-density co-channel environ- ments. By contrast, because the quantum-inspired annealing in RA-QAGC minimizes a global cost function J driven 6VOLUME 00, 2026 FIGURE 4: Proposed RA-QAGC per user throughput com- pared with existing schemes. directly by real-time SINR constraints, that optimized search space, generating highly stable sum-rate. Figure5 (a) compares the aggregate network through- put achieved by different user-association and clustering schemes. The proposed RA-QAGC framework achieves the highest total throughput of 59.4 Mbps, outperforming all benchmark methods. Compared with Random association (33.9 Mbps), K-means (45.2 Mbps), GC-km (51.7 Mbps), GC-SNRP (37.8 Mbps), iGCVis (36.7 Mbps), and JUTAP (28.2 Mbps), RA-QAGC provides substantial throughput gains. These improvements is from its ability to jointly opti- mize UAV positioning and user association while accounting for real-time channel conditions and interference dynamics. By continuously adapting UAV trajectories to traffic de- mands and network conditions, RA-QAGC enhances spatial resource utilization, improves signal quality, and reduces interference, resulting in superior network-wide spectral ef- ficiency and throughput. Similarly, Figure5 (b) evaluates the per-user throughput achieved by users, highlighting the effectiveness of each scheme in supporting users with strin- gent QoS requirements. The proposed RA-QAGC attains the highest priority-user throughput of 23.9 Mbps, significantly exceeding GC-km (17.9 Mbps), K-means (15.4 Mbps), GC- SNRP (12.6 Mbps), iGCVis (11.6 Mbps), Random (9.7 Mbps), and JUTAP (8.4 Mbps). The considerable gain demonstrates that RA-QAGC not only maximizes overall network performance but also effectively prioritizes critical users. This is achieved through its adaptive reinforcement learning-based coordination mechanism, which dynamically adjusts UAV trajectories and service regions to maintain favorable communication links for high-priority users. Con- sequently, the proposed framework delivers enhanced QoS guarantees while simultaneously preserving high network throughput. Figure 6 shows the optimized trajectories learned by the IQL agent. The learned trajectories exhibit several intelli- gent behaviors: (1) UAVs maintain coordinated spacing to minimize interference, (2) trajectories follow predicted user density patterns, and (3) paths are smooth without abrupt direction changes, confirming that the IQL agent learned efficient mobility patterns. VI. Conclusion This work demonstrates that intelligent state-space reduc- tion can substantially improve the practicality of multi- UAV trajectory optimization in interference-limited wire- less networks. The proposed framework effectively balances network-wide throughput and user-level service require- ments, achieving 59.4 Mbps aggregate throughput and 23.9 Mbps priority-user throughput, outperforming all considered benchmark schemes. The results indicate that incorporating rate-awareness into the environment abstraction process en- ables UAVs to identify more favorable operating regions and utilize network resources more efficiently. Furthermore, the decentralized learning strategy provides a scalable alterna- tive to centralized optimization, whose complexity rapidly increases with network size. These findings suggest that state abstraction combined with distributed decision-making offers a promising direction for supporting dense UAV deployments in future wireless systems. Future research will focus on extending the framework to dynamic user mobility scenarios, continuous control models, and large-scale heterogeneous aerial networks. REFERENCES [1] Z. Kaleem, F. A. Orakzai, W. Ishaq, K. Latif, J. Zhao, and A. Ja- malipour, “Emerging trends in UAVs: From placement, semantic communications to generative AI for mission-critical networks,” IEEE Transactions on Consumer Electronics, vol. 71, no. 3, p. 7412–7438, 2025. [2] C. C. Ekechi, T. Elfouly, A. Alouani, and T. M. Khattab, “A survey on uav control with multi-agent reinforcement learning,” Drones, 2025. [Online]. Available: https://api.semanticscholar.org/CorpusID: 280142119 [3] A. Henshall and S. Karaman, “Generalized multiagent reinforcement learning for coverage path planning in unknown, dynamic, and haz- ardous environments,” AIAA SCITECH 2024 Forum, 2024. [Online]. Available: https://api.semanticscholar.org/CorpusID:267343711 [4] K. Gogineni, P. Wei, T. Lan, and G. Venkataramani, “Scalability bottlenecks in multi-agent reinforcement learning systems,” ArXiv, vol. abs/2302.05007, 2023. [Online]. Available: https://api.semanticscholar. org/CorpusID:256808299 [5] A. A. Shamsabadi, C. Mwaba, T. Nugent, J. Gao, P. Madoery, H. Yanikomeroglu, and S. Pal, “Dqn-based joint uav trajectory and association planning in ntn-assisted networks,” arXiv preprint arXiv:2603.22127, 2026. [6] S. Kim and J. Park, “Path planning with multiple uavs considering the sensing range and improved k-means clustering in wsns,” Aerospace, vol. 10, no. 11, p. 939, 2023. [7] T. V. Tung, T. T. An, and B. M. Lee, “Joint resource and trajectory op- timization for energy efficiency maximization in uav-based networks,” Mathematics, vol. 10, no. 20, p. 3840, 2022. [8] C. Kim, H.-H. Choi, and K. Lee, “Joint optimization of trajectory and resource allocation for multi-uav-enabled wireless-powered communication networks,” IEEE Transactions on Communications, vol. 72, p. 5752–5764, 2024. [Online]. Available: https://api. semanticscholar.org/CorpusID:268789122 [9] M. Misbah, Z. Kaleem, W. Khalid, C. Yuen, and A. Jamalipour, “Phase and 3-D placement optimization for rate enhancement in ris-assisted UAV networks,” IEEE Wireless Communications Letters, vol. 12, no. 7, p. 1135–1138, 2023. [10] T. Q. Dinh, S. H. Dau, E. Lagunas, S. Chatzinotas, D. N. Nguyen, and D. T. Hoang, “Quantum annealing for complex optimization in satellite communication systems,” IEEE Internet of VOLUME 00, 20267 Khoula et al.: Quantum-Aware Trajectory Optimization (a) Total throughput(b) Per-user throughput FIGURE 5: Throughput performance comparison of the proposed RA-QAGC framework. FIGURE 6: Optimized UAV trajectories over one complete mobility cycle. UAVs follow smooth, coordinated paths to maintain coverage of moving users. Things Journal, vol. 12, p. 3771–3784, 2025. [Online]. Available: https://api.semanticscholar.org/CorpusID:273414150 [11] S.-G. Jeong, P. D. A. Duc, Q. V. Do, D.-I. Noh, N. X. Tung, T. V. Chien, Q.-V. Pham, M. Hasegawa, H. Sekiya, and W. J. Hwang, “Quantum-annealing-based sum rate maximization for multi-uav-aided wireless networks,” IEEE Internet of Things Journal, vol. 12, p. 21 225–21 239, 2025. [Online]. Available: https://api.semanticscholar.org/CorpusID:276580836 [12] H. Lyu, J. Jang, H. Lee, and H. J. Yang, “Non-iterative optimization of trajectory and radio resource for aerial network,” IEEE Transactions on Wireless Communications, vol. 24, p. 1555–1567, 2024. [Online]. Available: https://api.semanticscholar.org/CorpusID:269502522 [13] Y. Li, A. H. Aghvami, and D. Dong, “Path planning for cellular- connected UAV: A DRL solution with quantum-inspired experience replay,” IEEE Transactions on Wireless Communications, vol. 21, no. 10, p. 7897–7912, 2022. [14] Silvirianti, B. Narottama, and S. Y. Shin, “Layerwise quantum deep reinforcement learning for joint optimization of UAV trajectory and resource allocation,” IEEE Internet of Things Journal, vol. 11, no. 1, p. 430–443, 2024. [15] Z. Kaleem, M. Afaq, C. Yuen, O. A. Dobre, and J. M. Cioffi, “Quantum-driven state-reduction for reliable UAV trajectory optimiza- tion in low-altitude networks,” IEEE Wireless Communications Letters, p. 1–1, 2026. [16] X. Zhang, H. Zhao, J. Wei, C. c. Yan, J. Xiong, and X. Liu, “Cooperative trajectory design of multiple uav base stations with heterogeneous graph neural networks,” IEEE Transactions on Wireless Communications, vol. 22, no. 3, p. 1495–1508, 2023. Khoula Khalid received Engineering Degree in Computer Programming and Specific Applications from the National School of Applied Sciences of Marrakech, Morocco, in 2021. She earned Master of Science (M.S.) degree in Computer Engineer- ing from King Fahd University of Petroleum and Minerals (KFUPM), Dhahran, Saudi Arabia, in 2026. Currently, she is pursuing her Ph.D. de- gree in Computer Science, specializing in Artifi- cial Intelligence, at the National School of Ap- plied Sciences of Kenitra (ENSA-K), Morocco. Her research interests include wireless communications, Unmanned Aerial Vehicles (UAVs), and advanced computational intelligence. Specifically, her work focuses on vehicle path planning, trajectory optimization, and intelligent frameworks combining reinforcement learning, machine learning, and quantum computing. 8VOLUME 00, 2026 MUHAMMAD AFAQ received a B.S. degree in Electrical Engineering from the University of Eng. and Technology, Pakistan in 2007. He received an MS degree in Electrical Engineering with an emphasis on Telecom from Blekinge Institute of Technology (Sweden) in 2010 and a Ph.D. degree in Computer Engineering from Jeju National Uni- versity (Korea) in 2017. Currently, he is working as an Assistant Professor at the Department of Computer Engineering, King Fahd University of Petroleum and Minerals, Saudi Arabia. His re- search interests are cloud computing, SDN, NFV, computer networks and protocols, and machine learning. Ali Arshad Nasir received the Ph.D. degree in telecommunications engineering from the Aus- tralian National University, Australia, in 2013, where he worked as a Research Fellow from 2012 to 2015. From 2015 to 2016, he was an Assistant Professor with the School of Electrical Engineer- ing and Computer Science, National University of Sciences and Technology, Pakistan. He joined the Department of Electrical Engineering, King Fahd University of Petroleum and Minerals, Dhahran, Saudi Arabia, in 2016, where he is currently work- ing as an Associate Professor. His research interests are in the area of signal processing in wireless communication systems. He served as an Editor for IEEE Wireless Communications Letters from 2021 to 2023 and for IEEE Communications Letters from 2024 to 2025. He received the Exemplary Editor Award from IEEE Communications Letters in 2025. He has been serving as an Editor for IEEE Transactions on Communications since 2026. Zeeshan Kaleem (Senior Member) is serving as an Associate Professor in the Computer En- gineering Department, King Fahd University of Petroleum and Minerals (KFUPM), Saudi Arabia. Prior to joining KFUPM he served for 8 Years at COMSATS University Islamabad. He received his BS in Electrical Engineering from Univer- sity of Engineering and Technology, Peshawar in 2007. He received MS and Ph.D. in Electron- ics Engineering from Hanyang University, and Inha University, South Korea in 2010 and 2016, respectively. Dr. Zeeshan consecutively received the National Research Productivity Award (RPA) awards from the Pakistan Council of Science and Technology (PSCT) in 2017 and 2018. We won the Runner-up Award in the National Hackathon 23 competition for Project to develop Drone Detection system. He won the Higher Education Commission (HEC) Best Innovator Award in 2017, with a single award from all over Pakistan. He received the 2021 Top Reviewer Recognition Award for IEEE Transactions on Vehicular Technology. He has published 100+ technical journal papers, including 21 as 1st author papers, books, book chapters, and conference papers in reputable journals/venues, and holds 21 US and Korean patents. He has also received research grants of around 90k US$. He is a co-recipient of the best research proposal award from SK Telecom, Korea. He is currently serving as Technical Editor of several prestigious Journals/Magazines like IEEE Transactions on Vehicular Technology, IEEE Transactions on Network and Service Management, Elsevier Computer and Electrical Engineering, Springer Nature Wireless Personal Communications, Human-centric Com- puting and Information Sciences, Journal of Information Processing Sys- tems, and Frontiers in Communications and Networks. He has served/serving as Guest Editor for special issues in IEEE Wireless Communications, IEEE Communications Magazine, IEEE Access, Sensors, IEEE/KICS Journal of Communications and Networks, and Physical Communications, and served as a Track Chair in VTC-Fall 2024 and VTC-Spring 2025. He also regularly serves as TPC Member for world-distinguished conferences like IEEE Globecom, IEEE VTC, IEEE ICC, and IEEE PIMRC. VOLUME 00, 20269