Paper deep dive
Separation Assurance between Heterogeneous Fleets of Small Unmanned Aerial Systems via Multi-Agent Reinforcement Learning
Iman Sharifi, Hyeong Tae Kim, Maheed Hatem Ahmed, Mahsa Ghasemi, Peng Wei
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 7/8/2026, 8:10:29 AM
Summary
This paper investigates multi-agent reinforcement learning (MARL) for tactical deconfliction and separation assurance among heterogeneous fleets of small unmanned aerial systems (sUASs) operating in dense urban airspace. The authors propose an attention-enhanced Proximal Policy Optimization-based Advantage Actor-Critic (PPOA2C) framework that enables independent fleet training while preserving policy privacy. Experimental results demonstrate that heterogeneous fleets can reach equilibrium for safe separation, with PPOA2C outperforming rule-based baselines and adapting safely to mixed-policy environments. However, the study highlights a fairness challenge: converged equilibria tend to favor fleets with stronger configurations, underscoring the need for fairness-aware conflict management.
Entities (8)
Relation Signals (7)
George Washington University â affiliatedwith â Iman Sharifi
confidence 96% ¡ I. Sharifi and P. Wei are with the Department of Mechanical and Aerospace Engineering, George Washington University, Washington, DC, USA.
Multi-Agent Reinforcement Learning (MARL) â usedfor â Tactical Deconfliction
confidence 95% ¡ To address these inherent safety challenges in multi-agent UASs, various methods have been proposed... multi-agent reinforcement learning (MARL) approaches have demonstrated superior reliability in maintaining separation assurance under highly dense traffic scenarios.
Heterogeneous fleets â reachequilibriumfor â Safe separation
confidence 94% ¡ Experimental results show that two fleets with distinct, shared PPOA2C policies can reach an equilibrium to maintain safe separation.
PPOA2C â isvariantof â Multi-Agent Reinforcement Learning (MARL)
confidence 93% ¡ We employ an advantage actor-critic (A2C) framework combined with a loss function derived from proximal policy optimization (PPO)... We refer to the combination of PPO and A2C as PPOA2C.
PPOA2C â incorporates â Attention Mechanism
confidence 91% ¡ To enhance the scalability and adaptability of the PPOA2C framework in high-density airspace environments, we incorporate a multiplicative attention mechanism into the A2C neural network architecture to capture high-level relationships among states.
PPOA2C â outperforms â Rule-based baselines
confidence 90% ¡ While two PPOA2C policies outperform two strong rule-based baselines in terms of conflict resolution, a PPOA2C policy exhibits safer interaction with a rule-based policy, indicating adaptive capabilities of PPOA2C policies.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:In the envisioned future dense urban airspace, multiple companies will operate heterogeneous fleets of small unmanned aerial systems (sUASs), where each fleet includes several homogeneous aircraft with identical policies and configurations, e.g., equipage, sensing, and communication ranges, making tactical deconfliction highly complex for the aircraft. This paper aims to address two core questions: (1) Can tactical deconfliction policies converge or reach an equilibrium to ensure a conflict-free airspace when companies operate heterogeneous fleets of homogeneous aircraft? (2) If so, will the converged policies discriminate against companies operating sUASs with weaker configurations? We investigate a multi-agent reinforcement learning paradigm in which homogeneous aircraft within heterogeneous fleets operate concurrently to perform package delivery missions over Dallas, Texas, USA. An attention-enhanced Proximal Policy Optimization-based Advantage Actor-Critic (PPOA2C) framework is employed to resolve intra- and inter-fleet conflicts, with each fleet independently training its own policy while preserving privacy. Experimental results show that two fleets with distinct, shared PPOA2C policies can reach an equilibrium to maintain safe separation. While two PPOA2C policies outperform two strong rule-based baselines in terms of conflict resolution, a PPOA2C policy exhibits safer interaction with a rule-based policy, indicating adaptive capabilities of PPOA2C policies. Furthermore, we conducted extensive policy-configuration evaluations, which reveal that equilibria between similar policy types tend to favor fleets with stronger configurations. Even under similar configurations but different policy types, the equilibrium favors one of the heterogeneous policies, underscoring the need for fairness-aware conflict management in heterogeneous sUAS operations.
Tags
Links
- Source: https://arxiv.org/abs/2605.01041v3
- Canonical: https://arxiv.org/abs/2605.01041v3
Trouble viewing inline? Open PDF directly â
Full Text
48,576 characters extracted from source content.
Expand or collapse full text
Separation Assurance between Heterogeneous Fleets of Small Unmanned Aerial Systems via Multi-Agent Reinforcement Learning Iman Sharifi1, Hyeong Tae Kim2, Maheed Hatem Ahmed2, Mahsa Ghasemi2, and Peng Wei1 1I. Sharifi and P. Wei are with the Department of Mechanical and Aerospace Engineering, George Washington University, Washington, DC, USA. i.sharifi,pwei@gwu.edu2H. T. Kim, M. H. Ahmed, and M. Ghasemi are with the Department of Electrical and Computer Engineering, Purdue University, West Lafayette, IN, USA. kim4741,ahmed237,mahsa@purdue.edu Abstract In the envisioned future dense urban airspace, multiple companies will operate heterogeneous fleets of small unmanned aerial systems (sUASs), where each fleet includes several homogeneous aircraft with identical policies and configurations, e.g., equipage, sensing, and communication ranges, making tactical deconfliction highly complex for the aircraft. This paper aims to address two core questions: (1) Can tactical deconfliction policies converge or reach an equilibrium to ensure a conflict-free airspace when companies operate heterogeneous fleets of homogeneous aircraft? (2) If so, will the converged policies discriminate against companies operating sUASs with weaker configurations? We investigate a multi-agent reinforcement learning paradigm in which homogeneous aircraft within heterogeneous fleets operate concurrently to perform package delivery missions over Dallas, Texas, USA. An attention-enhanced Proximal Policy Optimization-based Advantage Actor-Critic (PPOA2C) framework is employed to resolve intra- and inter-fleet conflicts, with each fleet independently training its own policy while preserving privacy. Experimental results show that two fleets with distinct, shared PPOA2C policies can reach an equilibrium to maintain safe separation. While two PPOA2C policies outperform two strong rule-based baselines in terms of conflict resolution, a PPOA2C policy exhibits safer interaction with a rule-based policy, indicating adaptive capabilities of PPOA2C policies. Furthermore, we conducted extensive policy-configuration evaluations, which reveal that equilibria between similar policy types tend to favor fleets with stronger configurations. Even under similar configurations but different policy types, the equilibrium favors one of the heterogeneous policies, underscoring the need for fairness-aware conflict management in heterogeneous sUAS operations. I Introduction Unmanned aerial systems (UASs) are increasingly being adopted across a wide range of commercial and industrial domains, including infrastructure inspection, aerial surveying, environmental monitoring, and package delivery [2]. Specifically, small unmanned aerial systems (sUASs) represent a rapidly expanding class of low-altitude UASs whose compact design, low cost, and operational flexibility make them particularly well-suited for high-frequency, localized missions [3]. Companies such as Google Wing, Amazon Prime Air, and Zipline are increasingly deploying fleets of sUASs to perform time-sensitive and spatially distributed delivery tasks. As these operations scale, UAS Traffic Management (UTM) service providers are expected to accommodate sUAS flight operations in low-altitude urban airspace, particularly in densely populated areas. In such shared airspace, each company often deploys vehicles with distinct aircraft performance, sensing and communication ranges, and proprietary control policies, making tactical deconfliction among sUASs highly challenging. Tactical deconflictionâalso known as conflict resolutionârequires rapid, reactive decision-making based on local and dynamic state information to ensure safe separation [22]. To address these inherent safety challenges in multi-agent UASs, various methods have been proposed [15, 9, 18], among which multi-agent reinforcement learning (MARL) approaches have demonstrated superior reliability in maintaining separation assurance under highly dense traffic scenarios [22]. In an MARL framework, each agent typically optimizes its policy in either a fully centralized [7, 4, 13, 6] or fully decentralized manner [1, 8]. However, in real-world urban aerial traffic, a hybrid combination of centralized and decentralized learning often provides a more practical solution to the challenges of a high-density shared airspace with multiple operator companies. While aircraft (agents) within a company typically employ identical deconfliction logic or policies, companies do not share their policies with other stakeholders due to privacy and safety concerns. To this end, a hybrid decentralized-centralized approach not only allows homogeneous agents within a company to share experiences for improved training efficiency but also enables heterogeneous fleets between different companies to preserve policy confidentiality, enhancing overall safety. In this paper, we investigate an MARL paradigm within a two-fleet, high-density scenario in which heterogeneous fleets of sUASs, each with multiple homogeneous aircraft, operate concurrently in a complex urban airspace to perform package delivery missions. Each fleet, representing a specific company such as Google Wing or Amazon Prime Air, consists of autonomous agents with distinct aircraft capabilitiesâsuch as maximum speed, acceleration, and sensor or communication range (e.g., one company uses radar detection for intruders, while the other employs Remote ID [11] for vehicle-to-vehicle information sharing). Each fleet independently trains its own MARL policy using an attention-based PPO-driven Advantage Actor-Critic (PPOA2C) framework [7, 8] to learn optimal tactical deconfliction strategies for both homogeneous and heterogeneous agents. While the architectural backbone of the method remains identical, the policies are trained separately, resulting in heterogeneous behaviors that reflect both differing vehicle dynamics and independent optimization processes. This setup mirrors realistic constraints found in competitive and proprietary operational environments. In particular, we seek to answer the following essential questions: (1) Will two fleets with proprietary PPOA2C policies lead to converged tactical deconfliction policies or reach an equilibrium for a conflict-free airspace, given that companies operate heterogeneous aircraft types with distinct sensing and communication ranges? (2) If so, will the equilibrium discriminate against a company operating sUASs with weaker performance, inferior equipage, or shorter sensing and communication ranges? We aim to investigate both operational safety and fairness challenges across heterogeneous sUAS fleets. The main contributions of this paper are as follows: ⢠We employ PPOA2C in a dense air traffic scenario including two heterogeneous fleets of sUASs, striving to ensure safe separation given the heterogeneity in both distributed policies and agentsâ physical configurations. We demonstrate that heterogeneous MARL policies for mixed aircraft types can reach an equilibrium. ⢠We show that the PPOA2C method outperforms a strong rule-based method in a dense scenario over Dallas, Texas. Moreover, PPOA2C not only cooperates safely with another PPOA2C policy but also learns to interact safely with a rule-based method. ⢠Through extensive policy-configuration evaluations, we show that the equilibrium with converged policies tends to favor aircraft with stronger configurations. Furthermore, even with identical configurations, training between two different policy types can exhibit discrimination against one of them. The remainder of this paper is organized as follows. Section I reviews the related literature. Section I presents the problem formulation and technical approach. Section IV details the experimental setup and analyzes the results. Finally, Section V concludes the paper. I Related Work Two types of tactical deconfliction strategies based on MARL have been proposed to improve separation assurance in air traffic control [22]: centralized training with decentralized execution (CTDE) and decentralized training and execution (DTE). The CTDE method leverages either global or local information from agents to train a centralized unit, such as a global critic or a shared actor-critic neural network, which governs agentsâ behaviors during deconfliction in high-density en-route airspace. Brittain and Wei [7] introduced a multi-agent actorâcritic framework that combines the advantage actor-critic (A2C) algorithm with elements of proximal policy optimization (PPO) under a centralized training, decentralized execution paradigm. To accommodate a variable number of agents, Brittain and Wei [4] proposed a long short-term memory (LSTM) network, enabling each agent to process a flexible number of intruder inputs. More recent approaches address the challenge of handling varying numbers of intruders by employing graph convolutional neural networks [16] and attention mechanisms [5, 6]. Groot et al. [12] compared three different attention mechanismsâscaled dot-product, additive, and context-aware attentionâintegrated with the soft actor-critic (SAC) algorithm to manage separations in high-density traffic scenarios. Other centralized reinforcement learning approaches integrate classical modelsâsuch as the 3D reciprocal velocity obstacle principle [25] and the solution space diagram method [24]âto enhance the quality of the observations provided to the reinforcement learning agent. In contrast, the DTE framework assumes that each agentâs training and policy are independent of those of other agents. While the CTDE framework may suffice for homogeneous fleets, heterogeneous UAS operations demand more flexible architectures, which can be supported by DTE. In a fully decentralized setting, Brittain and Wei [8] presented a distributed deep reinforcement learning approach using an off-policy SAC algorithm enhanced with attention networks. Our method combines the advantages of centralized learning with the scalability of decentralization: within a single sUAS fleet, agents share parameters and data to improve sample efficiency, while training remains decentralized across different fleets to ensure both policy flexibility and privacy. I Problem Formulation and Methodology I-A Tactical Deconfliction Strategy In future airspace operations, a combination of CTDE and DTE strategies is more realistic than fully centralized or fully decentralized approaches, as it accommodates different policies for each fleet while allowing policy sharing among agents within the same fleet. The primary objective of such mixed frameworks is to resolve both intra- and inter-fleet conflicts through an intelligent autonomous framework in which each fleetâs control policy is trained independently and remains private from other entities. This hybrid strategy distributes a shared policy among homogeneous agents while distinguishing between heterogeneous fleetsâ policies. It aims to ensure safe separation among both homogeneous and heterogeneous agents in dense aerial traffic environments where multiple companies simultaneously utilize shared airspace. Although a common MARL framework is adopted across fleets, each fleetâs agents have distinct configurations and train independently using their own collected experiences. Consequently, the learning processes and policy parameters remain fully independent, and information sharing occurs only through publicly available data. Through MARL, agents can adapt to dynamic environments and improve over time by interacting not only with homogeneous teammates but also with heterogeneous agents from other fleets. I-B Multi-Agent PPO-driven Advantage Actor-Critic To jointly train homogeneous agents within heterogeneous fleets, we employ an advantage actor-critic (A2C) framework [17] combined with a loss function derived from proximal policy optimization (PPO) [20, 23], similar to [5]. A2C is a policy-gradient method that often uses a unified neural network architecture to estimate both the policy (actor) and the value (critic) functions. PPO complements A2C by introducing a clipping-based loss function that stabilizes learning, ensuring that policy updates remain within a controlled range of the previous policy. We refer to the combination of PPO and A2C as PPOA2C, which enables efficient exploration of the action space and allows agents to refine their strategies by focusing on rewarding behaviors, ultimately improving robustness and convergence performance within the decentralized multi-agent setting. To enhance the scalability and adaptability of the PPOA2C framework in high-density airspace environments, we incorporate a multiplicative attention mechanism [6, 12] into the A2C neural network architecture to capture high-level relationships among states. This extension enables each aircraft agent to selectively attend to the most relevant intrudersâan essential capability in scenarios where the number of nearby intruders varies dynamically. The attention module encodes a variable number of neighboring agents into a fixed-length context vector, which is subsequently passed to the policy and value networks. After computing action probabilities and the state value using the A2C network, the optimization process minimizes two primary loss components, including the policy loss: âĎ= _Ď= tâ[minâĄ(Îśtâ(θ)â At,clipâ(Îśtâ(θ),1âĎľ,1+Ďľ)â At)] \ E_t [ ( _t(θ)¡ A_t,\;clip( _t(θ),1-Îľ,1+Îľ)¡ A_t ) ] âβâHâ(Ďâ(st)), -β\ H(Ď(s_t)), and the value loss: âv=At2L_v=A_t^2. The advantage function AtA_t quantifies the relative quality of an action compared to the expected behavior under the current policy. To compute AtA_t more accurately and robustly, we employ the Generalized Advantage Estimation (GAE) technique [19]. The term Îśtâ(θ) _t(θ) is the likelihood ratio between the new and old policies. The parameter Ͼξ defines the clipping range that restricts the magnitude of policy updates. The entropy regularization term βâ Hâ(Ďâ(st))β¡ H(Ď(s_t)) encourages exploration by penalizing premature convergence to deterministic policies, where Hâ(Ďâ(st))H(Ď(s_t)) denotes the entropy of the current policy distribution and β controls its influence during training. Figure 1: Use-case scenario based on Frisco, a suburb area in Dallas, Texas, simulated in BlueSky. Merging points WP3 and WP7 and the Intersection WP9 correspond to M1, M2, and IN bottlenecks, respectively. I-C MARL Components The core MARL components are outlined as follows: I-C1 State Space In reinforcement learning, the state space represents the set of all possible environmental configurations that an agent can observe at a given time. Here, we assume that aircraft state and dynamic information are openly shared among agents. This assumption reflects realistic operational considerations, as the Federal Aviation Administration (FAA) mandates Remote ID [11] to broadcast essential information for all participating sUASs. According to FAA standards, all agents broadcast basic information, including aircraft ID, location, altitude, and speed, to neighboring aircraft and the nearest ground control station. Thus, we assume each agent has access to the basic information of nearby agents. The states for each agent are divided into two components: the ownship state and the intruders state, defined as: sto=[d(o),v(o),θ(o),vâ˛âŁ(o)],sti=[do(i),v(i),θ(i),vâ˛âŁ(i)],s_t^o=[d^(o),\ v^(o),\ θ^(o),\ v (o)],\;\;s_t^i=[d_o^(i),\ v^(i),\ θ^(i),\ v (i)], where stoââ|so|s_t^o ^|s^o| denotes the ownship state, consisting of the distance to the next waypoint d(o)d^(o), current speed v(o)v^(o), heading θ(o)θ^(o), and acceleration vâ˛âŁ(o)v (o). Except do(i)d_o^(i) which is the distance between ownship o and intruder i, the intruder state stiââ|si|s_t^i ^|s^i| mirrors the ownship state structure. Intruder is an aircraft within a certain range around the ownship. To avoid complexity in the decision-making, we only consider front intruders, as aircraft arriving at the next bottleneck earlier than the ownship, since they have the most direct influence on the ownshipâs immediate tactical decisions. Following [10], this approach improves both computational and operational efficiency. After extracting the states, each component of stos_t^o and stis_t^i undergoes a normalization step before being fed into the PPOA2C neural network. I-C2 Action Space In this study, the action space is defined as the set of discrete speed adjustments that an aircraft can apply at each decision step. The agent can choose to decelerate, hold its current speed, or accelerate: =âÎâv, 0,+ÎâvA=\- v,\;0,\;+ v\, where Îâv v is the speed adjustment magnitude specific to each fleet of agents and varies depending on their configurations. After selecting an action atoâa_t^o , it is applied to the ownshipâs current speed and transmitted to the simulation environment. I-C3 Reward Function In reinforcement learning, the reward function provides a scalar feedback signal that reflects the desirability of the stateâaction pair executed by the agent. In our framework, the total reward rtor_t^o for each ownship agent at each time step is composed of five distinct components: rto=RLoS+RV+RA+RM+RT,r_t^o=R_LoS+R_V+R_A+R_M+R_T, where RLoSR_LoS, known as the loss of separation (LoS) reward, penalizes unsafe separation; RVR_V penalizes undesirable velocities; RAR_A penalizes undesirable flight behaviors such as abrupt speed changes; RMR_M incentivizes task completion; and RTR_T encourages time efficiency. Since maintaining safe separation is the central objective of this study, RLoSR_LoS constitutes the primary term in the reward function. It is defined as: RLoS=â1,if âdo(i)<dNMAC,Îąâ(â1+do(i)âdNMACdLoWCâdNMAC),if âdNMACâ¤do(i)â¤dLoWC,0,otherwise,R_LoS= cases-1,&if d_o^(i)<d_NMAC,\\ Îą (-1+ d_o^(i)-d_NMACd_LoWC-d_NMAC ),&if d_NMAC⤠d_o^(i)⤠d_LoWC,\\ 0,&otherwise, cases where do(i)d_o^(i) denotes the distance between the ownship and the i-th intruder, dNMACd_NMAC is the near mid-air collision (NMAC) threshold, and dLoWCd_LoWC is the loss of well clear (LoWC) threshold. If the separation distance falls below dNMACd_NMAC, a severe penalty of â1-1 is imposed, immediately removing the ownship and the corresponding intruder agent from the simulation. If the distance lies between dNMACd_NMAC and dLoWCd_LoWC, the penalty is linearly scaled with distance using the coefficient Îąâ[0,1]Îąâ[0,1]. To ensure that the agentsâ speeds remain within a specified range, another reward component penalizes violations of the speed limits, as defined below: RV=âĎ1vâ â[v(o)<vmin(o)+Ρ1v]âĎ2aâ â[v(o)>vmax(o)âΡ2v],R_V=- _1^v¡I\!\ [v^(o)<v_min^(o)+ _1^v]- _2^a¡I\!\ [v^(o)>v_max^(o)- _2^v], where â[â ]I[¡] equals 11 if the condition holds, and 0 otherwise. vmin(o)v_min^(o) and vmax(o)v_max^(o) are the nominal minimum and maximum speeds of the ownship. The parameters Ρ1v _1^v and Ρ2v _2^v are offset values that prevent agents from reaching the exact minimum and maximum velocities, respectively. The hyperparameters Ď1v _1^v and Ď2v _2^v assign small penalties when these limits are approached. These constraints prevent agents from becoming stuck in the environment and encourage energy efficiency. In real-world operations, frequent or abrupt speed adjustments in urban environments can destabilize aircraftâparticularly those carrying payloadsâand increase risks for nearby traffic, thereby compromising safety. To discourage such undesirable behaviors, we introduce an action penalty term RAR_A, defined as: RA=âĎ1aâ â[at(o)â atâ1(o)]âĎ2aâ â[at(o)â HOLD],R_A=- _1^a¡I\!\ [a_t^(o)â a_t-1^(o)]- _2^a¡I\!\ [a_t^(o)â HOLD], where at(o)a_t^(o) and atâ1(o)a_t-1^(o) denote the current and previous actions of the ownship agent, respectively. The hyperparameters Ď1a _1^a and Ď2a _2^a control the penalties for frequent speed changes and deviations from steady flight. This formulation encourages agents to maintain smoother and more consistent speed profiles, thereby promoting both energy conservation and operational safety. We also incorporate a reward component to promote mission accomplishment, represented by: RM=Ďmâ â[df(o)<Ρm]R_M= _m¡I\!\ [d_f^(o)< _m], which provides a bonus for reaching the final destination. df(o)d_f^(o) denotes the distance between the ownship and its final waypoint, and Ρm _m is a predefined distance threshold. If the aircraft reaches within a specified proximity to its destination, it receives a terminal reward of Ďm _m and is withdrawn from the simulation; otherwise, it receives no penalty. To encourage timely mission completion, a small negative penalty âĎt- _t is applied at every time step. This component incentivizes efficient traversal toward the goal, as: RT=âĎtâ â[t<T]ââ[tâĽT]R_T=- _t¡I\!\ [t<T]-I\!\ [t⼠T], where T is a time threshold after which the agent can no longer operate in the environment and is terminated from the simulation. Then, the agent receives a penalty of â1-1. Overall, the reward function prioritizes safety, efficiency, and smoothness in flight operations. I-D Training Approach After collecting all state information for each agent within a fleet, the corresponding PPOA2C network architecture first processes stos_t^o and stis_t^i through separate 6464-unit fully connected layers, followed by a 64Ă6464Ă 64 attention module to capture interaction features. The resulting representation is further transformed by two 128128-unit fully connected layers before branching into the actor and critic heads to generate the action probability distribution and the state-value estimate. Based on this probability distribution, an action atoa_t^o is sampled for each ownship agent; a reward rtor_t^o is then assigned accordingly, and the tuple (sto,sti,ato,rto)(s_t^o,s_t^i,a_t^o,r_t^o) is stored in the fleet replay buffer. This process continues until the episode terminates. After each episode, discounted returns and advantage estimates are computed from the collected trajectories and appended to the corresponding samples. Training then proceeds by evaluating the PPOA2C loss over a batch of 512512 samples. The network parameters are updated using the Adam optimizer, and this optimization procedure is repeated for 88 epochs per episode for each fleet. Once the policy and value networks have been optimized, the resulting policy is shared among all homogeneous agents within that fleet. This process is executed independently for each fleet, after which the learned policies are deployed to their respective agents, thereby preserving policy coherence within fleets while maintaining heterogeneity across fleets. The learning rate, entropy coefficient β, and clipping threshold Ͼξ are 10â410^-4, 10â310^-3, and 0.20.2, respectively. The separation thresholds dNMACd_NMAC and dLoWCd_LoWC are set to 100100 m and 500500 m, respectively, and the goal tolerance Ρm _m is 5050 m. The simulation time step Îât t and mission horizon T are 33 s and 1818 min, respectively. The speed offsets Ρ1v _1^v and Ρ2v _2^v are 5.145.14 m/s and 2.572.57 m/s, respectively. The reward-shaping coefficients Ď1v _1^v and Ď2v _2^v are 10â310^-3 and 10â410^-4, respectively, while Ď1a _1^a and Ď2a _2^a are 10â510^-5 and 10â410^-4, respectively. The remaining coefficients are Îą=0.1Îą=0.1, Ďt=10â4 _t=10^-4, and Ďm=0.1 _m=0.1. IV Experimental Results IV-A Simulation Setup To model and evaluate the performance of our tactical deconfliction framework, we employ BlueSky [14], an open-source, fast-time air traffic simulator widely recognized in the aviation research community for its versatility and scalability. IV-A1 Use-case Scenario To train MARL policies in a realistic and structured environment, we design a custom scenario based on the airspace over Frisco, a suburban area in Dallas, Texas, as illustrated in Fig. 1. The scenario consists of four fixed routes representing typical drone delivery paths: ⢠Route I: WP1 â WP3 â WP9 â WP4 ⢠Route I: WP2 â WP3 â WP9 â WP4 ⢠Route I: WP5 â WP7 â WP9 â WP8 ⢠Route IV: WP6 â WP7 â WP9 â WP8 The origin points of these routes (marked by blue and orange stars) are selected approximately based on the locations of major commercial hubs such as Walmart and Walgreens. The destination points (green stars) correspond to residential areas where delivery demand is expected to be high during the day. The routes are designed to have approximately the same total length of 10.3310.33 km for both companies. To simulate a realistic and challenging operational environment, the routes are deliberately configured to create bottlenecksâspecifically at WP3 and WP7 as merging waypoints and WP9 as an intersection waypointâwhich reflect common congestion points in low-altitude airspace networks. In this scenario, we simulate a total of 2020 agents, with 1010 agents assigned to each company. Company A (Co. A) agents operate on Route I and Route I (orange stars in Fig. 1), while Company B (Co. B) agents operate on Route I and Route IV (blue stars in Fig. 1). This configuration results in five agents per route. Agent spawn times (in seconds) are defined as 35+5âk35+5k, where k is randomly selected from 0,1,âŚ,10\0,1,âŚ,10\ to introduce temporal variability. This design not only generates numerous conflict situations at bottlenecks but also provides sufficient temporal spacing for agents to make informed and strategic decisions. Conflicts are expected to arise at merging points WP3 and WP7, where agents from both companies converge. However, the most complex and congested interactions occur at the intersection point WP9. This setup reflects realistic operational conditions in shared urban airspace, where agents must resolve both intra- and inter-company conflicts. IV-A2 sUAS Configurations To model heterogeneity, we consider distinct configurations for each fleet of agents, including different speed limits, acceleration ranges, and sensory capabilities, reflecting the diversity of hardware and onboard systems across drone operators. In general, we consider two configurations: X and Y, where configuration X possesses stronger capabilities than configuration Y. The speed limits and acceleration ranges for configurations X and Y are selected based on the performance specifications of the Google Wing Hummingbird drone and the Amazon MK30 drone, respectively. The sensory ranges are chosen to align with current technological standards for drone communication via Remote ID [11] or radar detection, ensuring that agents can detect and respond to nearby intruders while maintaining distinct sensing ranges. Configurations X (strong) and Y (weak) have speed ranges [0,44.88][0,44.88] and [0,30.12][0,30.12] m/s, acceleration sets â1.71,0,1.71\-1.71,0,1.71\ and â1.02,0,1.02\-1.02,0,1.02\ m/s2, and sensing ranges of 10001000 and 750750 m, respectively. The ultimate goal is to train all agents with both homogeneous and heterogeneous configurations such that they learn to minimize the occurrence of NMACs during operation while successfully completing their mission. The desired outcome is a cooperative and adaptive system in which all agents fly in the shared airspace safely and efficiently despite differences in control policies, aircraft performance, and sensor capabilities. In particular, agents are expected to coordinate not only with teammates trained under the same policy but also with agents operating under independently trained policies from competing companies. Figure 2: Average total rewards gained by both Co. A and Co. B agents with various policy-configuration models. (a) Average Reward (b) Successes and NMACs Figure 3: Average total reward and number of successful agents with average NMACs per episode for both fleets. IV-B Experimental and Numerical Analyses The PPOA2C framework is applied to the aforementioned complex use-case scenario to resolve intra- and inter-fleet conflicts. IV-B1 Baselines To evaluate the performance of the PPOA2C framework, we adopt a Rule-based policy, same as [21], that operates similarly to conventional air traffic decision-making systems. Unlike the rule-based method used by Chen et al. [10], which underrepresents the true performance of such methods, we enhance the rule-based approach with three additional features to make it more human-like and well-rounded. These enhancements include: (1) considering aircraft in merging/intersecting routes, not just the same route, when making decisions; (2) incorporating the distance of both the ownship and the closest intruder from the next waypoint; and (3) accounting for the nearest following intruder when determining maneuvers. This agile and informed method serves as a strong standard baseline for comparison. Additionally, we include a weak Random baseline to demonstrate the performance of a poorly performing policy. IV-B2 Policy-Configuration Combinations For a comprehensive evaluation, we consider seven different combinations of policies and sUAS configurations, each expressed in the following format: policy A (configuration A) ++ policy B (configuration B), where A and B represent two distinct fleets of agents. Policy A and policy B can be either Random, Rule-based, or PPOA2C, while configuration A and configuration B can be either X or Y. The first group of policy-configuration models considers different configurations for each fleet (X for fleet A and Y for fleet B), including: Random(X) ++ Random (Y), Rule-based(X) ++ Rule-based(Y), PPOA2C(X) ++ PPOA2C(Y), and PPOA2C(X) ++ Rule-based(Y). In contrast, the second group considers identical configurations (X) for both fleets, including: Rule-based(X) ++ Rule-based(X), PPOA2C(X) ++ PPOA2C(X), and PPOA2C(X) ++ Rule-based(X). These diverse combinations allow us to examine the influence of policy and configuration types on the overall performance of each model. During training, the PPOA2C models are updated, while the Random and Rule-based policies remain fixed. IV-B3 Training Hardware All experiments were conducted using PyTorch on an NVIDIA RTX 30903090 graphics card. Each training process consisted of 2,5002,500 episodes (approximately 600,000600,000 time steps), with each episode terminating once all agents were either truncated or removed from the simulation environment. Each episode simulated approximately 1212 minutes of flight operations, corresponding to 240240 environment time steps. Policy weights were updated at the end of each episode. On average, each complete training process required more than seven hours to finish. To ensure the reliability of results, we trained with five different random seeds (approximately 140140 hours of total training) and report the averaged outcomes across all seeds. After training, 300300 evaluation episodes were conducted to assess the effectiveness of the learned policies. TABLE I: Evaluation results for all policy-configuration combinations. A and B represent Co. A and Co. B, and X and Y are the strong and weak configurations, respectively. XY Configurations X Configurations Policy A (Config. A): Policy B (Config. B): Random(X) Random(Y) Rule-based(X) Rule-based(Y) PPOA2C(X) PPOA2C(Y) PPOA2C(X) Rule-based(Y) Rule-based(X) Rule-based(X) PPOA2C(X) PPOA2C(X) PPOA2C(X) Rule-based(X) Average NMAC (â ) A 3.30 0.24 0.03 0.000 0.19 0.02 0.01 AB 1.23 0.25 0.15 0.005 0.30 0.26 0.13 B 3.44 0.26 0.02 0.001 0.04 0.02 0.00 M1 3.92 0.00 0.00 0.005 0.16 0.13 0.06 M2 3.96 0.00 0.00 0.000 0.13 0.17 0.08 IN 0.10 0.77 0.22 0.001 0.24 0.00 0.00 Total 7.99 0.77 0.22 0.006 0.53 0.30 0.14 Success (â ) NsN_s / 20 3.96Âą 1.77 18.53Âą 1.65 19.55Âą 0.98 19.98Âą 0.10 19.42Âą 0.7 19.79Âą 0.60 19.90Âą 0.34 Reward (â ) RÂŻ(Ă103) R\,(Ă 10^3) â- -3.15 -1.41 -1.76 -3.27 -0.80 -1.87 Mission Time TÂŻA T_A (min) â- 5.88 7.23 6.58 5.05 7.42 5.46 TÂŻB T_B (min) â- 7.13 8.94 7.60 5.05 7.54 5.02 Fairness (â ) Ft(%)F_t\,(\%) â- 82.4 80.8 86.5 100 98.4 91.9 IV-B4 Training Analysis Due to the large number of policy-configuration models, we only present the average total reward for models with learnable policies. We then focus on visualizing and analyzing the results of the PPOA2C(X)++PPOA2C(Y) model. During the evaluations and comparisons, the remaining models are also considered. Reward Analysis: Fig. 2 depicts the average total reward obtained by both fleets for the trainable models. All average rewards consistently increased and converged to small negative valuesârepresenting approximate maxima. Notably, the PPOA2C(X)++PPOA2C(Y) and PPOA2C(X)++PPOA2C(X) models converged to higher final values compared to the PPOA2C(X)++Rule-based(Y) and PPOA2C(X)++Rule-based(X) models, indicating that two heterogeneous PPOA2C policies outperformed the configurations combining one PPOA2C policy with a Rule-based policy. To examine the reward behavior in more detail, we focus on the most relevant model: PPOA2C(X)++PPOA2C(Y). Fig. 3 illustrates the average instantaneous rewards obtained by Co. A and Co. B agents for this model. Both average rewards converged to small negative valuesârepresenting approximate maximaâafter about 1,2001,200 training episodes. According to the defined reward function, the maximum value is expected to be near zero, confirming that the rewards converged close to the optimal point. This outcome suggests that the learned policies not only reduce the number of NMACs but also enable smooth and efficient flight behavior consistent with the reward design. Moreover, as shown in Fig. 3, the number of successful agents (NsN_s), which finished their missions without NMACs, steadily increases during training and eventually converges to the maximum possible value of 2020, consistent with the observed reward trends. NMAC Analysis: As shown in Fig. 3, the number of NMACs (NcN_c) per episode decreases toward zero for the PPOA2C(X)++PPOA2C(Y) model, although occasional NMACs still occur. We hypothesize that these rare collisions arise under highly dense traffic conditions, where agentsâ actions become interdependent, and one agentâs maneuver may inadvertently cause confusion or delayed reactions in others. IV-C Evaluation Analysis and Comparison During the evaluation process, the average numbers of NMACs (NcN_c) and successful missions (NsN_s) in each category were recorded for all policy-configuration models, as shown in Table I. The PPOA2C(X)++PPOA2C(Y) and PPOA2C(X)++PPOA2C(X) models achieved mission success rates of 97.75%97.75\% and 98.95%98.95\%, respectively (an overall average of 98.35%98.35\%). For the XY (X) configuration settings, the PPOA2C(X)++PPOA2C(Y) (PPOA2C(X)++PPOA2C(X)) model increased NsN_s by 5.5%5.5\% (1.9%1.9\%) compared to the Rule-based(X)++Rule-based(Y) (Rule-based(X) ++ Rule-based(X)) model, while reducing the average number of NMACs by 71.4%71.4\% (43.3%43.3\%), respectively. To further analyze NMAC occurrences, we examine them from two perspectives: (1) company-to-company (C2C) NMACs and (2) bottleneck NMACs. C2C NMACs are categorized into three types: A, B, and AB, representing collisions between two Co. A agents, two Co. B agents, and one Co. A with one Co. B agent, respectively. Similarly, bottleneck NMACs are divided into three categories: M1, M2, and IN, corresponding to the merging points WP3 and WP7, and the intersection point WP9, respectively. Table I summarizes the C2C and bottleneck NMACs for all policy-configuration models. For the PPOA2C(X)++PPOA2C(Y) model, approximately 75%75\% of the remaining NMACs fall under the AB category, while A and B NMACs are rare. NMACs still occur across all bottlenecks, but most unresolved conflicts are concentrated at M1 and M2. This occurs because many aircraft fail to reach the intersection, as NMACs typically arise earlier at the merging points. Over time, IN NMACs disappear, whereas occasional M1 and M2 NMACs persist. Across nearly all NMAC categories, the PPOA2C(X)++ PPOA2C(Y) (PPOA2C(X)++PPOA2C(X)) model outperformed the Rule-based(X)++Rule-based(Y) (Rule-based(X) ++ Rule-based(X)) model. This result indicates that two heterogeneous PPOA2C policies cooperate more effectively and safely than two rule-based policies in this scenario, regardless of whether the agents use different or identical configurations. Notably, the PPOA2C policy and the Rule-based method interact even more safely than when two PPOA2C policies interact with each other. Based on Table I, the PPOA2C(X) ++ Rule-based(Y) (PPOA2C(X) ++ Rule-based(X)) model increased NsN_s by 2.19%2.19\% (0.55%0.55\%) compared to the PPOA2C(X) ++ PPOA2C(Y) (PPOA2C(X) ++ PPOA2C(X)) model. However, the average evaluation reward RÂŻ R for the PPOA2C(X)++PPOA2C(Y) (PPOA2C(X)++PPOA2C(X)) model is higher than that of the PPOA2C(X)++Rule-based(Y) (PPOA2C(X)++Rule-based(X)) model. This indicates that while a PPOA2C policy and a Rule-based policy interact more safely, the interaction is not necessarily efficient. The reason lies in the Rule-based methodâs tendency to adjust speed inefficiently to avoid conflicts. Implications: Based on the policy-configuration models, the results suggest that two PPOA2C policies can still learn to reach an equilibrium while maintaining near-optimal safety and efficiency. This interaction is significantly safer and more efficient than that between two well-engineered Rule-based methods. Moreover, the PPOA2C policy can also interact safely with the Rule-based method; however, such interactions are less efficient than those between two PPOA2C policies. IV-D Fairness Analysis When two independently trained policies interact, the learned behaviors may favor one policy over the other, potentially introducing unfair interactions. The goal here is to evaluate how the MARL framework manages this issue and whether it ensures fairness. We assess fairness based on mission time by initializing agents from both fleets under similar conditions, such as identical initial velocities. Since the travel distances are nearly equal, both fleets are expected to complete their missions in approximately the same amount of time. Given the average mission times for Co. A and Co. B as TÂŻA T_A and TÂŻB T_B, respectively, we propose a standard time-based fairness metric, defined as Ft(%)=F(TÂŻA,TÂŻB)F_t(\%)=F( T_A, T_B), where F:â+2â[0,100]F:R_+^2â[0,100] is expressed as: Fâ(t1,t2)=(1â|t1ât2|maxâ(t1,t2))Ă100.F(t_1,t_2)= (1- |t_1-t_2|max(t_1,t_2) )Ă 100. Higher FtF_t values indicate greater time-based fairness between fleets. Table I presents TÂŻA T_A, TÂŻB T_B, and FtF_t for all evaluation episodes. When the policy and configuration types are identical (e.g., PPOA2C(X)++PPOA2C(X)), TÂŻA T_A and TÂŻB T_B are nearly equal, resulting in a high FtF_t of approximately 99.2%99.2\% on average for the corresponding policy-configuration models. However, when either the policy or configuration types differ, TÂŻA T_A and TÂŻB T_B diverge significantly, yielding lower FtF_t values ranging from 80.8%80.8\% to 91.9%91.9\%. When policy types are identical but configurations differ, such as in the PPOA2C(X)++PPOA2C(Y) model, TÂŻB T_B is 22.4%22.4\% greater than TÂŻA T_A, indicating that Co. A, with the stronger configuration, completes the mission faster than Co. B. This results in an average FtF_t of 81.6%81.6\% for the corresponding policy-configuration models, reflecting significant bias against Co. B with weaker configurations. Similarly, when configurations are identical but policy types differ (e.g., PPOA2C(X)++Rule-based(X)), TÂŻA T_A is 8.76%8.76\% greater than TÂŻB T_B, with an FtF_t of 91.9%91.9\%. This suggests that, given equal configurations, Co. B with a Rule-based policy operates faster than Co. A with a PPOA2C policy. Hence, policy differences can also contribute to mission time disparities in heterogeneous settings. Implications: Based on the policy-configuration models, both policy heterogeneity and configuration heterogeneity can lead to lower time-based fairness. Furthermore, the fairness evaluation reveals that interactions consistently favor agents with stronger configurations. Even under identical configurations, differing policies may introduce discriminatory interactions among agents in terms of mission completion time. V Conclusion This work investigated safe separation in dense urban airspace involving heterogeneous fleets of small unmanned aerial systems (sUASs). We employed a multi-agent reinforcement learning framework built on an attention-enhanced PPO-driven Advantage Actor-Critic (PPOA2C) algorithm and evaluated it in a realistic scenario based on the airspace over Dallas, Texas, USA. This study aimed to answer two core questions: (1) Can tactical deconfliction policies with distinct configurations reach an equilibrium to maintain a conflict-free airspace? (2) Do these policies behave fairly across fleets? Experimental results show that heterogeneous PPOA2C policies for different companies, with either similar or distinct configurations, are capable of reaching an equilibrium while outperforming strong rule-based policies in terms of mission success rate. Moreover, a PPOA2C policy demonstrates safer interaction with a Rule-based policy compared to another PPOA2C policy, although two PPOA2C policies exhibit greater efficiency based on the achieved rewards. Through extensive policy-configuration evaluations, we observed that equilibria between similar policy types exhibit bias favoring fleets with stronger configurations. Even under similar configurations but differing policy types, the equilibrium still tends to favor one of the heterogeneous policies. References [1] L. E. Alvarez, I. Jessen, M. P. Owen, J. Silbermann, and P. Wood (2019) ACAS sXu: Robust decentralized detect and avoid for small unmanned aircraft systems. In 2019 IEEE/AIAA 38th Digital Avionics Systems Conference (DASC), p. 1â9. Cited by: §I. [2] G. Ariante and G. Del Core (2025) Unmanned aircraft systems (UASs): current state, emerging technologies, and future trends. Drones 9 (1), p. 59. Cited by: §I. [3] R. W. Beard and T. W. McLain (2012) Small unmanned aircraft: theory and practice. Princeton University Press. Cited by: §I. [4] M. W. Brittain and P. Wei (2021) One to any: distributed conflict resolution with deep multi-agent reinforcement learning and long short-term memory. In AIAA Scitech 2021 Forum, p. 1952. Cited by: §I, §I. [5] M. W. Brittain, X. Yang, and P. Wei (2021) Autonomous separation assurance with deep multi-agent reinforcement learning. Journal of Aerospace Information Systems 18 (12), p. 890â905. Cited by: §I, §I-B. [6] M. W. Brittain, L. E. Alvarez, and K. Breeden (2024-Mar.) Improving autonomous separation assurance through distributed reinforcement learning with attention networks. Proceedings of the AAAI Conference on Artificial Intelligence 38 (21), p. 22857â22863. External Links: Link, Document Cited by: §I, §I, §I-B. [7] M. Brittain and P. Wei (2019) Autonomous separation assurance in a high-density en route sector: a deep multi-agent reinforcement learning approach. In 2019 IEEE Intelligent Transportation Systems Conference (ITSC), p. 3256â3262. Cited by: §I, §I, §I. [8] M. Brittain and P. Wei (2022) Scalable autonomous separation assurance with heterogeneous multi-agent reinforcement learning. IEEE Transactions on Automation Science and Engineering 19 (4), p. 2837â2848. Cited by: §I, §I, §I. [9] Y. Cai and Y. Shen (2019) An integrated localization and control framework for multi-agent formation. IEEE Transactions on Signal Processing 67 (7), p. 1941â1956. Cited by: §I. [10] S. Chen, A. D. Evans, M. Brittain, and P. Wei (2024) Integrated conflict management for UAM with strategic demand capacity balancing and learning-based tactical deconfliction. IEEE Transactions on Intelligent Transportation Systems 25 (8), p. 10049â10061. Cited by: §I-C1, §IV-B1. [11] Federal Aviation Administration (2020) FAA remote identification of unmanned aircraft. Note: Accessed: Aug 30, 2025 Cited by: §I, §I-C1, §IV-A2. [12] D. Groot, J. Ellerbroek, and J. Hoekstra (2025) Comparing attention-based methods with long short-term memory for state encoding in reinforcement learning-based separation management. Engineering Applications of Artificial Intelligence 159, p. 111592. Cited by: §I, §I-B. [13] W. Guo, M. Brittain, and P. Wei (2021) Safety enhancement for deep reinforcement learning in autonomous separation assurance. In 2021 IEEE International Intelligent Transportation Systems Conference (ITSC), p. 348â354. Cited by: §I. [14] J. M. Hoekstra and J. Ellerbroek (2016) Bluesky ATC simulator project: an open data and open source approach. In Proceedings of the 7th International Conference on Research in Air Transportation, Vol. 131, p. 132. Cited by: §IV-A. [15] G. Hunter and P. Wei (2019) Service-oriented separation assurance for small UAS traffic management. In 2019 Integrated Communications, Navigation and Surveillance Conference (ICNS), p. 1â11. Cited by: §I. [16] R. Isufaj, M. Omeri, and M. A. Piera (2022) Multi-UAV conflict resolution with graph convolutional reinforcement learning. Applied Sciences 12 (2), p. 610. Cited by: §I. [17] V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu (2016) Asynchronous methods for deep reinforcement learning. In International Conference on Machine Learning (ICML), p. 1928â1937. Cited by: §I-B. [18] H. Y. Ong and M. J. Kochenderfer (2017) Markov decision process-based distributed conflict resolution for drone air traffic management. Journal of Guidance, Control, and Dynamics 40 (1), p. 69â80. Cited by: §I. [19] J. Schulman, P. Moritz, S. Levine, M. Jordan, and P. Abbeel (2015) High-dimensional continuous control using generalized advantage estimation. arXiv preprint arXiv:1506.02438. Cited by: §I-B. [20] J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov (2017) Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347. Cited by: §I-B. [21] I. Sharifi, A. Zongo, and P. Wei (2026) Fine-tuning large language models for cooperative tactical deconfliction of small unmanned aerial systems. arXiv preprint arXiv:2603.28561. Cited by: §IV-B1. [22] Z. Wang, W. Pan, H. Li, X. Wang, and Q. Zuo (2022) Review of deep reinforcement learning approaches for conflict resolution in air traffic control. Aerospace 9 (6), p. 294. Cited by: §I, §I, §I. [23] C. Yu, A. Velu, E. Vinitsky, J. Gao, Y. Wang, A. Bayen, and Y. Wu (2022) The surprising effectiveness of ppo in cooperative multi-agent games. Advances in Neural Information Processing Systems (NeurIPS) 35, p. 24611â24624. Cited by: §I-B. [24] P. Zhao and Y. Liu (2021) Physics informed deep reinforcement learning for aircraft conflict resolution. IEEE Transactions on Intelligent Transportation Systems 23 (7), p. 8288â8301. Cited by: §I. [25] G. Zhong, Y. Liu, S. Du, F. Wang, J. Zhou, and H. Zhang (2025) 3D RVO-enhanced multi-agent deep reinforcement learning for collision avoidance in urban structured airspace. Aerospace Science and Technology 164, p. 110378. Cited by: §I.