Paper deep dive
Cyber Deception for Mission Surveillance via Hypergame-Theoretic Deep Reinforcement Learning
Zelin Wan, Jin-Hee Cho, Mu Zhu, Ahmed H. Anwar, Charles Kamhoua, Munindar P. Singh
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 97%
Last extracted: 3/26/2026, 2:21:58 AM
Summary
The paper introduces HT-DRL, a hybrid framework combining hypergame theory and deep reinforcement learning to defend Unmanned Aerial Vehicle (UAV) mission systems against Denial-of-Service (DoS) attacks. By deploying honey drones (HDs) as decoys and using hypergame theory to model the evolving, perception-based strategies of both attackers and defenders, the approach optimizes mission performance and energy consumption while reducing learning convergence time compared to standard DRL.
Entities (6)
Relation Signals (4)
HT-DRL â integrates â Hypergame Theory
confidence 100% · HT-DRL identifies optimal solutions... by taking the solutions of hypergame theory into the neural network of deep reinforcement learning.
HT-DRL â integrates â Deep Reinforcement Learning
confidence 100% · HT-DRL... integrating DRL with game theory
Honey Drone â mitigates â DoS Attack
confidence 95% · We adopt cyber deception as a defense strategy, in which honey drones (HDs) are proposed to bait and divert attacks.
HT-DRL â optimizes â UAV
confidence 90% · HT-DRL-based HD approach outperforms existing non-HD counterparts up to two times better in mission performance
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Unmanned Aerial Vehicles (UAVs) are valuable for mission-critical systems like surveillance, rescue, or delivery. Not surprisingly, such systems attract cyberattacks, including Denial-of-Service (DoS) attacks to overwhelm the resources of mission drones (MDs). How can we defend UAV mission systems against DoS attacks? We adopt cyber deception as a defense strategy, in which honey drones (HDs) are proposed to bait and divert attacks. The attack and deceptive defense hinge upon radio signal strength: The attacker selects victim MDs based on their signals, and HDs attract the attacker from afar by emitting stronger signals, despite this reducing battery life. We formulate an optimization problem for the attacker and defender to identify their respective strategies for maximizing mission performance while minimizing energy consumption. To address this problem, we propose a novel approach, called HT-DRL. HT-DRL identifies optimal solutions without a long learning convergence time by taking the solutions of hypergame theory into the neural network of deep reinforcement learning. This achieves a systematic way to intelligently deceive attackers. We analyze the performance of diverse defense mechanisms under different attack strategies. Further, the HT-DRL-based HD approach outperforms existing non-HD counterparts up to two times better in mission performance while incurring low energy consumption.
Tags
Links
- Source: https://arxiv.org/abs/2603.20981v1
- Canonical: https://arxiv.org/abs/2603.20981v1
Trouble viewing inline? Open PDF directly â
Full Text
112,587 characters extracted from source content.
Expand or collapse full text
1 Cyber Deception for Mission Surveillance via Hypergame-Theoretic Deep Reinforcement Learning Zelin Wan, Jin-Hee Cho,Senior Member, IEEE, Mu Zhu, Ahmed H. Anwar, and Charles Kamhoua,Senior Member, IEEE, Munindar P. Singh,IEEE Fellow AbstractâUnmanned Aerial Vehicles (UAVs) are valuable for mission-critical systems like surveillance, rescue, or delivery. Not surprisingly, such systems attract cyberattacks, including Denial-of-Service (DoS) attacks to overwhelm the resources of mission drones (MDs). How can we defend UAV mission systems against DoS attacks? We adopt cyber deception as a defense strategy, in which honey drones (HDs) are proposed to bait and divert attacks. The attack and deceptive defense hinge upon radio signal strength: The attacker selects victim MDs based on their signals, and HDs attract the attacker from afar by emitting stronger signals, despite this reducing battery life. We formulate an optimization problem for the attacker and defender to identify their respective strategies for maximizing mission performance while minimizing energy consumption. To address this problem, we propose a novel approach, calledHT-DRL. HT-DRL identifies optimal solutions without a long learning convergence time by taking the solutions of hypergame theory into the neural network of deep reinforcement learning. This achieves a systematic way to intelligently deceive attackers. We analyze the performance of diverse defense mechanisms under different attack strategies. Further, the HT-DRL-based HD approach outperforms existing non-HD counterparts up to two times better in mission performance while incurring low energy consumption. Index TermsâCyber deception, deep reinforcement learning, game theory, unmanned aerial vehicle, mission effectiveness I. INTRODUCTION Unmanned Aerial Vehicles (UAVs) have been widely adopted in mission systems to improve energy efficiency, foster autonomous control, and enhance communication ca- pabilities [1]. However, ensuring the security of UAVs is an ongoing challenge. Denial-of-Service (DoS) attacks are espe- cially harmful because they can result in data invalidation, data leakage, and physical damage through drone crashes [2]. To solve this problem, we take a cyber deception-based defense This research was partly sponsored by the Army Research Laboratory and was accomplished under Cooperative Agreement Number W911NF-19-2-0150 and W911NF-23-2-0012. In addition, this research is also partly supported by the Army Research Office under Grant Contract Numbers W911NF-20-2- 0140 and W911NF-17-1-0370. The views and conclusions contained in this document are those of the authors and should not be interpreted as representing the official policies, either expressed or implied, of the Army Research Laboratory or the U.S. Government. The U.S. Government is authorized to reproduce and distribute reprints for Government purposes, notwithstanding any copyright notation herein (Corresponding author: Zelin Wan). Zelin Wan and Jin-Hee Cho are with the Department of Computer Science, Virginia Tech, Falls Church, VA, USA. Email:zelin, jicho@vt.edu. Mu Zhu is with the Computer Network Information Center, Chinese Academy of Sciences, Beijing, China. Email: zhumu@cnic.cn. Munindar P. Singh is with the Department of Computer Science, North Carolina State University, Raleigh, NC 27695. Email: mpsingh@ncsu.edu. Ahmed H. Anwar and Charles A. Kamhoua are with the US Army Research Laboratory, Adelphi, MD, USA. Email: a.h.anwar@knights.ucf.edu; charles.a.kamhoua.civ@mail.mil. approach, calleddefensive deception(D). The D strategy can confuse and mislead attackers into choosing sub-optimal attack strategies [3]. We aim to design a surveillance mission system using honey drones (HDs) to combat DoS attacks [4]. HDs, equipped with lightweight virtual machines (VMs) run- ning vulnerable software, can lure potential cyberattacks at higher signal strengths and collect attack intelligence to update system settings in realtime. Deep reinforcement learning (DRL) has been applied to cybersecurity problems [5â7]. However, DRL algorithms are challenged by long convergence times to solutions under non- stationary settings. On the other hand, game theory has been substantially used in cybersecurity to model strategic attack- defense decision-making processes. The synergistic effect of combining game theory and deep learning has been recognized as a promising direction [3]. However, such hybrid approaches have been significantly less studied. We propose a novel hybrid approach,hypergame theory- guided DRL(HT-DRL), integrating DRL with game theory to identify the optimal settings of the proposed honey drone (HD)-based system. By taking the merit of DRL for autonomy and hypergame theory for strategic decision-making, HT-DRL can empower UAVs to make proactive, intelligent decisions in uncertain and adversarial environments. Hypergame the- ory [8] can better handle an agentâs perceived uncertainty than conventional game theory by estimating its utility,hypergame expected utility(HEU). Hypergame theory helps reduce the number of strategies an agent takes based on the subgames each agent perceives towards the opponentâs moves. A game- theoretic solution is more efficient than its DRL counterparts due to no training time needed in game theory. However, in terms of solution optimality, DRL can perform better than game theory. Therefore, taking a hybrid approach by com- bining these two is a natural direction to tackle this problem. In our UAV DoS setting, the attacker does not know which nodes are honey drones versus mission drones and must in- stead infer targets from noisy radio-signal patterns and limited mission information. The defender, in turn, only partially observes how the attacker thresholds signal levels and adapts to the current honeyâdrone configuration. These perceptions evolve over time and may be inaccurate. Conventional game- theoretic defenses for UAVs and cyberâphysical systems, including our prior honeyâdrone study [11], typically assume a single, fully specified game with a fixed payoff structure and optimize strategies within that static model. DRL-based defenses can adapt in high-dimensional mission states but often treat the opponent as part of a stationary environment, arXiv:2603.20981v1 [cs.CR] 21 Mar 2026 2 which leads to instability when the attacker is itself adaptive. Hypergame theory is therefore directly relevant, as it pro- vides a principled framework for representing and updating the differing perceived games held by the attacker and defender. Building on this insight, our HT-DRL framework formulates a dynamic hypergame between an adaptive attacker and a honeyâdrone defender, derives HEUs for both players, and uses the defenderâs HEU to shape the actorâcritic policy logits through a filter layer. This integration enables early exploration and subsequent policy adaptation to be guided by the defenderâs evolving perception of the opponent, rather than assuming a fixed or stationary adversarial model. In sum, ourkey contributionsare as follows: âąWe propose a surveillance mission system using honey drones to effectively thwart cyber threats, protect valuable resources, and maintain mission integrity, which has not been explored in the literature. âąWe provide a novel way of integrating game theory with DRL, called hypergame theory-guided DRL (HT-DRL), to enable a shorter convergence time for better solutions. Prior studies [9, 10] using game theory and DRL have not studied this hybrid approach. âąWe model an attack-defense game where the attacker and defender take intelligent strategies via HT-DRL to evaluate defense strategies under various attack scenarios. âąWe validate the superiority of HT-DRL defense strategy through performance analyses under various attack scenar- ios. Via extensive experiments, we compare the performance of HT-DRL using HDs over non-HD counterparts. Our results demonstrate the outperformance of our approach in mission performance and energy consumption. Due to space constraints, we demonstrate the additional comparative per- formance results and analyses in the supplement document. In our preliminary work [11], we investigated the effectiveness of using honey drones in a UAV-based mission system for a surveillance mission. And the prior considered attack-defense interactions based on game theory. Our prior work [11] investigated a honeyâdrone (HD) defense mechanism that adjusts the signal strengths of mis- sion and honey drones using either a game-theoretic strategy selector or a DRL-based selector. That study focused on determining the appropriate HD signal level and comparing the resulting performance against existing baselines (e.g., HD- based IDS and ContainerDrone) in terms of mission comple- tion, energy consumption, and connectivity. The underlying adversarial interaction assumed a standard game model in which the attacker and defender share a common and accurate view of the environment. Moreover, the DRL agent (using a conventional A2C policy) and the game-theoretic agent each optimized a hand-crafted expected-utility function; the two approaches were evaluated independently rather than being integrated into a unified framework. In contrast, the present work introduces a fundamentally different modeling and algorithmic perspective. We formulate adynamic hypergamein which the attacker and defender may hold distinct and evolving perceptions of the game state and payoff structure. Building on this formulation, we derive hypergame-expected utilities (HEUs) for both players and leverage the defenderâs HEU to construct a filter layer that shapes the DRL policy logits. This coupling of hyper- game analysis with policy learning yields the proposed HT- DRL framework, which explicitly addresses cold-start and non-stationarity challenges posed by adaptive DoS attackers. Rather than offering another empirical comparison between DRL and game-theoretic defenses, this work integrates the two viewpoints into a principled, perception-aware learning architecture that is robust to adversarial adaptation. In this work, we introducedHypergame Theory (HT)[8] to address each playerâs different view about the same game, which reflects real-world scenarios more appropriately. Hence, playersâ action choices are based on the Hypergame Expected Utility (HEU), representing their utility depending on the level of perceived uncertainty. Further, the merit of using HT is to reduce the size of an action space by utilizing the concept of a subgame to avoid a full game with all possible actions. Our prior work [11] introduces the basic UAV surveillance system and simulation environment. In contrast, the present work advances the field by developing a hypergame-theoretic DRL framework that supports scalable HD deployment, incorporates refined energy and threat models, and conducts substantially expanded and more rigorous experimental evaluations. Further, we substantially extended our experimental results and analy- ses based on four key metrics (see Section IV-B) to investigate the effectiveness of HD-based deception defenses. Our proposed HT-DRL framework adapts to varying levels of prior knowledge. When no prior knowledge is available, HT-DRL gracefully degrades to operate as a standard DRL model, ensuring applicability even in scenarios with limited historical data. The frameworkâs flexibility allows it to leverage any available domain knowledge while remaining functional without it. With no prior knowledge, it operates as a full DRL model, while with complete prior knowledge, it fully leverages HT-based decision-making. This flexibility ensures its applicability across diverse scenarios. I. SYSTEMMODEL A. Network Model This study examines a drone fleet deployed for surveillance in a target area. The network consists of a Charging Station (CS) for drone recharging, a Ground Control Station (GCS) responsible for mission assignments, and UAVs for mission execution. We consider three types of UAVs: Regional Leader Drone (RLD), Mission Drones (MDs), and Honey Drones (HDs). Considering the limited transmission range of a GCS, we propose a network architecture where an RLD is connected to the GCS through a satellite network, and MDs and HDs es- tablish connections with the RLD via WiFi, forming a âflying ad hoc networkâ (FANET) [12]. The FANET facilitates multi- hop communications between drones coordinated by the RLD. MDs send real-time data to the RLD. We assume the GCS possesses high computational power and a firewall to filter out malicious streams, making it a trusted entity not vulnerable to DoS attacks (see our threat model in Section I-D). As Fig. 1b shows, each drone maintains two tables: The Neighbor Table (NT) and the Fleet Table (FT). The NT 3 Charging Station HD MD HD Attacker MD MD MD MD MD GCSRLD satellite Low Battery NTFT NTFT NTFT NTFT NTFT NTFT NTFT NTFT lure (a) Example multi-hop communication network of a drone fleet. Neighbor Table IDlatitudelongitudealtitudeTTL 0x ...... ...... ...... ...... ... 10x Fleet Table IDTypeStatus 0RLDON 1MDOFF ... ...... ...... ... 10HDON ... ...... ...... ... (b) Example neighbor table (NT) and fleet table (FT). Fig. 1.(1) Example drone fleet with multi-hop communicationsconsisting of MDs, HDs, RLD, and GCS.(2) Example of the Neighbor Table (NT) and Fleet Table (FT)where NT provides information about the neighboring drones of a given drone, while FT tracks the membership of the mission team, including drones executing missions, leaving the team to recharge batteries, or compromised by the attacker. Drones with a battery level lower than the energy thresholdT e will leave the mission team and move to the charging station. contains location information (i.e., latitude, longitude, and altitude) and the Time-to-Live (TTL) values of neighboring drones [12]. The FT includes a droneâs ID, type, and member- ship to monitor the mission teamâs status. To ensure continuous network connectivity, each drone (i.e., MD or HD) broadcasts a periodic hello message containing its ID and location. Upon receiving a hello message from dronej, droneisearches for dronejâs status information in its NT. If dronejis not found in the NT, droneiinitiates connection establishment and updates its NT accordingly. If dronejis present in the NT, droneiupdates the location information and TTL values based on dronejâs hello message [12]. The connection process begins with a TCP handshake where drones exchange JSON records to determine a communication channel. Droneisends its record to dronej, and dronejresponds with acceptance or rejection based on whether droneiâs ID is found in the FT (for authentication). Once both drones accept the connection, they add each other to their respective NTs. Then, data collection and transmission occur via UDP. We consider the target region composed of multiple discrete cells [13], allowing MD and HD to move between them. To mitigate packet loss in unstable wireless environments, our network model integrates TTI bundling [14]. The receiver can perform soft combining by transmitting repeated copies of a packet in short, consecutive intervals, significantly improving decoding reliability under adverse channel conditions. To facilitate scalability in larger or more complex mis- sions, our network architecture was designed to be inherently modular and could be extended by integrating additional Re- gional Leader Drones (RLDs). These RLDs could individually manage sub-fleets of Mission Drones and Honey Drones, thereby ensuring efficient coordination across an expanded operational area. This scalability approach was supported by existing works: Widhalm et al. [15] demonstrated scalable drone operations via shared service models, and Qin and Pournaras [16] highlighted the effectiveness of decentralized, energy-aware coordination in managing large drone fleets. The computational complexity of our HD deployment al- gorithm (Algorithm 1) isO(|L H |Ă|L M |), where|L H |is the number of honey drones and|L M |is the number of mission drones. Our network architectureâs modular design inherently supports larger deployments through additional RLD coordi- nation. Our hierarchical RLD structure is designed to support scalable deployments through its modular architecture. The algorithm complexity is primarily determined by the product of HD and MD numbers, making it suitable for moderate-scale deployments while maintaining mission effectiveness through parallel RLD coordination. Because the present study employs the same physical UAV platform and communication topology as in [11], certain elements of the system description and simulator configuration unavoidably overlap with our prior work. B. Node Model At the beginning of the mission, the drone fleet is fully charged and takes off from the GCS. Once a droneâs battery is depleted, it leaves the mission team and returns to the CS for recharging. The CS and GCS are not mobile. Fig. 1a illustrates the network model, characterized by the following node types. 1)Charging Station (CS):A drone returns fully recharged before rejoining the mission team by takingT C time. 2)Ground Control Station (GCS):The GCS assigns missions in the target region and sets a maximum mission duration,T max M . The GCS tracks the drone fleet, monitoring mission progress reported by RLDs, e.g., via FBCB2âs BFT (Blue Force Tracker) satellite network [17]. However, due to the high latency, low bandwidth, and instability of satellite communication, the GCS may be unable to track the droneâs locations and signal strengths in real-time. Yet, the RLD is authorized to have timely control over the UAVsâ settings. 3)Regional Leader Drone (RLD):The RLD is a high- altitude long-endurance drone [18]. It possesses sufficient energy to handle heavy computations and stores sensed data over the entire mission duration,T M (â€T max M ). The RLD dynamically configures the dronesâ trajectories and transmits 4 them to MDs. The RLD identifies the optimal signal strength to minimize energy consumption while accomplishing the mission. The RLD moves to the center of the target area to maximize signal coverage. When an MD is offline (i.e., compromised, exhausted in battery, or disconnected from un- reliable wireless connections) or online (i.e., non-compromised with sufficient battery), the RLD adjusts the trajectory for all drones to accommodate the new conditions. The RLD can analyze attack intelligence collected by HDs and update the fleet networkâs configuration by modifying open ports for attacked MDs, thereby avoiding DoS attacks. Although DoS attacks can potentially compromise the RLD, we assume it is equipped with robust defense mechanisms to detect attacks and has sufficient computing power to handle a large volume of requests. In our system, a single RLD exists and may expose a single point of failure. To deal with this, when the RLD is compromised, a backup RLD, which stands by in the GCS, will replace it. During the RLDâs downtime, as it serves as a data storage unit, the mission process is paused. The DRL agent, running on the RLD, controls the signal strengths of MDs and HDs to maximize mission performance. Section I-C specifies how to identify the optimal signal strength to be used by HDs and MDs. 4)Mission Drone (MD):MDs are equipped with Himax HM01B0 ULP monochrome QVGA cameras [19]. They trans- mit sensed data to the RLD through the FANET. Each MD ini- tially follows an assigned trajectory upon deployment. When an MD goes to recharge, it notifies the RLD offline. However, if the RLD does not receive such a signal but detects that the MD is offline, it assumes the MD has been compromised. Since an MD can go offline due to terrain, obstacles, or unreliable wireless connections, a non-compromised MD may appear compromised. When such drones reconnect, NT and FT (see Fig. 1b) are updated accordingly. We assume the drones have basic knowledge of the target region [13]. Each drone is loaded with an optimized trajectory to complete the mission. If a drone goes offline, the trajectory is recalculated for others based on the scanned and completed cells. 5)Honey Drone (HD):Each HD consists of two logi- cally isolated components: a lightweight honey VM and an infrastructure VM. The honey VM exposes multiple open ports to attract DoS attacks. The infrastructure VM, running on a lightweight Linux environment, monitors the memory and log of the honey VM. Upon detecting a DoS attack based on abnormal computing consumption (e.g., CPU and memory usage), the infrastructure VMâs backchannel informs the RLD of the port being used by the attacker for communication. Subsequently, the RLD reconfigures the open port of the MD to prevent further attacks. In case the decoy system in the honey VM is compromised or malfunctions, the honey VM can restore it. This honeypot can be effectively implemented on a Samsung Galaxy S2 smartphone, released in 2011 [20], indicating it would not overload the drone. HDs are deployed in a greedy manner following Algo- rithm 1, where thed(sg HD ) = 10 sg HD +100 η·10 is derived from Eq. (3). The deployment of HDs in Algorithm 1 considers dronesâ energy concerns by limiting MDs assigned to each HD for monitoring. That is, we consider an upper bound that Algorithm 1Honey Drone Deployment 1:L M âA set of active and not in GCS MD locations 2:L H âA set of active HD locations 3:P H r âd(sg HD ),sg HD âDS 5 â·The signal radius/range of an HD when RLD selectssg HD =DS 5 for the HD 4:D(x,y)âThe distance between two dronesxandy 5:S=â â·A set of deployed HD locations 6:[Ï l ,Ï u ]âThe lower and upper bounds of the number of MDs 7:forl H âL H do 8:if|L M |= 0then 9:SâSâȘl H 10:else 11:N(l H ) =D(l H ,l M )< P H r :l M âL M â·A set of MDs in the protect range of HDl H 12:if|N(l H )|< Ï l then 13:Find a new positionl âČ H such thatÏ l â€|N(l âČ H )|â€Ï u 14:whereN(l âČ H ) =D(l âČ H ,l M )< P H r :l M âL M â·A set of MDs detected/protected by HDl âČ H 15:L M âL M (l âČ H )â·Remove protected MDs from set L M 16:SâSâȘl âČ H â·Add deployed HD to setS 17:else ifÏ l â€|N(l H )|â€Ï u then 18:L M âL M (l H ) 19:SâSâȘl H 20:else 21:N âČ (l H )âN(l H ), where|N âČ (l H )|=Ï u â·Select the nearestÏ u MDs fromN(l H )and assign them to HDl H 22:L M âL M âČ (l H ) 23:SâSâȘl H 24:end if 25:end if 26:end for 27: Output: HD location setS one HD can simultaneously protectÏ u number of MDs. Each HD is assigned a set of MDs betweenÏ l andÏ u to protect and monitor. Some HDs may not be assigned any MDs and may move around to find available MDs. In this case (i.e., line 13), the HD finds a location whenÏ l â€|N(l âČ H )|â€Ï u . We search only MDsâ positions to reduce computational complexity instead of all cells. This will lead to reducing the complexity of Algorithm 1 fromO(|L H ||N cell |)toO(|L H ||L M |), where |N cell |is much higher than|L M |. C. Energy Model Our simulation (see Section IV) considers Crazyflie 2.X quadrotor drones [21] with the âgym-pybullet-dronesâ sim- ulator [22]. Both MDs and HDs consume energy as follows: E MD =E P +E C + E R ·DS j 10 , E HD =E P + E R ·DS j 10 , (1) whereE C represents the consumption rate of the Himax camera [23].E P corresponds to the power consumed by the drone platform, including the Standard Operating Conditions (SoCs) and four motors. The estimation ofE P is derived from the platformâs flight time.E R represents the maximum consumption rate of the radio. As the defense strategy controls the signal level,E R Ă DS j 10 represents the real-time energy consumed by taking a given defense strategy. The energy consumption model is based on real-world hardware specifications, with power values calibrated from Bitcraze Crazyflie 2.X [21] and Himax camera data [19] and validated against empirical measurements reported in the literature [23]. Although not exhaustive, this model offers a 5 reliable baseline for typical UAV operations, with future work planned to incorporate additional factors such as battery aging. D. Threat Model Recently, security threats and their serious impact on UAV systems have been recognized [2, 24â27]. Particularly, the serious adverse impact of DoS attacks on UAV systems has been a serious security issue, while the DoS attacks are simple to launch but can cause serious data leakage and crashing drones [26, 27]. Since UAVs often should perform real-time communications under high dynamics (e.g., node join, leave, or failure) and resource constraints, draining UAVsâ energy and preventing their communications by DoS attacks can in- troduce a critical impact to the mission system and easily cause mission failure [2]. Therefore, this work primarily focuses on providing a security solution to address DoS attacks. To estimate a droneâs software vulnerability, we adopt the Common Vulnerability Scoring System (CVSS) [28]. We model each droneâs software vulnerability using the CVSS score of an Android device [29]. We represent a drone Îșâs software vulnerability as a real value,vul Îș â[0,1], indicating the likelihood of a successful attack. We do not assume attackers can capture dronesâ vulnerability through reconnaissance attacks, which is possible when the attackers can observe a target system for an extended period because this concerned mission system will be assigned a short-term mission. Therefore, we consider intelligent attack strategies in choosing their signal strengths in Section I-B. We will consider DoS attackers sitting on the ground to ensure their infrastructure needs, such as high-power antennas and a stable power source. In addition, for the attackers to avoid physical detection of their presence, we do not consider attackers who are physically proximate to the UAV mission team. Each droneâs presence and its connectivity with the UAV network can be detected by the signal strength level [30]. In addition, recipients receiving such strong signals can save their energy [31]. A Denial-of-Service (DoS) attacker is more likely to target UAVs with stronger signal strengths to maximize the impact of their attacks while conserving energy. Due to the inverse relationship between signal-to-noise ratio (SNR) and packet error rate (PER) [32], stronger signals enable the attacker to deliver malicious packets more reliably and efficiently, ensuring that a higher proportion of these packets reach the UAV and effectively overwhelm its communication system [33]. This increased reliability enhances the overall effectiveness of the attack, leading to greater disruption of the UAVâs operations. Additionally, targeting UAVs with robust signal strengths allows the attacker to use lower transmis- sion power [34, 35], thereby conserving energy compared to attacking UAVs with weaker signals, which would require higher power levels to achieve the same packet delivery success. Attackers can optimize their resources by focusing on UAVs with stronger signals to sustain prolonged attacks with maximum disruptive potential. We assume an attacker is also limited to its computing power and energy resources. Hence, we consider the attackerâs budget,ζ, representing the maximum number of drones to launch its attack. Largerζ means more severe attack strength, whose impact is analyzed in Section V. While our current focus is on DoS attacks due to their prevalence and severe impact on UAV systems, the underlying principles of our HT-DRL approach can be adapted to address other cyber threats with suitable modifications in utility models and defense strategies. The modular design of our framework facilitates such extensions. Future work will extend our frame- work to address additional attack types such as data tampering and spoofing. I. STRATEGYSELECTION INHT-DRL A challenge in mission systems is the lack of prior knowl- edge about attack behaviors. Real-time defense cannot rely on instructions from a central entity due to the high delay of satellite networks. Therefore, an adaptive mechanism is needed to learn and make decisions. DRL algorithms support autonomous interactions with the environment and learning in- dependently without prior knowledge. We use DRL to identify an optimal signal strength to maintain network connectivity while minimizing vulnerabilities to DoS attacks. We formulate attack-defense interactions as a series of simulations by the attacker and defender to achieve their respective goals during mission execution. The missionâs com- pletion time isT M .T max M is the maximum time allowed for mission completion. IfT M > T max M , it indicates the partial completion, which means mission failure. The mission consists of multiple rounds of interaction between the attacker and defender, resulting in mission success or failure. The mission team of RLD, multiple MDs, and multiple HDs takes off from the GCS and continues until the mission concludes. In this work, two players, the attacker and defender, are given 10 strategies representing 10 different levels of signal strengths. This number was determined based on our experi- ments and showed a sufficient level of diverse strategies while it does not introduce too high computational complexity. As our work adopts hypergame theory, we use the concept of a subgame each player can use to reduce its action space, lowering solution search complexity. However, when there is uncertainty, they will use a full game with all 10 strategies as the action space. We described each playerâs hypergame theoretic decision-making process in Sections I-B and I-C. A. Key Procedures of HT-DRL The balance between exploration and exploitation in solu- tion search by DRL agents is critical. The dynamicΔ-greedy exploration allows initial random decision-making under a lack of knowledge while more exploitation is used as the agent accrues knowledge. However, this strategy often introduces a long convergence time to a close-to-optimal solution, partic- ularly in scenarios without prior knowledge. Therefore, we introduce HT-DRL, which uses Hypergame Expected Utility (HEU) to guide a DRL agent to choose its best action. That is, instead of randomly exploring solutions at the beginning, the agent can use HEU-based action distributions for DRLâs policy function at the beginning to resolve the cold start problem without being stuck at local optima. The hypergame outcomes 6 Filter Probability Distribution from HT Agent Input Output ) Create a HT Agent Probability Distribution of HT Agent's Action Interact with environment Create a DRL Agent Combine with HT's Distribution HT-DRLAgent Fig. 2.The procedures generating the solutions by a HT-DRL agent:S t is the state at roundtandÏ(a t |S t )is the probability of all actions. are integrated into the DRL framework to guide exploration, ensuring that the agent learns under conditions that reflect real- world uncertainty. We analyze HT-DRLâs performance and compare it with the performance of other baseline and existing counterparts in Section V. HT-DRL achieves its outperformance by incorporating a unique HT-guided filter layer into the neural network and taking the following procedures: 1) Utilize prior knowledge to train an HT agent to choose its action based on the HEU. Due to the characteristics of HT, only a small amount of data is required for this training. 2) Collect the HT agentâs action choices to construct its action probability distribution (APD). 3) Instantiate a DRL agent using the Advantage Actor-Critic (A2C) technique. We add a filter to the output layer to alter the action distribution by multiplying each output value by a corresponding weight. 4) Apply the APD from the pre-trained HT agent to each filter weight so that the DRL agent can perform a more strategic rather than completely randomized initial exploration. 5) Train this HT-DRL as a standard DRL for further fine- tuning parameters. We summarize the procedures above in Fig. 2. a) Hypergame-Theoretic Game Formulation:We model the attackerâdefender interaction as a finite-horizon stochastic gameG=âšP,S,A p pâP ,T,u p pâP â©, where: âąP=A,Dis the set of players: the attackerAand defenderD. âąSis the state space, containing the mission progress, drone connectivity, and scan map, e.g.,S A/D t = (R t MC ,M t SP ,N t TR ). âąA A andA D are the action spaces of attacker and defender, where each action corresponds to a discretized signal- strength range (AS i for the attacker,DS j for the defender). âąT:SĂA A ĂA D ââ(S)is the transition kernel induced by the UAV dynamics, energy consumption, and DoS effects in the simulator. âąu A andu D are the instantaneous utilities of the attacker and defender, derived from the gains and losses in (5) and (18). Classical stochastic games typically assume that both players share the same gameGand have correct beliefs about each otherâs strategies. In contrast, our setting is naturally modeled as ahypergame, where each playerpâA,Dmaintains its own perceived game b G p with a perceived utilityu p and belief C p ÎŁ over the opponentâs actions. The hypergame expected utili- ties (HEUs) in (15) and (29) are computed on these perceived games and used to construct practical HEU-guided policies, which we subsequently integrate into the DRL framework. Hyper Nash Equilibrium (HNE) [36] was analyzed in detail in our prior work [37], and a full HNE analysis is beyond the scope of the present study. B. Attacker Model 1)Attackerâs Action Space based on Signal Strengths: The attacker observes the dronesâ signal strengths and selects its attack strategy accordingly. We define the attack strategy asAS i â AS 1 ,...,AS 10 , where each action corresponds to a range of received signal strengths[sg l i ,sg u i ]chosen by the attacker to identify the target drone set,S target,i , for the DoS attack. IfS target,i =â , it implies no attack. Thus,S target,i is given by: S target,i =Îș|sg l i â€sg Îș â€sg u i ,(2) whereÎșis a droneâs ID, andsg Îș is the signal strength re- ceived by the attacker from droneÎș. The resource-constrained attacker can target at mostζdrones. We map the attackerâs strategies into 10 signal strength ranges indBmby: [sg l i ,sg u i ]â(â100,â98.1],(â98.1,â96.1],(â96.1,â93.8], (â93.8,â91.1],(â91.1,â87.9],(â87.9,â84.0],(â84.0,â79 .0],(â79.0,â72.0],(â72.0,â60],(â60,20]. These ranges are formed based on Eq. (3), which evenly groups drones based on the distance between the defender and attacker and converts it to the received signal strength. We estimate signal attenuation by [38]: P dBm (d) =P dBm (d 0 )âη·10·log 10 ( d d 0 ),(3) whereη= 4is a path loss exponent, andP dBm (d)and P dBm (d 0 )are the observed signal strengths at distancedand d 0 , respectively. According to the droneâs specifications [21], P dBm (d 0 ) = 20dBm(decibel-milliwatts) whend 0 = 1 m(meter). Sinceâ60dBmandâ100dBmare typical values for the strongest and lowest signal strength that a drone uses [39], we consider a signal to be strong whend <100 m, and signal weak whend= 1000m. Based on real-world test results [21], the maximum control range of1000mwell reflects a real-world scenario. The distance,d, between an attacker (A) and a drone (Îș) is calculated based on their coordinates,(x,y,z): d(A,Îș) = q |x A âx Îș | 2 +|y A ây Îș | 2 +|z A âz Îș | 2 ,(4) 7 2)Hypergame Theoretic Attack Strategy Selection: This section discusses how an attacker selects its best strategy based on the hypergame expected utility (HEU). Our frame- work builds upon established hypergame theory to provide practical equilibrium solutions suitable for the dynamic UAV environment, where traditional game-theoretic assumptions may not fully apply due to resource constraints and real-time operational requirements. a) Attackerâs Utility Calculation:An attacker estimates the HEU as a function of its utility. When the defender takes strategyj, the attackerâs utility,u A ij , by taking strategyiis: u A ij =G A ij âL A ij ,(5) G A ij = ai A ij + dc A ij , L A ij = di A ij + ac A ij ,(6) whereG A ij andL A ij are gains and losses of the attackerG A ij includes attack impact,ai A ij , and defense cost,dc A ij .L A ij is based on defense impact,di A ij and attack cost,ac A ij . Attack impact,ai A ij , is computed by: ai A ij = P Îș j âS A target,i ASR âČ Îș j C Îș j ζ ,(7) whereS A target,i is the target drone set when attacker selects strategyi, andζis the attack budget.ASR âČ Îș j is the expected attack success ratio for droneÎș j . It is calculated as the ratio of successful DoS attacks to the total number of DoS attacks performed by the attacker up to roundt.C Îș j is the criticality of droneÎș j . Since the network topology of the drone fleet is unknown to the attacker, it determines a droneâs criticality based on the observed signal strength. Attack cost,ac A ij , is determined by the size of the target drone set,S target,i , and estimated by: ac A ij =e |S target,i |âζ .(8) Defense impact,di A ij , is the opposite of the attack impact: di A ij = 1âai A ij .(9) The defender takesj, the defense cost perceived by the attacker,dc A ij , is calibrated by the droneâs energy consumption and the impact by a compromised drone, and estimated by: dc A ij = j sig max + ai A ij ,(10) wherejis the defense strategy perceived by attacker, and sig max is the maximum signal (i.e.,10in this work). b) Attackerâs Column-Mixed Strategies (CMSs):The probabilities of the CMSs are derived from the attackerâs experience when each subgamekconsists of multiple defense strategies [8]: CMS A k = [c A k1 ,...,c A km ],(11) where m X j=1 c A kj = 1, c A kj = Îł A kj P jâDS k Îł A kj . Herec A kj follows the Dirichlet distribution andmis the number of the defenderâs strategies.Îł A kj is the number of times the defender takesDS j when the attacker plays subgamek. Since the attacker cannot capture dronesâ true signal strength levels due to attenuation,Îł A kj andÎł A j are estimated based on the received signal strength levels. c) Attackerâs Belief Contexts:The attackerâs belief con- texts are modeled by a subgamek[8]. The attackerâs set of probabilities taking each subgame is formulated by: P A = [P A 0 ,P A 1 ,P A 2 ,P A 3 ],where 3 X k=0 P A k = 1.(12) HereP A 0 is the probability of taking a full game with all possible observed signals by the attacker, whereP A 1 , andP A 2 , P A 3 are the probabilities of taking a game with signal strengths (â100,â93.8],(â93.8,â79.0], and(â79.0,20], respectively. Those ranges are designed as discussed in Section I-B1. A subgame reduces the solution space for efficiency. In hyper- game theory, the concept of a subgame (e.g.,P A 1 ,P A 2 ,P A 3 ) is used to reduce the cost of computing all utilities associated with each strategy. When an attacker is uncertain about the game (i.e.,P r < g A , whereP r is a predefined value, andg A is the attackerâs perceived uncertainty), the attacker takes a full game,P A 0 . Otherwise, each subgameâs probability,P A i , is the number of drones using a signal strength in the signal range over the total number of drones observed in a round. Across all subgames, the attackerâs belief in the defender takingjstrategy is estimated by: S A j = 3 X k=0 P A k ·c kj ,where m X j=1 S A j = 1.(13) HereP A k andc kj are explained in Eq. (12) and Eq. (11), respectively. A set of the attackerâs belief in the defender takingmstrategy is denoted byC A ÎŁ =S A 1 ,...,S A m . d) Attackerâs Uncertainty:We model an attackerâs per- ceived uncertainty,g A , as a function of the degree of de- tectability toward given deception (i.e., honey drones) and the amount of successful experience in launched attacks. More specifically, we formulate this using an exponential decay function which shows the reduction in uncertainty as the attacker experiences more attack success and deception detection by: g A =e âλ A ·ad·N AS ,(14) whereλ A is a parameter for the attacker to control the range of uncertainty,adis the attackerâs deception detectability, randomly selected in the range of[0,0.5], andN AS is the number of attack successes since the mission begins. The modeling of the attackerâs perceived uncertainty accounts for its increase as the attacker detects whether the defender employs deception and successfully executes attacks. This approach is well-supported by existing literature [40, 41]. e) Attackerâs Hypergame EU (AHEU):The AHEU by taking strategyAS i is formulated as: HEU(AS i ,g A ,C A ÎŁ ,DS A w )(15) = (1âg A )·EU A (AS i ,C A ÎŁ ) +g A ·EU A (AS i ,DS A w ), whereg A is given in Eq. (14),EU A (C ÎŁ )is the attackerâs ex- pected utility (AEU), calculated by Eq. (16), andEU A (DS w ) 8 is the AEU when the defender takeswstrategy, producing the lowest utility to the attacker, computed by Eq. (17). The attackerâs EU,EU A (AS i ,C ÎŁ ), is estimated based on the attackerâs belief,S D j , and utility,u A ij , where the attacker and defender takeiandjstrategies, respectively, and given by: EU A (AS i ,C A ÎŁ ) = m X j=1 S A j ·u A ij ,(16) whereC A ÎŁ is a set of the attackerâs beliefs toward all defense strategies.S A j is estimated in Eq. (13) andu A ij is given in Eq. (5).EU A (AS i ,C A ÎŁ )is considered when the attacker is certain about the game. However, when the attacker is uncertain about the game (i.e.,g A ), it considers the worst situation and uses the following AEU: EU A (AS i ,DS A w ) =m·S A w ·u A iw ,(17) where the defender takes strategywwith the lowest EU to the attacker, andmis the number of defense strategies. 3)Attack Strategy Selection in DRL:We employ the DRL strategy selection algorithm for the attacker. Using a single DRL agent, the attacker identifies the optimal attack strategyAS i that maximizes its accumulated reward,G A , by: âąState(S A t ):S A t = (N t TR ), whereN t TR is the number of drones in each signal strength range at roundt. âąAction Set(A A ):A A =a 1 ,...,a i ,...,a n , wherea i is AS i that determines the set of target drones,S target,i . Action ithe attacker DRL agent takes at roundtis represented bya t i . Each action in the set is aligned with the attacker strategies as discussed in Section I. âąReward Function(R A t (a t i )):R A t (a t i ) =N t MNC , where N t MNC represents the number of mission tasks not com- pleted in roundt. The attackerâs DRL agent aims to maxi- mize its accumulated reward,G A = P â t=0 Îł A t ·R A t , where Îł A is the decay factor. These designs are also employed by the DRL agent in our proposed HT-DRL. Specifically, HT-DRL begins by utilizing HT during the initial stages of the learning process and subsequently transitions to a DRL approach for solution opti- mization, as described in Fig. 2. C. Defender Model 1)Defenderâs Action Space based on Signal Strength: We use a defense strategy,DS j âDS 1 ,...,DS 10 , to adjust the signal strength of HDs,sg HD . MDsâ signal strength levels are determined bysg MD =sg HD âÏ, whereÏis a predefined integer to give a higher signal strength to HDs (see Table I) than MDs. We evenly split the signal transmission range from 100m to 1000m based on Eq. (3) and mapsg HD â â20,â7.9,â0.9,4.0,7.9,11.1,13.8,16.1,18.1,20.We leverage DRL to identify the optimal defense strategy,sg HD , for adjusting MDsâ and HDsâ signal strength levels. 2)Hypergame Theoretic Defense Strategy Selection: The defender will estimate its HEU and choose defense strategy (DS j ) to determinesg HD andsg MD =sg HD âÏ. a) Defenderâs Utility Calculation:The defenderâs utility by taking strategyjwhen the attacker takes strategyiis: u D ji =G D ji âL D ji ,(18) G D ji = di D ji + ac D ji , L D ji = ai D ji + dc D ji ,(19) whereG D ji andL D ji are the defenderâs gains and losses. Defense impact,di D ji , is measured by: di D ji = 1â P ÎșâS D target,i vul Îș ζ + N âČ connect,j N drone ,(20) wherevul Îș refers to the vulnerability of droneÎșas a real number in[0,1], as in Section I-D. The number of target drones perceived by the defender inS D target,i is based on experience. The defender keeps track of which drones are targeted when the attacker selects strategyi. Theζis the attack budget.N âČ connect,j is the expected number of connected drones after selecting defense strategyj, andN drone is the total number of drones initially assigned to the mission team. Defense cost,dc D ji , is: dc D ji =e jâsig max ,(21) wherejis the defense strategy andsig max refers to the maximum signal (i.e., 10). The attack impact,ai D i , is given by: ai D ji = 1âdi D ji (22) wherevul Îș j is used as the probability for droneÎș j to be compromised, as discussed in Section I-D. The attack cost,ac D ji , is estimated as: ac D ji = |S D target,i | ζ ,(23) where|S D target,i |is based on the defenderâs prediction towards attack strategyi. b) Defenderâs Column-Mixed Strategies (CMSs):The probability of CMS when the defender playsk-th subgame is obtained by: CMS D k = [c D k1 ,...,c D kn ],where n X i=1 c D ki = 1,(24) wherec D ki follows the Dirichlet distribution withc D ki = Îł D ki P jâAS k Îł D ki andnis the count of the attackerâs strategies. Here Îł D ki is the number of times the attacker takesAS i based on the defenderâs observations in subgamekduring the observation window fromt= 0totâ1at roundt. c) Defenderâs Belief Contexts:The defenderâs beliefs are formulated by: P D = [P D 0 ,P D 1 ,P D 2 ,P D 3 ],where 3 X k=0 P D k = 1.(25) P D 0 is the probability of taking a full game with all strategies andP D 1 ,P D 2 , andP D 3 are the probabilities of the defender taking subgame 1, 2, and 3 with1,2,3,4,5,6, and 7,8,10, respectively. When the defender is uncertain about the game (i.e.,P r < g D whereP r is a random number with 9 MD MD MD MD Honeypot Honeypot RLD Ground Truth MD MD MD MD MD Honeypot Honeypot RLD Defender's View MDs' software vulnerabilities Historical record of compromised MDs Expected target set Defense utility Defender's column mixed strategies Defender's subgame Defender's belief in attacker DHEU Defender's perceived uncertainty Choose the optimal signal threshold for Take strategy MD MD Honeypot Attacker's View Defender Model Attacker Attacker HEU i MD Attacker's subgame Accumulated attack success/failure result Attack utility subgame-based HEU Attacker's belief in defender Attacker's perceived uncertainty Attacker's deception detectability AHEU Choose signal threshold, Choose target set Take attack strategy Historical records of received signals Attacker Model HEU i observation action result observation subgame-based HEU action result action add Attacker's Received Signal drone 15 drone 28 drone 34 action Fig. 3. Hypergame-Theoretic Strategy Selection Methods by the Attacker and Defender. uniform distribution in[0,1]andg D is estimated by Eq. (28)), it takes the full game,P D 0 . Otherwise (i.e.,P r â„g D ), it takes a subgame with probabilityP D k , estimated by: P D k = X jâB k n X i=1 c D ki ·u D ji ,(26) whereB k is the set of strategies available to the defender in subgamekandiis an attack strategy. The defenderâs belief in the attacker takingistrategy is estimated by: S D i = 3 X k=0 P D k ·c D ki ,where 9 X i=0 S D i = 1,(27) where the attacker has 10 strategies, representing integer signal strength in[0,9]. The defenderâs belief on attackerâs nstrategies is denoted byC D ÎŁ =S D 1 ,...,S D n . d) Defenderâs Uncertainty:The defenderâs perceived un- certainty is denoted byg D . The defenderâs uncertainty reduces as more attack alerts are received in honeypots. These alerts give the defender insights into attack patterns. Accordingly, g D is estimated by: g D =e âλ D ·N alert ,(28) whereλ D is a parameter for the defender to control the uncertainty andN alert refers to the number of alerts provided by honeypots about DoS attacks. e) Defenderâs Hypergame Expected Utility (DHEU):The DHEU of each defense strategyDS j is calculated by: HEU(DS j ,g D ,C D ÎŁ ,AS D w )(29) = (1âg D )·EU D (DS j ,C D ÎŁ ) +g D ·EU D (DS j ,AS D w ), whereg D is the uncertainty perceived by the defender and calculated by Eq. (28).EU D (DS j ,C D ÎŁ )is the expected utility of the strategy calculated by Eq. (30).EU D (DS j ,AS D w )is the expected utility when the attacker selects the strategy that gives the defender the worst result and is estimated by Eq. (31). A defenderâs EU,EU D (DS j ,C D ÎŁ ), is estimated based on its beliefs about attacker moves,S D i , and its utility,u D ji , where the defender and attacker takejandi, respectively. The defenderâs EU is formulated by: EU D (DS j ,C D ÎŁ ) = n X i=1 S D i ·u D ji ,(30) whereC D ÎŁ is a set of the defenderâs beliefs toward attack strategies in all possible subgames, whereS D i is estimated by Eq. (27) andu D ji by Eq. (18). When the defender is uncertain about the game, it estimates the expected utility under the worst case: EU D (DS j ,AS D w ) =n·S D w ·u D wj ,(31) whereAS D w is attacker strategy,w, providing the lowest expected utility to the defender andnis the total number of attack strategies. 3)Defense Strategy Selection in DRL:The defenderâs DRL agent optimizes the signal strength of the HDs to maximize its total accumulated reward,G D . This optimization procedure involves the following components: âąState(S D t ): The state at roundt,S D t , is a tuple composed of the mission completion ratio and the scan progress map, defined asS D t = (R t MC ,M t SP ). Here,R t MC represents the ratio of completed mission tasks at roundt. It is a real number between 0 (indicating no tasks have been completed) and 1 (indicating all tasks have been completed). M t SP is a map indicating the scan progress for each cell in the target area at roundt. Each cell value in this map reflects the level of scanning progress, providing a detailed snapshot of the surveillance status of the target area. âąAction Set(A D ): The action set is denoted byA D = a 1 ,·,a j ,·,a m where each actiona j is a defense strategyDS j indicating the signal strength of the HDs. For MDs, we follow in Section I-C1. The actionjselected by the defender in roundtis denoted bya t j . Each action in the set is aligned with the defender strategies as discussed in Section I. âąReward Function(R D t (a t j )): The immediate reward for the defender upon executing actiona t j is given byR D t (a t j ) = 10 Ground Truth MD MD MD MD MD HD HD RLD Defender's View Attacker's View Attacker Attacker State observation A3C reward Choose signal threshold Choose target set Take attack action Attacker Model State observation A3C reward Defender Model Choose signal level, Take defense action MD MD MIssion Drone (MD) Honey Drone (HD) HD Regional Leader Drone (RLD) MD MD HD MD MD action action Fig. 4. Conceptual model of the proposed honey-drone mission system in which both the DoS attacker and the defender employ Deep Reinforcement Learning to select their optimal signal strengths. This figure is adapted from our prior design [11]. TABLE I KEYDESIGNPARAMETERS ANDDEFAULTVALUES SymbolMeaningDefault ÏSignal strength decrement interval5 T C Time duration of battery being charged30 T max M Maximum mission duration150 [Ï l ,Ï u ]The maximum number of MDs that an HD can protect simultaneously [2,4] E P Energy consumption rate by a drone7,900 mW E C Energy consumption rate by a camera4 mW ζMaximum number of targeted drones by the attacker in a single round 5 N t MC . In this equation,N t MC represents the number of mis- sion tasks completed in roundt. The defenderâs DRL agent will maximize its accumulated reward,G D = P â t=0 Îł D t · R D t , whereÎł D is the decay factor. These designs are also employed by the DRL agent in our proposed HT-DRL. The procedure of HT-DRL is described in Fig. 2. Fig. 3 shows how the attacker and defender take their best strategies based on the estimated HEU. Fig. 4 describes the high-level description of how the attacker and defender use their DRL agents to take strategies. Regarding the transition functions and the simulation envi- ronment, our approach employs a model-free DRL method- ology, which does not require explicit knowledge of the environmentâs transition dynamics. Instead, the DRL agents learn optimal policies through continuous interaction with a simulated environment that accurately reflects the operational dynamics of the UAV fleet and the adversarial nature of DoS attacks. The simulation environment, detailed in Section IV, encompasses key factors such as drone mobility, signal attenu- ation, and energy consumption, providing a realistic setting for the agents to explore and adapt their strategies. By leveraging a model-free approach, we ensure that the agents can effectively learn and optimize their strategies in response to evolving attack patterns without predefined transition models. This design choice reinforces the practicality and adaptability of our proposed HT-DRL framework, demonstrating its capability to intelligently deceive attackers and maintain mission integrity in dynamic and uncertain environments. IV. EXPERIMENTALSETUP A. Simulation Environment Setup We evaluated the efficacy of our proposed mission sys- tem through extensive experiments in a simulated environ- ment based on PyTorch deep learning library and the Net- workX graph library. Our simulation was executed on a high- performance computing cluster with 128-thread AMD EPYC 7702 CPUs. Our simulation environment reflects operational UAV dynamics, including realistic mobility, signal attenuation, and energy usage patterns derived from actual hardware spec- ifications. This approach ensures that our simulation outcomes closely approximate those expected in live deployments. Our experiments utilized the A2C algorithm. This synchronous approach maintains an ideal balance between efficiency and stability, unlike its asynchronous variant, A3C, which often in- troduces unnecessary complexity and unpredictability through multiple parallel workers. We allocated a 750mĂ750m target area, subdivided into 25 cells of 150mĂ150m each. The GCS and CS were stationed outside the target area. The drone fleet was composed of 15 MDs and 5 HDs. A drone was classified as having crashed if its altitude dropped below 0.1 meters. Several factors can trigger such a crash: navigation errors induced by DoS attacks, inter- drone collisions, loss of balance due to improper operations, and battery depletion. We generate the trajectory of MDs with OR-Tools library [42], a Python-based efficient, close- to-optimal path-finding algorithm [43]. Using this algorithm, the drone fleet can optimally utilize a subset of available MDs to accomplish a mission if the number of MD exceeds the requirements. In situations where energy depletion occurred mid-mission, the remaining drones at the GCS were deployed to replace the energy-depleted ones. However, if no drones were available, the mission team had to pause operations and wait for a recharged drone. Additionally, if a drone crashes, RLD will reactivate the path-finding algorithm to assign a new drone from GCS. If no replacement drones were present at the GCS during such an event, the mission performance could be adversely impacted. We considered an energy consumption model as detailed in Section I-C. For high experimental validity, we conducted 100 runs for each scenario. The mean values of the relevant results were used for further analysis. We summarize key design parameters of the simulation environment and their default values based on real drone specifications and behaviors in Table I. We adopted the Advantage Actor-Critic (A2C) algorithm for the deep learning agent, a synchronous approach that effectively balances efficiency and stability. We set the reward decay fac- 11 HD-FIDSCDHD-HT-DRLNo-Defense 020406080100120140 Episode 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1.0 MC (a)R MC under fixed attack. 020406080100120140 Episode 0.2 0.3 0.4 0.5 0.6 0.7 0.8 MC (b)R MC under HT attack. 020406080100120140 Episode 0.4 0.5 0.6 0.7 0.8 MC (c)R MC under DRL attack. 020406080100120140 Episode 0.5 0.6 0.7 0.8 MC (d)R MC under HT-DRL attack. Fig. 5. Performance comparison of the considered defense approaches with respect to the ratio of completed mission tasks (R MC ), given an attack strategy. HD-FIDSCDHD-HT-DRLNo-Defense 020406080100120140 Episode 500 600 700 800 900 1000 (a)ECunder fixed attack. 020406080100120140 Episode 400 500 600 700 800 (b)ECunder HT attack. 020406080100120140 Episode 500 600 700 800 (c)ECunder DRL attack. 020406080100120140 Episode 500 600 700 800 900 (d)ECunder HT-DRL attack. Fig. 6. Performance analysis of HD-HT-DRL, IDS, CD, and fixed defense, given an attack strategy with respect to energy consumption (EC). tor (Îł) at0.99and initialized the learning rate at0.0005with a subsequent decay factor of0.9. The neural networks em- ployed LeakyReLU activation functions for all hidden layers, promoting non-linearity and mitigating the vanishing gradient problem. Network parameters were initialized using the Xavier normal distribution to maintain appropriate variance across layers, facilitating faster convergence. Further, we incorporated layer normalization and a dropout rate of 0.01 for each layer to enhance training stability and prevent overfitting. The policy and value networks of the A2C are structured as multilayer perceptrons with internal layers arranged as (128, 128, 64, 32). To ensure stable training, we implemented gradient clipping, which helps prevent excessively large updates that could desta- bilize the learning process. We did not employ a target network during training, as the synchronous nature of A2C obviates the need for one. We used a memory buffer size of10,000 and a mini-batch size of32, applying prioritized experience replay [44] to efficiently learn from key experiences. Since the DRL agent cannot be trained until the memory is full, all the results are generated after the memory population. While our HT-DRL framework is method-agnostic and can be integrated with other policy gradient methods such as PPO or SAC, in this work, we employ A2C due to its demonstrated stability and convergence in our simulation environment. The key novelty of our approach lies in the hypergame-theoretic guidance mechanism, which mitigates the cold-start problem independently of the underlying DRL backbone. Future work will explore comparative evaluations of HT-guided PPO, SAC, and related algorithms. We made the source code available at [45]. B. Metrics We use the following metrics to evaluate the proposed approach and the existing counterparts: âąRatio of Completed Mission Tasks (R MC )is the ratio of cells completed to all assigned cells. âąEnergy Consumption (EC)calculates the total energy consumed by all drones, including HDs and MDs. âąNumber of Active, Connected Drones (N AC )is the number of non-compromised drones in the network. âąAccumulated Reward (G A orG D )by the attacker or de- fender are given in Sections I-B3 and I-C3, respectively. Due to space limitations, we show the experimental results usingN AC andG A /G D in the supplement document. C. Comparing Schemes For the comparative comparison of the proposed HT-DRL with the state-of-the-art (SOTA) non-HD counterparts in Sec- tion V-A, we consider two HD-based approaches, including HD-F using a fixed signal strength level and HD-HT-DRL using dynamic signal strength levels identified by HT-DRL. We consider two SOTA non-HD-based approaches, using an intrusion detection system (IDS) [46] and ContainerDrone (CD) [25] technique, along with a baseline mode with No Defense in handling DoS attacks. We also conduct the ablation study in our proposed HT-DRL by investigating the effect of using the fixed signal strength level, HT, and DRL in Section V-B. V. NUMERICALRESULTS ANDANALYSES Due to space constraints, additional ablation studies, sen- sitivity analyses, and supplementary metric evaluations are provided in the supplemental material. 12 FHTDRLHT-DRL 020406080100120140 Episode 0.5 0.6 0.7 0.8 0.9 1.0 MC (a)R MC under fixed attack. 020406080100120140 Episode 0.45 0.50 0.55 0.60 0.65 0.70 0.75 0.80 MC (b)R MC under HT attack. 020406080100120140 Episode 0.50 0.55 0.60 0.65 0.70 0.75 0.80 0.85 MC (c)R MC under DRL attack. 020406080100120140 Episode 0.50 0.55 0.60 0.65 0.70 0.75 0.80 0.85 MC (d)R MC under HT-DRL attack. Fig. 7. Performance analysis of honey drone defense strategies with respect to the ratio of completed mission tasks (R MC ), given an attack strategy. FHTDRLHT-DRL 020406080100120140 Episode 500 600 700 800 900 (a)ECunder fixed attack. 020406080100120140 Episode 500 550 600 650 700 750 (b)ECunder HT attack. 020406080100120140 Episode 500 550 600 650 700 750 800 850 (c)ECunder DRL attack. 020406080100120140 Episode 500 550 600 650 700 750 800 850 (d)ECunder HT-DRL attack. Fig. 8. Performance analysis of honey drone defense strategies with respect to energy consumption (EC), given an attack strategy. A. Comparative Performance Analysis with Non-HD Defenses We compare HD defenses (HD-F and HD-HT-DRL) with other defenses (i.e., IDS, CD), and no defense with a fixed signal under a given attack strategy. 1)Ratio of Mission Completion:Fig. 5 shows the per- formance of different defenses against DoS attacks in terms ofR MC . (1) HD-HT-DRL outperforms all. HD-F performs well only using HDs. This shows the effectiveness of defensive deception. The trends are well aligned withN AC results shown in Fig. 1 in Appendix A of the supplement document. (2) The effect of the HT attack is stronger than other attacks based on its lowerR MC . This shows that the strategic HT attack has a more adverse impact on mission performance. 2)Energy Consumption:Fig. 6 demonstrates the perfor- mance of different types of defenses against DoS attacks in terms ofEC. The key observations are: (1) The drone fleet protected by CD had the highestEC, meaning the CD can protect drones but lowers mission performance by stopping its operation. Such âzombie dronesâ hover on the target area for the maximum mission duration but do not perform mission operations. (2) As seen in Fig. 8, except Fig. 6a, HD-HT- DRL consumes less energy than other schemes for the same reason. More active, connected drones lead to higherR MC and naturally consume more energy. However, HD-HT-DRL consumes much less than CD and HD-F, despite its high per- formance inR MC . We also show the results using the metrics of accumulated reward and connected drones in Appendix A of the supplement document. B. Performance Analysis of Honey Drone Defenses In this section, we analyze the performance of HD-based defenses under various attacks in terms of the ratio of mission completion (R MC ) and energy consumption (EC). 1)Ratio of Mission Completion:Fig. 7 compares the performance of HD defenses, in terms ofR MC , given an attack strategy. The key observations are: (1) Overall we ob- serve the outperformance of HT-DRL over other HD defenses because the DRL agent in HT-DRL can explore solutions from its local optima identified by hypergame-theoretic, strategic decision-making without random exploration. We explain this process in Fig. 12 of the Appendix, demonstrating the agent using hypergame theory and later using RL to explore optimal solutions further. HT-DRL also leverages the benefit of using A2C as a DRL algorithm in exploring global optimal solutions adaptable to complex, non-stationary environments. Fig. 13 in the Appendix shows the distinct probability distributions in selecting strategies for the trained HT-DRL and DRL models. We can consistently observe the similar trends in Fig. 7 and sensitivity analyses shown in Fig. 7 of the Appendix. 2)Energy Consumption:Fig. 8 compares the perfor- mance of diverse HD-based defenses, inEC, given an attack strategy. We observe: (1) Except Fig. 8a, HT-DRL shows acceptable energy consumption compared to its highestR MC . Energy consumption relates to how many drones are active in the given network. Hence, more active, connected drones (N AC ; see Appendix B of the supplement document ) are aligned with higherR MC , resulting in higherEC. However, despite high performance inN AC andR MC , HT-DRLâs energy consumption is relatively low, meaning that HT-DRL considers energy efficiency in selecting its best strategy. (2) In Fig. 8a, the highestECis HT-DRL, showing much higher R MC (i.e., over 95%) under fixed attack than under other attacks in Figs. 8bâ8d. 13 VI. RELATEDWORK A. Defenses Against DoS Attacks in UAVs Chen et al. [25] discussed a defense mechanism to deal with DoS attacks in real-time UAV systems using containers, calledContainerDrone(CD), to prevent excessive resource usage for protecting critical resources (CPU, memory, and communication channel) from DoS attacks. Hence, if a certain level of resource utilization is reached, it stops providing ser- vices. Sedjelmaci et al. [47] proposed a hierarchical detection and response system to enhance security against DoS attacks, including attacks on the GPS module, in UAV networks. Ouiazzane et al. [24] proposed an IDS that utilizes the decision tree model to protect the multi-agent UAV system from DoS attacks. Gudla et al. [48] proposed a moving target defense (MTD) technique to protect the Parrot AR drone from DoS attacks by shuffling communication channels. Feng et al. [49] proposed a reinforcement learning approach that uses a multi-objective reward function to adaptively defend against application-layer DDoS attacks by balancing between mini- mizing false positives during low system load and maximizing attack mitigation during high load. Li et al. [50] proposed a transfer double deep Q-network-based DDoS detection method for the Internet of Vehicles that leverages traffic flow similarity between base stations to speed up DRL training for newly added base stations. Limitations: Although these works offer valuable defenses against DoS attacks in UAV systems, they may be limited when facing intelligent attackers empowered by machine learning or deep learning. For instance, the Moving Target Defense (MTD) technique [48] is susceptible to pattern recog- nition algorithms, potentially allowing attackers to anticipate communication channel shuffling. While Feng et al. [49] showed promising DDoS defense using reinforcement learn- ing, their work was limited to traditional server environments. Further, it did not consider the unique challenges of UAV systems, such as energy constraints, mobility, and dynamic network topology changes. While Li et al. [50] addressed the long training time issue of DRL through transfer learning, it still relies on collecting historical traffic data from existing base stations to calculate similarities, which may not be feasible in mission-critical UAV systems where immediate defense is needed and historical data collection opportunities are limited. Additionally, their approach focuses specifically on DDoS detection rather than comprehensive defensive de- ception strategies. B. Game Theory-based Defensive Deception Xiao et al. [51] examined the framing effect of an ad- vanced persistent threat (APT) attacker on its detection. They employedCumulative Prospect Theoryto consider playersâ subjective, uncertainty-aware decision-making for obtaining Nash Equilibrium (NE) solutions. Basak et al. [52] deployed honeypots using a multistage Stackelberg game and evaluated the accuracy of detecting the attackers collected by the honey- pots. P Ì Ä±bil et al. [53] used honeypots as a defensive deception technique to lure an attacker aiming to choose real servers. They formulated this as a zero-sum game with imperfect and incomplete information. Recently, Hypergame Theory (HT) has been used to develop defensive deception techniques. Cho et al. [54] developed a cyber deception hypergame (CDHG) using HT to examine the best strategy selection by an attacker and defender by developing the Stochastic Petri Nets models with the CDHG. Wan et al. [55] used HT to model the attack-defense inter- actions where an attacker and defender identify their best strategy using hypergame expected utility (HEU). Further, they investigated the effect of multiple attacks when the defender can choose a set of defense strategies based on HT [37]. Anwar et al. [56] considered HT to develop intelligent honey- pot strategies using either low-interaction or high-interaction honeypots for efficient and effective defensive deception. Limitations:However, the existing game-theoretic ap- proaches [51â53] have not considered the attackerâs and defenderâs perceived uncertainty in estimating their utilities for selecting the best strategies. In addition, the works using hypergame theory (HT) [37, 54â56] has mainly applied HT in static or low-mobility settings, such as enterprise or mobile cloud networks, which do not have the concerns of limited resources such as energy or wireless bandwidths. Further, they have not integrated game theory with DRL which can enhance the solution quality. C. DRL-based Defensive Deception (D) DRL has also been used to develop defensive deception strategies. Li et al. [5] used DRL to generate an optimal D strategy. They formulated a utility function to model the underlying threats related to common vulnerabilities in the virtual machine to deal with reconnaissance attacks. Li et al. [6] proposed a DRL-based D strategy selection approach to generate optimal placement strategies for the decoys and de- ceptive routing. Huang et al. [7] examined the vulnerability of systems where RL is used to choose the best defense strategy among MTD, D, and assistive human security technologies that aim to mitigate human-related vulnerabilities. Similarly, Huang and Zhu [57] investigated RLâs vulnerability under malicious falsification of cost signals based on the relationship between the falsified cost and the Q-factors. Charpentier et al. [58] used DRL (e.g., Deep Q-Network or DQN) to identify the best defense strategy for four types of MTD and D strategies. Olowononi et al. [59] studied how to deploy Intelli- gent Reflective Surfaces (IRS) in UAVs to strengthen wireless communications for battlefield contexts. They developed data- driven power allocation in communication channels using RL for obfuscating the attack surface by luring jammers into designated channels to mitigate DoS attacks. Limitations: The existing DRL-based defensive deception methods [5â7, 57â59] often requires extensive training and data to demonstrate its outperformance. However, they did not try to reduce the training time as the long training time is not allowed in resource-constrained environments. D. DRL Integrated with Game Theory Zheng et al. [9] presented a game-theoretic interpretation of actor-critic (AC) algorithms, modeled as a two-player general- 14 sum Stackelberg game, where the actor is the leader, and the critic is the follower. Cao and Xie [10] proposed a game- theoretic inverse RL to learn the parameters of the dynamic system and individual cost function of multistage games from the demonstrated sequences of system states. Limitations: Both studies [9, 10] did not address the long convergence time to learn in DRL using the integrated approach of DRL and game theory. Furthermore, these existing works did not explore the applicability of their methods in domains like cybersecurity or UAV systems. In summary, existing studies on hypergame-based cyber defense, DRL-based UAV defense, and HD deployment share three key limitations. First, most DRL-based defenses treat the attacker as part of a stationary environment and do not model asymmetric perception or misperception between attacker and defender. Second, DRL-based defenders generally begin from uninformed random exploration, resulting in slow and unstable learning under highly non-stationary DoS/DDoS attack behaviors. Third, HD deployment strategies and energy- consumption models are frequently simplified or static, lim- iting their applicability to realistic, energy-constrained multi- UAV surveillance missions. Our proposed HT-DRL framework is designed primarily to address the first and second limitations. By constructing a hypergame-based threat model and computing HEU-induced action probabilities, we explicitly account for perception asymmetry between attacker and defender and leverage it to guide the defenderâs early action selection. This mechanism mitigates cold-start effects and reduces inefficient exploration in the presence of adaptive DoS/DDoS attackers. In addition, our extended HD deployment and energy-consumption model partially addresses the third limitation by incorporating het- erogeneous HD/MD roles and per-UAV energy costs. More detailed hardware-level energy modeling and broader resource constraints remain important directions for future work. VII. CONCLUSION, LIMITATIONS, & FUTUREWORK Denial-of-Service (DoS) attacks on drones have become increasingly prevalent, representing a significant threat to UAV systems in real-world scenarios. Our defensive deception strategy using honey drones presented a practical solution to countermeasure DoS attacks effectively and efficiently. We integrated hypergame theory with deep reinforcement learning to provide a solution for enhancing the resilience and adaptability of the proposed cyber deception using honey drones, which can be well-suited for complex environments and real-world applications. In particular, we developed a hybrid approach integrating hypergame theory (HT) with deep reinforcement learning (DRL), namely HT-DRL, to avoid a cold start problem in DRL process of converging to close-to- optimal solutions. The key idea is how to intelligently select optimal signal strengths to minimize security vulnerabilities to DoS attacks while maximizing mission performance with acceptable energy consumption. A. Key Findings Via extensive experimental analysis, we obtained the fol- lowingkey findings: (1) Honey drone (HD) approaches outperformed all considered counterparts, particularly when attack strategies are highly intelligent by using game theory or DRL or both. Our proposed approach, HT-DRL, showed outperformance in mission performance and energy consump- tion. (2) Except for the case under a fixed attack strategy (F), HT-DRL (HD-HT-DRL) showed lower or comparable energy consumption while achieving high mission performance. This demonstrated the energy efficiency of HT-DRL. (3) When the attack uses F, the HT-DRL showed the best mission performance, compared to when other attack strategies. This is well explained by the trends observed in the number of active, connected drones (shown in the supplement document), consuming more energy. B. Limitations The presented work has the followinglimitationswhich can be addressed in our future work. First, a key characteristic of hypergame theory is its capability to deal with an agentâs per- ceived uncertainty. Although the defense system can estimate its own uncertainty, the estimation of the attackerâs uncertainty is based on our assumption that the system knows when it is exposed to the attacker and for how long. However, this may not be known in practice. In real-world deployments, this limi- tation can be addressed through practical approaches based on observed attack patterns or past records from similar scenarios. Second, our proposed HT-DRL leverages hypergame expected utility to tackle the cold-start problem. However, it requires a small amount of prior knowledge to train the HT agent. Hence, when no prior knowledge is available, the effectiveness of HT- DRL is reduced. Lastly, we assumed a static mission scenario. However, in real-world operations, mission parameters such as target area boundaries and surveillance time requirements may change dynamically. This limitation can be addressed in future work by developing adaptive mechanisms to adjust drone parameters in real-time when mission parameters change. We acknowledge that a well-trained hypergame theory agent and a static mission scenario may not fully reflect the dynamic nature of real-world environments. Future work will explore adaptive techniques to address these challenges. While our simulation-based DRL approach shows promis- ing results, the derived parameters and strategies are not directly transferable to real-world scenarios due to the inherent complexities of real-world environments. Nevertheless, this work provides high-level guidelines for leveraging DRL and hypergame-theoretic DRL in cyber deception drone scenarios, offering valuable insights for future real-world applications. The proposed techniques have not been validated with physical drones. Such an experiment could provide a more comprehensive evaluation of the techniques, taking into ac- count the complexities and uncertainties that arise in real- world scenarios. C. Future Research Considering all these limitations in mind, we have a plan for future researchas follows: (1) addressing additional types of attacks; (2) consideringtransfer learningto mitigate the cold start in addition to HT with respect to the above methods; and 15 (3) developing more realistic measures of agentsâ changing un- certainty reflecting real-time scenarios. (4) Implementing and evaluating our deception strategies using actual UAV hardware and environments to assess performance. (5) Incorporating strategies specifically designed to mitigate data manipulation and spoofing attacks, leveraging the flexibility of our HT-DRL approach to enhance the overall security and robustness of mission-critical UAV operations. An important direction for future research is extending our framework to multi-attacker and multi-defender scenarios. This would require developing n-player hypergame formulations and corresponding multi- agent reinforcement learning algorithms. The current single attacker-defender model provides the foundational framework that can be extended through coalition game theory and distributed multi-agent reinforcement learning (MARL) ap- proaches. (6) conducting a systematic sensitivity analysis of the utility trade-off weights(Ï,ζ)and the uncertainty param- eters(λ A ,λ D )to quantify how their values influence both mission coverage (RMC) and energy consumption (EC) under diverse attacker behaviors. ACKNOWLEDGMENT This work is partly supported by the Army Research Office and Army Research Laboratory under Grant Contract Num- bers W911NF-20-2-0140, W911NF-19-2-0150, W911NF-17- 1-0370, and W911NF-23-2-0012 and NSF award 2107450. REFERENCES [1] H. Kurunathanet al., âMachine learning-aided operations and communications of unmanned aerial vehicles: A contemporary survey,âIEEE Communications Surveys & Tutorials, 2023. [2] M. Hooperet al., âSecuring commercial wifi-based UAVs from common security attacks,â inMILCOM 2016-2016 IEEE Military Communications Conference. Baltimore, MD, USA: IEEE, 2016, p. 1213â1218. [3] M. Zhuet al., âA survey of defensive deception: Ap- proaches using game theory and machine learning,âIEEE Communications Surveys & Tutorials, vol. 23, no. 4, p. 2460â2493, 2021. [4] J. Daubert, D. Boopalanet al., âHoneyDrone: A medium- interaction unmanned aerial vehicle honeypot,â inNOMS 2018-2018 IEEE/IFIP Network Operations and Manage- ment Symposium. IEEE, 2018, p. 1â6. [5] H. Li, Y. Guo, S. Huo, H. Huet al., âDefensive de- ception framework against reconnaissance attacks in the cloud with deep reinforcement learning,âScience China Information Sciences, vol. 65, no. 7, p. 170305, 2022. [6] H. Li, Y. Guo, P. Sun, Y. Wang, and S. Huo, âAn optimal defensive deception framework for the container-based cloud with deep reinforcement learning,âIET Informa- tion Security, vol. 16, no. 3, p. 178â192, 2022. [7] Y. Huang, L. Huang, and Q. Zhu, âReinforcement learn- ing for feedback-enabled cyber resilience,âAnnual Re- views in Control, 2022. [8] R. R. Vane, âHypergame theory for DTGT agents,â in American Association for Artificial Intelligence, 2000. [9] L. Zhenget al., âStackelberg actor-critic: Game-theoretic reinforcement learning algorithms,â inProceedings of the AAAI Conference on Artificial Intelligence, vol. 36, 2022, p. 9217â9224. [10] K. Cao and L. Xie, âGame-theoretic inverse reinforce- ment learning: A differential pontryaginâs maximum principle approach,âIEEE Transactions on Neural Net- works and Learning Systems, 2022. [11] Z. Wanet al., âDeception in drone surveillance missions: Strategic vs. learning approaches,â inProceedings of the Twenty-Fourth International Symposium on Theory, Al- gorithmic Foundations, and Protocol Design for Mobile Networks and Mobile Computing, 2023, accepted. [12] G.-H. Kimet al., âMulti-drone control and network self- recovery for flying ad hoc networks,â in2016 Eighth International Conference on Ubiquitous and Future Net- works. Vienna, Austria: IEEE, 2016, p. 148â150. [13] W. Khiatiet al., âAir surveillance planning approach for large areas,â in2019 Third International Conference on Intelligent Computing in Data Sciences (ICDS).Mar- rakech, Morocco: IEEE, 2019, p. 1â6. [14] A. Ahmed, A. Al-Dweik, Y. Iraqi, H. Mukhtar, M. Naeem, and E. Hossain, âHybrid automatic repeat request (harq) in wireless communications systems and standards: A contemporary survey,âIEEE Communica- tions Surveys & Tutorials, vol. 23, no. 4, p. 2711â2752, 2021. [15] P. Widhalm, U. Ritzinger, N. Pr Ì uggler, W. Pr Ì uggler, D. Strelnikova, G. Paulus, F. dâApolito, and F. Eicken, âCommunity drones: A concept study on shared drone services,âDrones, vol. 9, no. 2, p. 107, 2025. [16] C. Qin and E. Pournaras, âCoordination of drones at scale: Decentralized energy-aware swarm intelligence for spatio-temporal sensing,âTransportation Research Part C: Emerging Technologies, vol. 157, p. 104387, 2023. [17] K. Chevliet al., âBlue force tracking network modeling and simulation,â inMILCOM 2006-2006 Military Com- munications Conference. Washington, DC, USA: IEEE, 2006, p. 1â7. [18] B. Siddappaji and K. Akhilesh, âRole of cyber security in drone technology,âSmart Technologies: Scope and Applications, p. 169â178, 2020. [19] I.HimaxTechnologies,âHM01B0ultralow powerCIS,â2015,accessed:03-08-2022.[On- line]. Available: https://w.himax.com.tw/en/products/ cmos-image-sensor/always-on-vision-sensors/hm01b0/ [20] S. Liebergeld, M. Lange, and C. Mulliner, âNomadic honeypots: A novel concept for smartphone honey- pots,â inProc. Wâshop on Mobile Security Technologies (MoSTâ13), together with 34th IEEE Symp. on Security and Privacy, vol. 4. Germany: Citeseer, 2013, p. 1â4. [21] Bitcraze, âBitcrazeâs Crazyflie 2.x nano-quadrotor,â 2022,accessed:03-08-2022.[Online].Avail- able: $https://w.bitcraze.io/documentation/hardware/ crazyflie\ 2\1/crazyflie\2\1-datasheet.pdf$ [22] J. Panerati, H. Zheng, S. Zhou, J. Xuet al., âLearning to flyâa gym environment with pybullet physics for re- inforcement learning of multi-agent quadcopter control,â 16 2021 International Conference on Intelligent Robots and Systems (IROS), vol. 0, p. 7512â7519, 2021. [23] D. Palossi, N. Zimmermanet al., âFully onboard ai- powered human-drone pose estimation on ultralow-power autonomous flying nano-uavs,âIEEE Internet of Things Journal, vol. 9, no. 3, p. 1913â1929, 2021. [24] S. Ouiazzane, M. Addou, and F. Barramou, âA multiagent and machine learning based denial of service intrusion detection system for drone networks,âGeospatial Intelli- gence: Applications and Future Trends, p. 51â65, 2022. [25] J. Chen, Z. Fenget al., âA container-based DoS attack- resilient control framework for real-time UAV systems,â in2019 Design, Automation & Test in Europe Conference & Exhibition (DATE). IEEE, 2019, p. 1222â1227. [26] B. Ly and R. Ly, âCybersecurity in unmanned aerial vehicles (UAVs),âJournal of Cyber Security Technology, vol. 5, no. 2, p. 120â137, 2021. [27] Y. Mekdadet al., âA survey on security and privacy issues of UAVs,âComputer Networks, p. 109626, 2023. [28] P. Mellet al., âCommon vulnerability scoring system,â IEEE Security & Privacy, vol. 4, no. 6, p. 85â89, 2006. [29] CVEdetails, âCVE-2021-39804,â 2021. [Online]. Avail- able: https://w.cvedetails.com/cve/CVE-2021-39804/ [30] M. K. Baazaoui, I. Ketata, A. Fakhfakh, and F. Derbel, âModeling of packet error rate distribution based on received signal strength indications in omnet++ for wake- up receivers,âSensors, vol. 23, no. 5, p. 2394, 2023. [31] A. Gachhadar, R. K. Maharjanet al., âPower optimiza- tion in multi-tier heterogeneous networks using genetic algorithm,âElectronics, vol. 12, no. 8, p. 1795, 2023. [32] A. Mahmood, M. A. Hossain, and M. Gidlund, âCross- layer optimization of wireless links under reliability and energy constraints,â in2018 IEEE Wireless Communica- tions and Networking Conference (WCNC). IEEE, 2018, p. 1â6. [33] W. contributors, âStochastic geometry models of wireless networks.â [34] E. Bj Ì ornson and E. G. Larsson, âHow energy-efficient can a wireless communication system become?â in2018 52nd Asilomar conference on signals, systems, and com- puters. IEEE, 2018, p. 1252â1256. [35] F. Mahmood, E. Perrins, and L. Liu, âEnergy-efficient wireless communications: From energy modeling to per- formance evaluation,âIEEE Transactions on Vehicular Technology, vol. 68, no. 8, p. 7643â7654, 2019. [36] T. Kanazawa, T. Ushio, and T. Yamasaki, âReplicator dy- namics of evolutionary hypergames,âIEEE Transactions on Systems, Man, and Cybernetics-Part A: Systems and Humans, vol. 37, no. 1, p. 132â138, 2006. [37] Z. Wan, J.-H. Cho, M. Zhu, A. H. Anwar, C. Kamhoua et al., âResisting multiple advanced persistent threats via hypergame-theoretic defensive deception,âIEEE Trans- actions on Network and Service Management, 2023. [38] W. Xue, W. Qiu, X. Hua, and K. Yu, âImproved Wi- Fi RSSI measurement for indoor localization,âIEEE Sensors Journal, vol. 17, no. 7, p. 2224â2230, 2017. [39] M. Sauter,From GSM to LTE: An introduction to mobile networks and mobile broadband.India: John Wiley & Sons, 2010. [40] S. Fugate and K. Ferguson-Walter, âArtificial intelligence and game theory models for defending critical networks with cyber deception,âAI Magazine, vol. 40, no. 1, p. 49â62, 2019. [41] L. Huang and Q. Zhu, âDynamic bayesian games for ad- versarial and defensive cyber deception,â inAutonomous Cyber Deception: Reasoning, Adaptive Planning, and Evaluation of HoneyThings.Springer, 2019, p. 75â 97. [42] L. Perron and V. Furnon, âOr-tools,â Google, 2022, version v9.5. [Online]. Available: https://developers. google.com/optimization/ [43] C. Voudouris and E. Tsang, âPartial constraint satisfac- tion problems and guided local search,âProc., Practical Application of Constraint Technology (PACTâ96), Lon- don, vol. 0, p. 337â356, 1996. [44] T. Schaul, J. Quan, I. Antonoglou, and D. Sil- ver, âPrioritized experience replay,âarXiv preprint arXiv:1511.05952, 2015. [45] Z. Wanet al., âSource code for drone simulation.â [Online]. Available: github.com/Wan-ZL/ARO-Foureye [46] J.-P. Condomines, R. Zhang, and N. Larrieu, âNetwork intrusion detection system for UAV ad-hoc communica- tion: From methodology design to real test validation,â Ad Hoc Networks, vol. 90, p. 101759, 2019. [47] H. Sedjelmaci, S. M. Senouci, and N. Ansari, âA hi- erarchical detection and response system to enhance security against lethal cyberattacks in UAV networks,â IEEE Transactions on Systems, Man, and Cybernetics: Systems, vol. 48, no. 9, p. 1594â1606, 2017. [48] C. Gudla, M. S. Ranaet al., âDefense techniques against cyber attacks on unmanned aerial vehicles,â inProceed- ings of the International Conference on Embedded Sys- tems, Cyber-physical Systems, and Applications (ESCS). The Steering Committee of The World Congress in Computer Science, Computer, 2018, p. 110â116. [49] Y. Feng, J. Li, and T. Nguyen, âApplication-layer ddos defense with reinforcement learning,â in2020 IEEE/ACM 28th International Symposium on Quality of Service (IWQoS). IEEE, 2020, p. 1â10. [50] Z. Li, Y. Kong, and C. Jiang, âA transfer double deep q network based ddos detection method for internet of vehicles,âIEEE Transactions on Vehicular Technology, vol. 72, no. 4, p. 5317â5331, 2023. [51] L. Xiao, D. Xu, N. B. Mandayam, and H. V. Poor, âAttacker-centric view of a detection game against ad- vanced persistent threats,âIEEE transactions on mobile computing, vol. 17, no. 11, p. 2512â2523, 2018. [52] A. Basak, C. Kamhoua, S. Venkatesanet al., âIdentifying stealthy attackers in a game theoretic framework using deception,â inInternational Conference on Decision and Game Theory for Security. Springer, 2019, p. 21â32. [53] R. P Ì Ä±bilet al., âGame theoretic model of strategic hon- eypot selection in computer networks,â inProceedings of the International Conference on Decision and Game Theory for Security. Springer, 2012, p. 201â220. [54] J. Cho, M. Zhuet al., âModeling and analysis of decep- 17 tion games based on hypergame theory,â inAutonomous Cyber Deception. Springer, 2019, p. 49â74. [55] Z. Wan, J.-H. Cho, M. Zhuet al., âFoureye: Defen- sive deception against advanced persistent threats via hypergame theory,âIEEE Transactions on Network and Service Management, vol. 19, no. 1, p. 112â129, 2021. [56] A. H. Anwar, M. Zhuet al., âHoneypot-based cyber de- ception against malicious reconnaissance via hypergame theory,â inGLOBECOM 2022-2022 Global Communica- tions Conference. IEEE, 2022, p. 3393â3398. [57] Y. Huang and Q. Zhu, âDeceptive reinforcement learning under adversarial manipulations on cost signals,â inDe- cision and Game Theory for Security.Springer, 2019, p. 217â237. [58] A. Charpentier, N. Boulahia Cuppenset al., âDeep re- inforcement learning-based defense strategy selection,â inProceedings of the 17th International Conference on Availability, Reliability and Security, 2022, p. 1â11. [59] F. O. Olowononiet al., âDeep reinforcement learning for deception in IRS-assisted UAV communications,â inMILCOM 2022-2022 IEEE Military Communications Conference (MILCOM). IEEE, 2022, p. 763â768. Zelin Wanis an AI Research Scientist at Bobyard in San Francisco, CA. He earned his Ph.D. in Computer Science from Virginia Tech in 2025, following an M.S. in Computer Science from Virginia Tech in 2021 and a B.S. in Computer Science from the University of Arizona in 2019. His research interests include game theory, deep learning, cybersecurity, and network science. He has been nominated for the Joseph Frank Hunkler Memorial Scholarship, received the College of Engineering Graduate Stu- dent Publication Fellowship, and was selected for the departmentâs Best Dissertation Award for 2024â2025 CS Ph.D. students. He is a student member of the ACM. Jin-Hee Cho(Mâ09; SMâ14) is an Associate Profes- sor in the Department of Computer Science at Vir- ginia Tech, where she has served since August 2018, and the director of the Trustworthy Cyberspace Lab. Prior to joining Virginia Tech, she was a computer scientist at the U.S. DEVCOM Army Research Lab- oratory (US DEVCOM ARL) in Adelphi, MD, be- ginning in 2009. Dr. Cho has published extensively in leading journals and conferences in the areas of cybersecurity, decision-making under uncertainty, and network science. She has received multiple best paper awards, including IEEE TrustCom 2009, BRIMS 2013, IEEE GLOBECOM 2017, ARLâs Publication Award in 2017, and IEEE CogSima 2018. She was awarded the 2015 IEEE Communications Society William R. Bennett Prize in Communications Networking and the 2023 IEEE ComSoc Network Operations & Management (CNOM) Test of Time Paper Award. In 2013, Dr. Cho received the Presidential Early Career Award for Scientists and Engineers (PECASE), the highest honor bestowed by the U.S. government on early-career researchers. She was also recognized with the 2022 Faculty Fellow Award from the College of Engineering (CoE) and named a 2025â2026 CoE Deanâs Fellow at Virginia Tech. In addition, she was awarded the Stephen and Cherye Tyndall Moore Junior Faculty Fellowship, which provides endowed research support. Dr. Cho earned her Ph.D. in Computer Science from Virginia Tech in 2008. She currently serves as an associate editor for the IEEE Transactions on Network and Service Management, IEEE Transactions on Services Computing, and The Computer Journal. She is a Senior Member of IEEE and a member of ACM. Mu Zhuis a postdoctoral researcher at the Com- puter Network Information Center of the Chinese Academy of Sciences as of 2024. His research fo- cuses on cybersecurity and defensive deception, with a particular interest in enhancing these fields through game theory and machine learning. He earned his Ph.D. in Computer Science from North Carolina State University in 2024, his M.S. in Computer Engineering from the University of Delaware in 2017, and his B.S. in Electronic Engineering from Zhengzhou University, Henan, China, in 2014. Ahmed H. Anwar(Sâ09; Mâ19) is currently a postdoctoral research scientist in the U.S. Army Research Lab in Adelphi, MD since 2019. His re- search interests include Network Security, Algorith- mic Game Theory and Machine Learning. Dr. Anwar earned his PhD degree in Electrical Engineering from the University of Central Florida in 2019. Before that he worked as a research assistant in Nile University, Egypt, and Qatar University be- tween 2011 and 2013. He received his BSc. degree (with highest honors) in electrical engineering from Alexandria University, Alexandria, Egypt, in 2011 and the MSc. degree in wireless Information Technology from Nile University, Egypt, in 2013. Charles A. Kamhouais a Senior Electronics Engi- neer at the Network Security Branch of the ARL in Adelphi, MD, where he conducts and directs basic research on applying game theory to cybersecurity. Before joining ARL, he spent six years as a re- searcher at AFRL, Rome, NY, and held visiting po- sitions at the University of Oxford and Harvard Uni- versity. He has co-authored over 200 peer-reviewed papers, earning five best paper awards, and co-edited four Wiley-IEEE Press books: Game Theory and Machine Learning for Cyber Security, Modeling and Design of Secure Internet of Things, Blockchain for Distributed System Security, and Assured Cloud Computing. His scholarship and leadership have been recognized with awards such as the 2020 Sigma Xi Young Investigator Award and the 2019 US Army Civilian Service Commendation Medal. He earned a B.S. in electronics from the University of Douala (ENSET), Cameroon, in 1999, an M.S. in Telecommunication and Networking from Florida International University (FIU) in 2008, and a Ph.D. in Electrical Engineering from FIU in 2011. He is a senior member of ACM and IEEE. Munindar P. Singh(Fâ09) is the SAS Institute Dis- tinguished Professor and an Alumni Distinguished Graduate Professor in the Department of Computer Science at North Carolina State University. Munin- darâs research interests include AI and multiagent systems and their applications, including in cyber- security. Munindar is a Fellow of AAAI (Associa- tion for the Advancement of Artificial Intelligence), AAAS (American Association for the Advancement of Science), ACM (Association for Computing Ma- chinery), and IEEE (Institute of Electrical and Elec- tronics Engineers), and was elected a foreign member of Academia Europaea. He has won the ACM/SIGAI Autonomous Agents Research Award, the IEEE TCSVC Research Innovation Award, and the IFAAMAS Influential Paper Award. He won NC State Universityâs Outstanding Graduate Faculty Mentor Award as well as the Outstanding Research Achievement Award (twice). He was selected as an Alumni Distinguished Graduate Professor and elected to NCSUâs Research Leadership Academy. He also won NC Stateâs Outstanding Faculty Mentor Award. Munindar was the editor-in-chief of the ACM Transactions on Internet Technology (2012â2018) and IEEE Internet Computing (1999â2002). Munindar served on the founding board of directors of IFAAMAS, the International Foundation for Autonomous Agents and MultiAgent Systems. Thirty-one students have received PhD degrees and forty-one students MS degrees under Munindarâs direction. 1 Supplement Document: Cyber Deception for Mission Surveillance via Hypergame-Theoretic Deep Reinforcement Learning Zelin Wan, Jin-Hee Cho,Senior Member, IEEE, Mu Zhu, Ahmed H. Anwar, and Charles Kamhoua,Senior Member, IEEE, Munindar P. Singh,IEEE Fellow APPENDIXA HT-DRL TRAININGALGORITHM This section provides the formal algorithmic presentation of the HT-DRL training procedure referenced in the main paper. Algorithm 1HT-DRL Training Procedure 1:Input:Environment, HT parameters, RL hyperparameters 2:Output:Trained HT-DRL agent 3:Initialize HT agent with utility functions 4:Generate action probability distribution (APD) using HEU 5:Initialize A2C agent and apply APD weights to output layer 6:forepisode = 1 to max episodesdo 7:Interact with environment using filtered policy 8:Update A2C networks with gradients 9:end for 10:returnTrained HT-DRL agent APPENDIXB COMPARATIVEPERFORMANCEANALYSIS WITHOTHER DEFENSES In this section, we explain the performance analyses of the defense mechanisms considered in this work, particularly in terms of honey drone (HD)-based vs. non-HD based. Specifically, we will compared the performance of HD-F, IDS, CD, HD-HT-DRL, and the No-Defense schemes (see Section IV.A of the main paper, discussing Comparing Schemes) in terms of the number of active, connected drones (N AC ) and accumulated rewards by the attacker and defender (G A and G D ), which are detailed in Section IV.B of the main paper. This research was partly sponsored by the Army Research Laboratory and was accomplished under Cooperative Agreement Number W911NF-19-2-0150 and W911NF-23-2-0012. In addition, this research is also partly supported by the Army Research Office under Grant Contract Numbers W911NF-20-2- 0140 and W911NF-17-1-0370. The views and conclusions contained in this document are those of the authors and should not be interpreted as representing the official policies, either expressed or implied, of the Army Research Laboratory or the U.S. Government. The U.S. Government is authorized to reproduce and distribute reprints for Government purposes notwithstanding any copyright notation herein (Corresponding author: Zelin Wan). Zelin Wan and Jin-Hee Cho are with the Department of Computer Science, Virginia Tech, Falls Church, VA, USA. Email:zelin, jicho@vt.edu. Mu Zhu and Munindar P. Singh are with the Department of Computer Science, North Carolina State University, Raleigh, NC 27695. Email:mzhu5, mpsingh@ncsu.edu. Ahmed H. Anwar and Charles A. Kamhoua are with the US Army Re- search Laboratory, Adelphi, MD, USA. Email: a.h.anwar@knights.ucf.edu; charles.a.kamhoua.civ@mail.mil. 1)Number of Active, Connected Drones:Fig. 1 com- pares the performance of HD-based defenses, including F, HT, DRL, and HT-DRL, in terms ofN AC , given an attack strategy. Our key observations are as follows: (1) HD-based defense techniques, specifically HD-F and HD-HT-DRL, outperform other defense techniques. This implies that honey drones are effective at luring attacker while simultaneously serving as a relay to maintain the connectivity of the drone fleet. (2) HD-HT-DRL exhibits the highest performance across different schemes, demonstrating that the synergistic, hybrid approach combining game theory and DRL can contribute to fast and autonomous learning for optimal solutions. This also can best utilize the merit of the honey drones and subsequently increasesN AC , which leads to higher mission performance as shown in Fig. 5 of the main section in this paper. 2)Accumulated Rewards:Figs. 2 and 3 compare the performance in terms ofG A andG D , respectively, under a given attack strategy. The figures show the following results: (1) No-Defense strategy results in the highestG A and lowest G D across all strategies because the system does not have any defense against DoS attacks. (2) In Figs. 2a and 3a, we observed the CD shows the highestG A and lowestG D among all. This is because CD prevents mission drones from being compromised but also disables them from performing the mission due to its protection mechanism. These drones can be analogous to âzombiesâ â they cannot be compromised but are also unproductive. As a result, the defender cannot complete new tasks, resulting in a lowG D . From the attackerâs perspective, as many cells remain unscanned, the attacker enjoys a high immediate reward. Since those âzombieâ drones are considered active, the mission does not end due to a lack of available drones. A prolonged mission duration thus allows the attacker to accumulate a higher reward,G A . APPENDIXC PERFORMANCEANALYSIS OFHONEYDRONE-BASED DEFENSES This appendix discusses the performance of HD-based de- fenses, including F, HT, DRL, and HT-DRL, under the number of active, connected drones (N AC ) and accumulated rewards by the attacker (G A ) or defender (G D ). 1)Number of Active, Connected Drones:Fig. 4 show- cases the comparative performance in terms ofN AC . We found: (1) The results in Fig. 4 are well aligned with the 2 HD-FIDSCDHD-HT-DRLNo-Defense 020406080100120140 Episode 2.5 5.0 7.5 10.0 12.5 15.0 17.5 AC (a)N AC under fixed attack. 020406080100120140 Episode 2 4 6 8 10 12 AC (b)N AC under HT attack. 020406080100120140 Episode 4 6 8 10 12 14 AC (c)N AC under DRL attack. 020406080100120140 Episode 4 6 8 10 12 14 AC (d)N AC under HT-DRL attack. Fig. 1. Performance analysis of HD-HT-DRL, IDS, CD, and fixed defense, given an attack strategy, with respect to the number of active, connected drones (N AC ). HD-FIDSCDHD-HT-DRLNo-Defense 020406080100120140 Episode 1600 1800 2000 2200 2400 2600 2800 A (a)G A under fixed attack. 020406080100120140 Episode 2200 2400 2600 2800 3000 3200 A (b)G A under HT attack. 020406080100120140 Episode 2000 2200 2400 2600 2800 A (c)G A under DRL attack. 020406080100120140 Episode 2000 2200 2400 2600 2800 A (d)G A under HT-DRL attack. Fig. 2. Performance analysis of HD-HT-DRL, IDS, CD, and fixed defense, given an attack strategy, with respect to the attackerâs accumulated reward (G A ). HD-FIDSCDHD-HT-DRLNo-Defense 020406080100120140 Episode 7.5 10.0 12.5 15.0 17.5 20.0 22.5 25.0 D (a)G D under fixed attack. 020406080100120140 Episode 4 6 8 10 12 14 16 18 20 D (b)G D under HT attack. 020406080100120140 Episode 10 12 14 16 18 20 22 D (c)G D under DRL attack. 020406080100120140 Episode 12 14 16 18 20 22 D (d)G D under HT-DRL attack. Fig. 3. Performance analysis of HD-HT-DRL, IDS, CD, and fixed defense, given an attack strategy, with respect to the defenderâs accumulated reward (G D ). results in Fig. 7 in the main paper, where the high mission performance is achieved by more active, connected nodes, showing the performance order of HT-DRLâ„DRLâ„HT â„F, given a sufficiently large number of episodes (e.g., after >60). (2) HT-DRL exhibits superior performance over the other HD-based defenses with the same reasons discussed with Fig. 7 in Section V-B1 of the main paper. 2)Accumulated Reward:Figs. 5 and 6 present a compar- ative study of HD-based defenses, specifically F, HT, DRL, and HT-DRL, in terms ofG A andG D , respectively, under a given attack strategy. From the results shown in those figures, we observed: (1) The performance order is reversed when comparing Figs. 5 and 6. This is mainly because the immediate reward for the attacker corresponds to the number of uncompleted mission tasks in roundt, whereas the defenderâs immediate reward relates to the number of completed mission tasks in the same round. This reward design effectively creates a zero-sum-like game scenario, resulting in the observed reverse order of rewards. (2) A special case is observed where DRL has a lowG A in Fig. 5b and also lowG D in Fig. 6b. This may be because HT-based attacker outperforms and compromises most drones, leading to early termination of the mission. Since bothG A andG D represent accumulated rewards, a reduced mission duration may lead to a lower performance under the accumulated rewards. APPENDIXD SENSITIVITYANALYSIS: EFFECT OFVARYINGATTACK BUDGET(ζ) This section investigates the impact of varying the attack budget (ζ) on system performance based on the metrics in Section IV-B. The results are presented based on the ones that converged at the 140th episode. 3 FHTDRLHT-DRL 020406080100120140 Episode 10 12 14 16 18 AC (a)N AC under fixed attack. 020406080100120140 Episode 9.5 10.0 10.5 11.0 11.5 12.0 12.5 13.0 AC (b)N AC under HT attack. 020406080100120140 Episode 10 11 12 13 14 AC (c)N AC under DRL attack. 020406080100120140 Episode 10 11 12 13 14 15 AC (d)N AC under HT-DRL attack. Fig. 4. Performance analysis of honey done-based defenses with respect to the number of active, connected drones (N AC ), under a given attack. FHTDRLHT-DRL 020406080100120140 Episode 1600 1800 2000 2200 2400 2600 2800 3000 A (a)G A under fixed attack. 020406080100120140 Episode 2200 2400 2600 2800 A (b)G A under HT attack. 020406080100120140 Episode 2000 2200 2400 2600 2800 A (c)G A under DRL attack. 020406080100120140 Episode 2000 2200 2400 2600 2800 A (d)G A under HT-DRL attack. Fig. 5. Performance analysis of honey done-based defenses with respect to the attackerâs accumulated reward (G A ), under a given attack. FHTDRLHT-DRL 020406080100120140 Episode 12 14 16 18 20 22 24 D (a)G D under fixed attack. 020406080100120140 Episode 12 14 16 18 20 D (b)G D under HT attack. 020406080100120140 Episode 14 16 18 20 22 D (c)G D under DRL attack. 020406080100120140 Episode 14 16 18 20 22 D (d)G D under HT-DRL attack. Fig. 6. Performance analysis of honey done-based defenses with respect to the defenderâs accumulated reward (G D ), under a given attack. 1)Ratio of Mission Completion:Fig. 7 represents how a different level of the attack budget (ζ, ranged from 1 to 9) impacts the ratio of mission completion (R MC ). From the results, our findings are: (1) Overall, as the attack budget increases,R MC decreases because a higher attack budget empowers the attacker to launch more attacks in each step. Consequently, fewer mission drones are available, resulting in a diminishedR MC . (2) However, we observed that after ζ= 5, the defense performance is mostly flat and can maintain its high performance. This is mainly because even if the attacker has higher budget, it may not find sufficient drones that have sufficient signal strength to perform attack. (3) As we observed in Section V-B1, HT attack is more effect in lowering down the performance of RL-based defenses due to the merit of strategic decision making by hypergame theory. 2)Energy Consumption:Fig. 8 shows the effect of varying the attack budget (ζ) on energy consumption (EC). We found the following from the results: (1)ECdecreases with higherζbecause more compromised (crashed) nodes can lead to fewer active nodes, resulting in low mission performance. (2) However, HT-DRL consumes relatively less compared to its outperformance in mission performance because of its intelligent energy-aware strategy selection. 3)Number of Active, Connected Drones:Fig. 9 demon- strates how a different level of the attack budget (ζ) affects the number of active, connected drones (N AC ). As observed in Fig. 7, in Fig. 9, HT-DRL outperforms among all showing the highestN AC . The same discussions apply as in Appendix D-1 with the result of the ratio of mission completion. 4)Accumulated Reward:Figs. 10 and 11 show the effect of varying the attack budget (ζ) on the attackerâs reward (G A ) and the defenderâs reward (G D ). Interestingly, we observed a reversed ordering betweenG A andG D . Specifically, Fig. 10 shows the performance order of F>HT>DRL>HT-DRL. On the other hand, Fig. 11 demonstrates the performance order of HT-DRL>DRL>HT>F. This is because the 4 attackerâs immediate reward is the number of mission tasks not completed in roundt(see more details in Section I- B3), while the defenderâs immediate reward is the number of mission tasks completed in roundt(see the details in Section I-C3). As a result, this reward design represents close to a zero-sum game based on the rationale that one playerâs gain is the opponent playerâs loss and vice-versa. APPENDIXE DISTRIBUTIONALANALYSIS OFEACHSTRATEGY This section delves into the evolution of the defenderâs strategy distribution across different schemes. Fig. 12 illustrates the probability distribution of taking defense strategies after the training with the first two episodes. These probability distributions of taking defense strategies explain: (1) DRL-based defense mainly explores based on the random selection, which shows a uniform probability distribu- tion of the strategy selection. This ensures a consistent output distribution of neural networks (NNs) for all schemes. (2) HT and HT-DRL exhibit similar distribution patterns caused by the N layer added to the original NNs before the output layer of HT-DRL. By initializing this particular layer using the action distribution from the HT agent, HT-DRLâs performance is close to that of HT. It can start its exploration to achieve global optimal solutions from the local optima identified by HT. Fig. 13 presents the probability distributions of taking de- fense strategies after the models are fully trained. The overall key findings are: (1) HT-DRL shows a skewed distribution with clearly dominant strategies compared to the distributions of the strategies taken in HT and DRL. We can expect lower entropies under HT-DRL compared to the ones under HT or DRL. (2) A notable trend emerges that a sharper distribution often correlates with a higherR MC , as shown in Fig. 7 of the main paper. This trend suggests that those frequently employed strategies likely strike a balance between minimizing detection probability by the attacker and maximizing the connectivity between mission drones to effectively execute the given mis- sion. 5 FHTDRLHT-DRL 13579 0.6 0.7 0.8 0.9 1.0 MC (a)R MC under fixed attack. 13579 0.60 0.65 0.70 0.75 0.80 0.85 MC (b)R MC under HT attack. 13579 0.65 0.70 0.75 0.80 0.85 0.90 MC (c)R MC under DRL attack. 13579 0.70 0.75 0.80 0.85 0.90 MC (d)R MC under HT-DRL attack. Fig. 7. Effect of varying the attack budget (ζ) on the ratio of completed mission tasks (R MC ). FHTDRLHT-DRL 13579 550 600 650 700 750 800 850 900 950 (a)ECunder fixed attack. 13579 500 550 600 650 700 750 800 850 (b)ECunder HT attack. 13579 550 600 650 700 750 800 850 900 (c)ECunder DRL attack. 13579 550 600 650 700 750 800 850 900 (d)ECunder HT-DRL attack. Fig. 8. Effect of varying the attack budget (ζ) on energy consumption (EC). FHTDRLHT-DRL 13579 12 14 16 18 AC (a)N AC under fixed attack. 13579 11.5 12.0 12.5 13.0 13.5 14.0 14.5 AC (b)N AC under HT attack. 13579 12 13 14 15 16 AC (c)N AC under DRL attack. 13579 12.5 13.0 13.5 14.0 14.5 15.0 15.5 16.0 AC (d)N AC under HT-DRL attack. Fig. 9. Effect of varying the attack budget (ζ) on the number of active, connected drones (N AC ). FHTDRLHT-DRL 13579 1700 1800 1900 2000 2100 2200 2300 2400 A (a)G A under fixed attack. 13579 2100 2200 2300 2400 2500 A (b)G A under HT attack. 13579 2000 2100 2200 2300 2400 A (c)G A under DRL attack. 13579 2000 2100 2200 2300 2400 A (d)G A under HT-DRL attack. Fig. 10. Effect of varying the attack budget (ζ) on the attackerâs accumulated reward (G A ). 6 FHTDRLHT-DRL 13579 14 16 18 20 22 24 D (a)G D under fixed attack. 13579 15 16 17 18 19 20 21 22 D (b)G D under HT attack. 13579 16 17 18 19 20 21 22 D (c)G D under DRL attack. 13579 17 18 19 20 21 22 23 D (d)G D under HT-DRL attack. Fig. 11. Effect of varying the attack budget (ζ) on the defenderâs accumulated reward (G D ). HTDRLHT-DRL 12345678910 Defender Strategy 0.000 0.025 0.050 0.075 0.100 0.125 0.150 0.175 0.200 Defender Strategy Frequency (a) Under fixed attack. 12345678910 Defender Strategy 0.000 0.025 0.050 0.075 0.100 0.125 0.150 0.175 0.200 Defender Strategy Frequency (b) Under HT attack. 12345678910 Defender Strategy 0.000 0.025 0.050 0.075 0.100 0.125 0.150 0.175 0.200 Defender Strategy Frequency (c) Under DRL attack. 12345678910 Defender Strategy 0.000 0.025 0.050 0.075 0.100 0.125 0.150 0.175 0.200 Defender Strategy Frequency (d) Under HT-DRL attack. Fig. 12. Probability distributions of the strategy selection in HT, DRL, and HT-DRL after the training with the two episodes. HTDRLHT-DRL 12345678910 Defender Strategy 0.00 0.05 0.10 0.15 0.20 0.25 0.30 0.35 0.40 Defender Strategy Frequency (a) Under fixed attack. 12345678910 Defender Strategy 0.00 0.05 0.10 0.15 0.20 0.25 0.30 0.35 0.40 Defender Strategy Frequency (b) Under HT attack. 12345678910 Defender Strategy 0.00 0.05 0.10 0.15 0.20 0.25 0.30 0.35 Defender Strategy Frequency (c) Under DRL attack. 12345678910 Defender Strategy 0.00 0.05 0.10 0.15 0.20 0.25 0.30 0.35 Defender Strategy Frequency (d) Under HT-DRL attack. Fig. 13. Probability distributions of the strategy selection in HT, DRL, and HT-DRL after the models are fully trained.