Paper deep dive
Hetero-Net: An Energy-Efficient Resource Allocation and 3D Placement in Heterogeneous LoRa Networks via Multi-Agent Optimization
Abdullahi Isa Ahmed, Ana Maria Drăgulinescu, El Mehdi Amhoud
Intelligence
Status: succeeded | Model: anthropic/claude-sonnet-4.6 | Prompt: intel-v1 | Confidence: 97%
Last extracted: 3/24/2026, 1:22:20 AM
Summary
This paper proposes Hetero-Net, a unified heterogeneous LoRa framework integrating ground-based WSNs and underground WUSNs with UAV-mounted gateways. The system jointly optimizes spreading factor, transmission power, and 3D UAV placement to maximize energy efficiency. The problem is modeled as a Partially Observable Stochastic Game (POSG) and solved using Multi-Agent Proximal Policy Optimization (MAPPO). Results show 55.81% and 198.49% energy efficiency improvements over isolated WSN-only and WUSN-only deployments respectively.
Entities (29)
Relation Signals (26)
Ana-Maria Drăgulinescu → affiliatedwith → Politehnica Bucharest
confidence 99% · Telecommunications Department, Politehnica Bucharest, Bucharest, Romania
El Mehdi Amhoud → affiliatedwith → Mohammed VI Polytechnic University (UM6P)
confidence 99% · College of Computing, Mohammed VI Polytechnic University (UM6P), Benguerir, Morocco
Abdullahi Isa Ahmed → affiliatedwith → Mohammed VI Polytechnic University (UM6P)
confidence 99% · College of Computing, Mohammed VI Polytechnic University (UM6P), Benguerir, Morocco
El Mehdi Amhoud → authored → Hetero-Net
confidence 99% · Abdullahi Isa Ahmed1, Ana-Maria Drăgulinescu2, El Mehdi Amhoud1 ... propose Hetero-Net
Ana-Maria Drăgulinescu → authored → Hetero-Net
confidence 99% · Abdullahi Isa Ahmed1, Ana-Maria Drăgulinescu2, El Mehdi Amhoud1 ... propose Hetero-Net
Abdullahi Isa Ahmed → authored → Hetero-Net
confidence 99% · Abdullahi Isa Ahmed1, Ana-Maria Drăgulinescu2, El Mehdi Amhoud1 ... propose Hetero-Net
Hetero-Net → optimizes → Energy Efficiency
confidence 99% · Our objective is to maximize system energy efficiency through the joint optimization of the spreading factor, transmission power, and 3D placement of the UAVs
Hetero-Net → uses →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The evolution of Internet of Things (IoT) into multi-layered environments has positioned Low-Power Wide Area Networks (LPWANs), particularly Long Range (LoRa), as the backbone for connectivity across both surface and subterranean landscapes. However, existing LoRa-based network designs often treat ground-based wireless sensor networks (WSNs) and wireless underground sensor networks (WUSNs) as separate systems, resulting in inefficient and non-integrated connectivity across diverse environments. To address this, we propose Hetero-Net, a unified heterogeneous LoRa framework that integrates diverse LoRa end devices with multiple unmanned aerial vehicle (UAV)-mounted LoRa gateways. Our objective is to maximize system energy efficiency through the joint optimization of the spreading factor, transmission power, and three-dimensional (3D) placement of the UAVs. To manage the dynamic and partially observable nature of this system, we model the problem as a partially observable stochastic game (POSG) and address it using a multi-agent proximal policy optimization (MAPPO) framework. An ablation study shows that our proposed MAPPO Hetero-Net significantly outperforms traditional, isolated network designs, achieving energy efficiency improvements of 55.81\% and 198.49\% over isolated WSN-only and WUSN-only deployments, respectively.
Tags
Links
- Source: https://arxiv.org/abs/2603.20404v1
- Canonical: https://arxiv.org/abs/2603.20404v1
Trouble viewing inline? Open PDF directly →
Full Text
38,356 characters extracted from source content.
Expand or collapse full text
Hetero-Net: An Energy-Efficient Resource Allocation and 3D Placement in Heterogeneous LoRa Networks via Multi-Agent Optimization Abdullahi Isa Ahmed1, Ana-Maria Drăgulinescu2, El Mehdi Amhoud1 1College of Computing, Mohammed VI Polytechnic University (UM6P), Benguerir, Morocco. 2Telecommunications Department, Politehnica Bucharest, Bucharest, Romania. Emails: abdullahi.isaahmed, elmehdi.amhoud@um6p.ma, ana.dragulinescu@upb.ro Abstract The evolution of Internet of Things (IoT) into multi-layered environments has positioned Low-Power Wide Area Networks (LPWANs), particularly Long Range (LoRa), as the backbone for connectivity across both surface and subterranean landscapes. However, existing LoRa-based network designs often treat ground-based wireless sensor networks (WSNs) and wireless underground sensor networks (WUSNs) as separate systems, resulting in inefficient and non-integrated connectivity across diverse environments. To address this, we propose Hetero-Net, a unified heterogeneous LoRa framework that integrates diverse LoRa end devices with multiple unmanned aerial vehicle (UAV)-mounted LoRa gateways. Our objective is to maximize system energy efficiency through the joint optimization of the spreading factor, transmission power, and three-dimensional (3D) placement of the UAVs. To manage the dynamic and partially observable nature of this system, we model the problem as a partially observable stochastic game (POSG) and address it using a multi-agent proximal policy optimization (MAPPO) framework. An ablation study shows that our proposed MAPPO Hetero-Net significantly outperforms traditional, isolated network designs, achieving energy efficiency improvements of 55.81% and 198.49% over isolated WSN-only and WUSN-only deployments, respectively. This paper has been accepted for publication in the 2026 IEEE International Conference on Communications (ICC) Workshop ©2026 IEEE. Please cite it as: A.I. Ahmed, A.M. Drăgulinescu and E. M. Amhoud, “Hetero-Net: An Energy-Efficient Resource Allocation and 3D Placement in Heterogeneous LoRa Networks via Multi-Agent Optimization,” in IEEE International Conference on Communications (ICC) Workshop, May 2026. I Introduction In recent years, low-power wide area networks (LPWANs) have emerged as a key enabling technology for autonomous Internet of Things (IoT) applications, offering extended range and energy efficiency for a wide array of devices [5]. Among these technologies, Long Range (LoRa) has proven particularly successful due to its low power consumption, low data rate, and long-range capabilities, making it ideal for large-scale, distributed sensor deployments. Traditionally, LoRa networks have been designed for ground-based wireless sensor networks (WSNs). However, emerging demands from critical sectors, such as smart agriculture, mining, and infrastructure monitoring, are driving the need for connectivity in more challenging environments. In particular, wireless underground sensor networks (WUSNs) have become essential for applications like real-time pipeline monitoring. In these scenarios, energy-efficient and reliable communication is vital despite the severe signal attenuation and path loss caused by underground propagation [8]. These evolving requirements highlight the growing importance of supporting both surface and subterranean connectivity within a unified framework. Extensive literature exists on optimizing LoRa networks for ground-based WSNs, with a strong focus on improving energy efficiency [4, 12, 2]. For instance, the authors of [12] proposed a multi-agent Q-learning algorithm to jointly optimize transmission power and spreading factor in uplink LoRa communication. Similarly, the authors in [2] introduced a distributed reinforcement learning-based scheme for transmission power and channel selection. Conversely, research on WUSNs has addressed the unique constraints of underground environments. In [13], an “Energy Saver” configuration method was proposed, which uses real-time soil moisture data to estimate path loss and adapt LoRa parameters for optimal energy usage. More recently, a multi-agent reinforcement learning (MARL) approach was adopted in [8] to allocate spreading factors in direct-to-satellite LoRa scenarios, aiming to reduce co-spreading factor interference in dense WUSN deployments. Although these studies provide valuable insights within their respective domains, they fall short of addressing heterogeneous LoRa deployments that span both ground and underground environments. Moreover, conventional ground-based gateway deployments suffer from limited coverage and non-line-of-sight (NLoS) challenges, whereas satellite-based solutions, such as those in [8], often introduce high latency, complex infrastructure requirements, and increased power consumption. These limitations highlight the need for a flexible, integrated approach capable of adapting to dynamic network conditions and diverse environmental constraints. In this paper, we propose Hetero-Net, a heterogeneous LoRa-based network architecture that utilizes unmanned aerial vehicles (UAVs) as mobile gateways to collect data from both WSN and WUSN LoRa end devices. Our objective is to maximize system energy efficiency by jointly optimizing the three-dimensional (3D) placement of UAV-mounted gateways, spreading factor, and transmission power across all LoRa end devices. To solve this complex coordination task, we model the optimization problem as a partially observable stochastic game (POSG) and leverage multi-agent proximal policy optimization (MAPPO) algorithm. To the best of our knowledge, this is the first study to address energy efficiency in heterogeneous LoRa deployments involving both ground and underground sensor networks in a multi-agent setting. Our main contributions are summarized as follows: • We propose Hetero-Net, a unified LoRa-based framework that supports both ground-based WSNs and WUSNs. This framework accounts for the physical disparities between these layers by incorporating realistic device positioning and channel modeling for ground-to-air (G2A) and underground-to-air (UG2A) communication links. • We formulate the joint problem of UAV 3D placement and LoRa parameter allocation as a single energy efficiency optimization task, modeled as a POSG. Unlike existing literature that optimizes these parameters in isolation, our framework incorporates tailored state representations, action spaces, and reward functions to balance the divergent requirements of surface and subterranean communication. • We demonstrate through extensive simulations that the proposed MAPPO-based approach outperforms existing deep reinforcement learning (DRL) and non-DRL baselines in terms of energy efficiency. Furthermore, ablation studies confirm that jointly optimizing this heterogeneous LoRa deployment significantly improves overall system performance. The remainder of the paper is structured as follows. Section I presents the system model and the formal problem formulation. Section I details the proposed MAPPO-based solution approach. Simulation results and performance analysis are provided in Section IV. Finally, Section V concludes the paper and outlines directions for future research. I System Model In this work, we consider a LoRa uplink heterogeneous communication system comprising multiple ground and underground LoRa end devices, as well as multiple UAVs. Specifically, the system includes U non-terrestrial UAVs equipped with LoRa gateways, denoted by the set ≜1,2,…,UU \1,2,…,U\, which serve as aerial data collection and forwarding stations. Multiple LoRa end devices are distributed across two terrestrial layers, denoted by the set ≜1,2,…,VV \1,2,…,V\. Each end device v∈v is deployed in one of two layers, i.e., k=0k=0 for the underground layer or k=1k=1 for the ground layer. The end devices are geographically organized into C clusters, represented by the set ≜1,2,…,CC \1,2,…,C\, where each cluster corresponds to a point of interest (POI) in the deployment area. We assume a one-to-one mapping between UAVs and clusters, such that UAV u∈u is assigned to serve cluster c=uc=u. We denote uV_u as the set of all end devices associated with UAV u, which is known a priori based on their proximity to the corresponding POI. Furthermore, uk⊆uV^k_u _u represents the subset of end devices served by UAV u that are located at layer k∈0,1k∈\0,1\. Clearly, u=u0∪u1V_u=V^0_u ^1_u and u0∩u1=∅V^0_u ^1_u= . An illustration of the studied system model is provided in Fig. 1. Figure 1: The studied system model. Moreover, the static wireless ground and underground LoRa end devices are distributed within a target area cA_c. Within each cluster u, the horizontal coordinates of all associated end devices at layer k, denoted by uk=(xvk,yvk):v∈ukG_u^k=\(x_v^k,y_v^k):v _u^k\, follow a bi-variate normal spatial distribution centered at the cluster centroid u μ_u with variance σu2 _u^2, i.e., uk∼(u,σu22)G_u^k ( μ_u, _u^2I_2). Consequently, each end device v∈ukv _u^k is located at 3D position vk=[xvk,yvk,hvk]g_v^k=[x_v^k,y_v^k,h_v^k], where the altitude component is hv0=−dbh_v^0=-d_b, where dbd_b denote the burial depth for underground devices, and hv1=0h_v^1=0 for ground-level devices. Furthermore, each UAV gateway u∈u operates at an altitude hu>0h_u>0, with a 3D position defined as u=[xu,yu,hu]q_u=[x_u,y_u,h_u]. I-A Channel Model I-A1 Underground-to-Air (UG2A) Model Following the model in [7], the total UG2A path loss from an underground LoRa end device v∈u0v ^0_u to UAV gateway u consists of three primary components: 1) the above-ground air attenuation Lu,vairL^air_u,v, 2) the underground soil attenuation LsoilL^soil, and 3) the refraction loss at the air-soil interface LrefL^ref. First, the underground soil attenuation LsoilL^soil model is expressed as [8] Lsoil=(2βdpe−αdp)2,L^soil= ( 2β d_pe^-α d_p )^2, (1) where the attenuation constant α and phase shifting constant β are given by α α =2πfμrμ0ϵ′ϵ02[1+(ϵ′ϵ′)2−1], =2π f _r _0ε _02 [ 1+ ( ε ε )^2-1 ], (2) β β =2πfμrμ0ϵ′ϵ02[1+(ϵ′ϵ′)2+1]. =2π f _r _0ε _02 [ 1+ ( ε ε )^2+1 ]. (3) Here, f is the operating frequency, μr _r is the relative permeability of the soil, μ0 _0 is the permeability of free space, and ϵ0 _0 is the permittivity of free space. The parameters ϵ′ε and ϵ′ε are the real and imaginary parts of the soil’s complex relative permittivity, respectively. Note that ϵr=ϵ′−jϵ′ _r=ε -jε , where the real part ϵ′ε represents energy storage capability, while ϵ′ε quantifies energy dissipation due to losses in the soil. The parameter dp=db/cos(arcsin(1/ϵ′))d_p=d_b/ ( (1/ ε ) ) in Eq.(1) represents the actual underground path length, where dbd_b is the burial depth assuming near-normal propagation. The specific values of these parameters are detailed in [9]. Next, to account for the above-ground air attenuation, we model Lu,vairL^air_u,v for underground device v∈u0v ^0_u as Lu,vair=(4πfc0)2du,vη0,L^air_u,v= ( 4π fc_0 )^2d_u,v^η^0, (4) where η0η^0 is the path loss exponent, c0c_0 is the speed of light, and du,vd_u,v denotes the Euclidean distance between underground device v at position v0g^0_v and UAV u at position uq_u. Lastly, following the approach in [8], we neglect the refraction loss at the air-soil interface, setting Lu,vref=1L^ref_u,v=1. This simplification is justified because most electromagnetic energy is successfully refracted when waves propagate from a high-density medium (soil) to a lower-density medium (air), making the refraction loss negligible. Consequently, we model the overall UG2A path loss as Lu,vUG2A=Lsoil⋅Lu,vair.L^UG2A_u,v=L^soil· L^air_u,v. (5) I-A2 Ground-to-Air (G2A) Model The G2A communication for ground LoRa end devices v∈u1v ^1_u follows a probabilistic path loss model that accounts for both Line-of-Sight (LoS) and Non-Line-of-Sight (NLoS) conditions. The occurrence of LoS or NLoS depends on the environmental characteristics, the elevation angle, and the presence of obstacles. According to [10], the average path loss of the G2A channel between ground end device v and UAV u is modeled as Lu,vG2A=20log(4πfdu,vG2Ac0)+Pu,vLoSηLoSG2A+(1−Pu,vLoS)ηNLoSG2A, splitL^G2A_u,v=20 ( 4π fd^G2A_u,vc_0 )+P^LoS_u,vη^G2A_LoS\\ + (1-P^LoS_u,v )η^G2A_NLoS, split (6) where du,vG2A=‖u−v1‖d^G2A_u,v=\|q_u-g^1_v\| is the Euclidean distance between UAV u and ground end device v, while ηLoSG2Aη^G2A_LoS and ηNLoSG2Aη^G2A_NLoS are environment-dependent parameters that representing the average additive losses for LoS and NLoS links, respectively. Additionally, Pu,vLoSP^LoS_u,v is the LoS probability between ground end device v and UAV u, which can be expressed as Pu,vLoS=11+ϕexp(−φ[θu,vG2A−ϕ]),P^LoS_u,v= 11+φ (- [θ^G2A_u,v-φ ] ), (7) where ϕφ and φ are environment-dependent parameters that characterize the propagation environment. Specifically, larger φ values indicate faster transitions from NLoS to LoS, typical of open or rural environments, while smaller φ reflects smoother transitions in dense urban areas. Similarly, larger ϕφ values correspond to scenarios requiring higher elevation angles to achieve LoS probability, as encountered in urban environments with tall buildings. The elevation angle θu,vG2Aθ^G2A_u,v between ground end device v and UAV u is given by θu,vG2A=arcsin(hu‖u−v1‖).θ^G2A_u,v= ( h_u\|q_u-g^1_v\| ). (8) Finally, we model the channel gain for both communication links. The UG2A channel gain between UAV u and underground device v∈u0v ^0_u is given by gu,v0=10−Lu,vUG2A/10g^0_u,v=10^-L^UG2A_u,v/10, while the G2A channel gain between UAV u and ground device v∈u1v ^1_u is expressed as gu,v1=10−Lu,vG2A/10g^1_u,v=10^-L^G2A_u,v/10. Consequently, the channel gain for any device v∈ukv ^k_u where k∈0,1k∈\0,1\ can be unified as gu,vk=10−Lu,vUG2A/10if k=0 (underground),10−Lu,vG2A/10if k=1 (ground).g^k_u,v= cases10^-L^UG2A_u,v/10&if k=0 (underground),\\ 10^-L^G2A_u,v/10&if k=1 (ground). cases (9) I-B Power Consumption Model I-B1 LoRa End Device Power Consumption Model The power consumption of a LoRa-enabled end device v∈ukv ^k_u at layer k consists mainly of two components: the transmission power PvtxP^tx_v and the circuit power PvcircuitP^circuit_v. Therefore, the total power consumption of end device v during uplink transmission is modeled as [4] Pv=Pvtx+Pvcircuit.P_v=P^tx_v+P^circuit_v. (10) I-B2 UAV Hovering Power Model In addition to the transmission power of LoRa end devices, a UAV also consumes power during various aerial operations such as hovering and steady flight. In our proposed scenario, we consider only the hovering power. The induced hover power for a multirotor UAV u∈u can be expressed as [3] Pu,hover=(1+kind)⋅Wu⋅Wu2ρj¯Arotor,P_u,hover=(1+k_ind)· W_u· W_u2ρ jA_rotor, (11) where WuW_u is the weight of UAV u in Newtons, kindk_ind is the incremental induced-power factor, ρ is the air density in kg/m³, j¯ j is the number of rotors, and ArotorA_rotor is the rotor disk area per rotor in m². The factor (1+kind)(1+k_ind) accounts for non-ideal effects, such as rotor wake interactions and tip losses, that increase the required power beyond the theoretical minimum. I-C Energy Efficiency Model To link the physical-layer transmission quality to the overall system performance, we define the energy efficiency of the LoRa UAV-assisted heterogeneous network by incorporating the signal-to-interference-plus-noise ratio (SINR), the achievable data rate, and the total system power consumption. As a first step, we characterize the uplink transmission quality between an end device and its associated UAV. Specifically, we consider an end device v∈ukv ^k_u at layer k served by UAV u∈u . Therefore, the signal-to-noise ratio (SNR) between v and u is given by ρu,vk=Pvgu,vkσ2ρ^k_u,v= P_vg^k_u,vσ^2, where gu,vkg^k_u,v denotes the channel gain from v to u, PvP_v is the transmit power of v, and σ2σ^2 is the noise power. The achievable uplink data rate from v to u follows the Shannon-Hartley model [1] and is given by ℜu,vk=BWvlog2(1+γu,v,nk), ^k_u,v=BW_v\, _2\! (1+γ^k_u,v,n ), (12) where BWvBW_v is the bandwidth allocated to end device v, n∈7,8,9,10,11,12n∈\7,8,9,10,11,12\ is the spreading factor assigned to device v, and the SINR is γu,v,nk=Pvgu,vk∑k′∈0,1∑v′∈uk′∖vδv′,vPv′gu,v′k′+σ2.γ^k_u,v,n= P_vg^k_u,v _k ∈\0,1\ _v ^k _u \v\ _v ,v\;P_v \,g^k _u,v +σ^2. (13) Here, uk′V^k _u is the set of end devices served by UAV u at layer k′∈0,1k ∈\0,1\, and δv′,v _v ,v is an indicator variable that equals 1 when device v′v uses the same spreading factor as device v, and 0 otherwise. We assume perfect orthogonality between different spreading factors, such that only co-SF devices contribute to interference. Additionally, we neglect inter-UAV interference due to sufficient spatial separation between UAV clusters. Finally, the overall energy efficiency of the system is defined as the sum of per-UAV energy efficiencies, representing the ratio of total throughput to total power consumption, and can be presented as ηE=∑u∈[∑k∈0,1∑v∈ukℜu,vk(∑k∈0,1∑v∈ukPv)+Pu,hover]. _E= _u [\; _k∈\0,1\ _v ^k_u ^k_u,v ( _k∈\0,1\ _v ^k_uP_v )+P_u,hover\; ]. (14) Our goal is to maximize the energy efficiency of the proposed heterogeneous LoRa network. To this end, we jointly optimize the spreading factor assignment, transmission power allocation, and 3D placement of the UAV-mounted gateways. Let the discrete sets of spreading factors and power levels be denoted by and P, respectively. We define the UAV placement decision as ≜(xu,yu,hu)u∈Q \(x_u,y_u,h_u)\_u , the spreading factor assignment as ≜nvv∈n \n_v\_v , and the power allocation as tx≜Pvv∈P_tx \P_v\_v . Our constrained optimization problem is formulated as follows max,tx,ηE _n,P_tx,Q _E (15a) s.t. 0≤xu≤xmax,∀u∈, 0≤ x_u≤ x^max, ∀ u , (15b) 0≤yu≤ymax,∀u∈, 0≤ y_u≤ y^max, ∀ u , (15c) hmin≤hu≤hmax,∀u∈, h_ ≤ h_u≤ h_ , ∀ u , (15d) Pv∈,∀v∈, P_v , ∀ v , (15e) nv∈,∀v∈, n_v∈ , ∀ v , (15f) ρu,vk≥℧¯nv,∀v∈uk,∀u∈,k∈0,1. ρ^k_u,v≥ _n_v,∀ v ^k_u,∀ u ,k∈\0,1\. (15g) Here, constraints (15b) and (15c) confine each UAV within the target area, while constraint (15d) restricts the altitude to the operational range [hmin,hmax][h_ ,h_ ]. Constraints (15e) and (15f) ensure that transmission power and spreading factor selections come from the allowed discrete sets. Finally, constraint (15g) guarantees feasibility by enforcing that the received SNR ρu,vkρ^k_u,v must exceed the minimum sensitivity threshold ℧¯nv _n_v corresponding to the assigned spreading factor for each device v. Note that the typical SNR thresholds for LoRa transceivers operating at BWv=125BW_v=125 kHz are summarized in Table I. The problem in (15) is a mixed-integer nonlinear program due to the discrete spreading factor and power selections, coupled with the non-convex nature of the wireless channel models. It is therefore known to be NP-hard in general. In the next section, we propose a MARL framework to obtain an efficient solution. I Proposed Approach We tackle the optimization problem in (15) using a MARL framework based on centralized training and decentralized execution (CTDE). Specifically, we adopt the MAPPO [11] algorithm, an actor-critic method suitable for partially observable environments. Furthermore, we model the environment as a POSG, defined by the tuple: ⟨,,u,u,,u,ℛu,γ,μ^0⟩ ,S,\A_u\,\O_u\,T,\Z_u\,\R_u\,γ, μ_0 , where U is the set of agents, S is the global state space, u\A_u\ and u\O_u\ are the individual action and observation spaces for each agent, T is the state transition function, u\Z_u\ are the observation functions, ℛu\R_u\ are the individual reward functions, γ is the discount factor, and μ^0 μ_0 is the initial state distribution. TABLE I: SNR Thresholds ℧¯n _n for Different Spreading Factors with BWv=125BW_v=125 kHz Spreading Factor (n) 7 8 9 10 11 12 SNR Threshold ℧¯nv _n_v (dB) -7.5 -10 -12.5 -15 -17.5 -20 The proposed heterogeneous network-based MAPPO algorithm, referred to as MAPPO-Hetero-Net, involves the following elements: • Agents (U): The finite set of agents. • Actions (uA_u): Each agent u∈u has an individual action space uA_u. The action of agent u is denoted as au=[Δxu,Δyu,Δhu,n,P]∈ua_u=[ x_u, y_u, h_u,n,P] _u. Here, (Δxu,Δyu,Δhu)( x_u, y_u, h_u) are the normalized movement deltas in [−1,1][-1,1] scaled by the maximum step size ssteps_step. That is, (Δxu,Δyu,Δhu)=sstep⋅~u,~u∈[−1,1]3( x_u, y_u, h_u)=s_step· d_u,\; d_u∈[-1,1]^3, while n and P are the spreading factor and power assignments, respectively. The joint action space is =∏u∈uA= _u A_u, with joint action =(1,2,…,U)a=(a_1,a_2,…,a_U). • Observations (uO_u): Each agent u receives a local observation ou∈uo_u _u based on its partial view of the environment. Therefore, the observation of agent u is given by ou=[xu,yu,hu,du,v,ρu,vk,nv,Pvv∈u]o_u=[x_u,y_u,h_u,\d_u,v,ρ^k_u,v,n_v,P_v\_v _u], where xu,yu,hux_u,y_u,h_u represent the UAV’s 3D position, du,vd_u,v is the Euclidean distance between UAV u and end device v∈uv _u, ρu,vkρ^k_u,v is the received SNR at layer k, nvn_v is the spreading factor, and PvP_v is the transmit power. The joint observation is denoted o=(o1,o2,…,oU)o=(o_1,o_2,…,o_U). Note that we normalized all distance, SNR, spreading factor, and power features using min-max scaling per cluster to [0,1][0,1] to enhance the learning process. • Rewards (ℛuR_u): Each agent u has an individual reward function ℛu:×→ℝR_u:S×A that balances system-wide performance with local cluster efficiency. We first define a local energy efficiency metric for agent u that reflects how well it manages the end devices in its associated cluster given by ηEElocal(u)=∑k∈0,1∑v∈ukℜu,vk(∑k∈0,1∑v∈ukPv)+Pu,hover _E^local(u)= _k∈\0,1\ _v ^k_u ^k_u,v ( _k∈\0,1\ _v ^k_uP_v )+P_u,hover, where Pu,hoverP_u,hover is the altitude-dependent hovering power of UAV u. Next, we incorporate the global energy efficiency ηE _E from Eq. (14), which captures the system-level trade-off between total throughput and total consumed power across all agents. The reward function for agent u combines both terms and is expressed as ℛu(s,)=ωηE+(1−ω)ηEElocal(u),R_u(s,a)=ω\, _E+(1-ω)\, _E^local(u), (16) where the scalar ω∈[0,1]ω∈[0,1] is a tunable weight that balances the global objective against local cluster efficiency. Each agent u follows a decentralized policy πu:u×u→[0,1] _u:O_u×A_u→[0,1], which maps its local observation to a probability distribution over actions. Therefore, the objective of each agent is to maximize its expected cumulative discounted reward, given as Ju(πu)=τu∼πu[∑t=0Tγtℛu([t],[t])],J_u( _u)=E_ _u _u [ _t=0^Tγ^tR_u(s[t],a[t]) ], (17) where the expectation is taken over the trajectories τu _u generated by the agent u. In the cooperative setting, agents aim to maximize the objective of the team J()=∑u=1UJu(πu)J( π)= _u=1^UJ_u( _u). MAPPO-Hetero-Net operates under the CTDE paradigm, where each agent u∈u maintains a decentralized policy network πθu(au|ou) _ _u(a_u|o_u) mapping local observations to actions, while a centralized critic Vϕ(s)V_φ(s) estimates state values using the global state s during training. Agents compute advantages via generalized advantage estimation (GAE) and update policies simultaneously by maximizing a clipped PPO objective with probability ratio ru(θu)=πθu(au|ou)/πθuold(au|ou)r_u( _u)= _ _u(a_u|o_u)/ _ _u^old(a_u|o_u), where πθuold(au|ou) _ _u^old(a_u|o_u) represents the old policy from the previous iteration with respect to which the ratio is computed. The ratio is constrained within [1−ϵclip,1+ϵclip][1- _clip,1+ _clip] for stability, where the clipping parameter ϵclip _clip prevents destructively large policy updates in a single step, ensuring monotonic improvement and training stability [11]. During deployment, each UAV operates autonomously using only its policy πθu(au|ou) _ _u(a_u|o_u) and local observation ouo_u, eliminating inter-agent communication requirements. A detailed discussion of the MAPPO algorithm and its computational complexity analysis can be found in [6]. (a) (b) (c) Figure 2: (a) 2D configuration with UAVs trajectories, (b) Final UAVs 3D placement, and (c) Final heights of UAVs. IV Simulation Results To evaluate the performance of our approach, we consider a square target area cA_c of size 2 km×2 km2 km× 2 km with 80 LoRa end devices organized into four spatial clusters around POIs. The device locations follow a Gaussian distribution centered at each POI with standard deviation of σu=150 _u=150 m. The four POIs (cluster centroids) are fixed at 1=(500,500)m μ_1=(500,500)\,m, 2=(1500,500)m μ_2=(1500,500)\,m, 3=(500,1500)m μ_3=(500,1500)\,m, and 4=(1500,1500)m μ_4=(1500,1500)\,m. At the start of each training episode, four UAVs are randomly positioned within their respective clusters. The simulation parameters are summarized in Table I. We evaluate our proposed MAPPO-HeteroNet against three groups of methods: (i) heterogeneous MARL state-of-the-art, (i) ablations of our own design, and (i) non-DRL baselines. First, for the MARL state-of-the-art, we include HAPPO [14], a heterogeneous-agent actor-critic with sequential policy updates and trust-region guarantees for monotonic joint improvement, and HAA2C [14], a heterogeneous asynchronous advantage actor-critic with type-specific networks and asynchronous updates for stability. Second, the ablation scenarios isolate the effect of vertical-layer assignment while keeping the device count and spatial distribution fixed. Specifically, MAPPO-UG-Net places all 80 end devices in the underground layer, whereas MAPPO-G-Net deploys all devices to the ground layer. These are compared against the proposed MAPPO-Hetero-Net configuration, which includes a total of 80 end devices equally divided into 40 underground and 40 ground LoRa end devices. Finally, we consider two non-DRL baselines. i.e., a Random approach, which samples UAV actions uniformly at random (movement and PHY parameters), providing a lower bound, and a Fixed+Heuristic approach, where UAVs remain fixed at their POI centers while spreading factor and transmission power parameters are assigned using distance-based allocation. TABLE I: Simulation setup. Sym. Value Sym. Value V 8080 U 44 [hmin,hmax][h_ ,h_ ] [70,150][70,150] m hv0h_v^0, hv1h_v^1 −0.4-0.4 m, 0 m f 868868 MHz BWvBW_v 125125 kHz ϕφ, φ 4.884.88, 0.430.43 ηLoSG2Aη^G2A_LoS, ηNLoSG2Aη^G2A_NLoS 0.10.1 dB, 2121 dB c0c_0 3×1083× 10^8 m/s σ2σ^2, σu _u −120-120 dBm, 150150 m 7,8,9,10,11,12\7,8,9,10,11,12\ P 2,5,8,11,14\2,5,8,11,14\ dBm kindk_ind, WuW_u 0.110.11, 20.020.0 N j¯ j, ρ, ArotorA_rotor 44, 1.1681.168 kg/m3, 0.2140.214 m2 ϵ0 _0 8.854×10−128.854× 10^-12 F/m η0η^0 2.02.0 ϵ′ε , ϵ′ε 18.203018.2030, 0.162870.16287 mvm_v, dpd_p 0.200.20, 0.411460.41146 m μ0 _0, μr _r 4π×10−74π× 10^-7 H/m, 1.01.0 Clay & Sand 10%10\% & 90%90\% αactor _actor, αcritic _critic 3×10−43× 10^-4, 5×10−45× 10^-4 ϵclip _clip, γ, T, TtotT_tot 0.20.2, 0.950.95, 5050, 3M3M βent _ent, ω 0.010.01, 0.30.3 Max steps/episode 100100 Seed 0,44,182,235\0,44,182,235\ Architecture MLP (2 layers, [128,128][128,128]) Optimizer Adam Activation ReLU (a) (b) (c) Figure 3: (a) Training rewards over environment steps. (b) Ablation study. (c) Energy efficiency with benchmarking schemes. Fig. 2 illustrates the initial and final configurations of the four flying gateways in the proposed Hetero-Net scenario. Specifically, Fig. 2(a) shows each UAV beginning from a random location within its cluster and moving towards the barycenter of its assigned LoRa end devices. As depicted, dashed lines indicate the associations between LoRa end devices and their serving flying gateways. Furthermore, Fig. 2(b) presents the final 3D placement. Here, the flying gateways settle near the centers of their respective end-device clusters. They position themselves to serve both ground WSN LoRa end devices and underground WUSN LoRa end devices, while also improving line-of-sight conditions for the ground nodes. Fig. 2(c) depicts the starting point, final point, and flying height of our deployed flying gateways in a 2D plane. As shown from the figure, the UAVs adopt different heights to improve coverage and reduce interference. Fig. 3(a) presents the convergence of cumulative training rewards for the DRL-based algorithms. As observed from the figure, the proposed MAPPO-Hetero-Net achieves the highest episode return, followed by the HAPPO algorithm. In contrast, HAA2C initially starts with higher rewards but begins to fluctuate around 0.5×1060.5× 10^6 steps and shows greater variability. This behavior is likely due to its asynchronous updates and less effective handling of the dynamic characteristics of the environment. In Fig. 3(b), the ablation study results are shown, comparing the heterogeneous setup against homogeneous configurations. Specifically, the performance of MAPPO-UG-Net (underground-only) and MAPPO-G-Net (ground-only) drops significantly, highlighting the importance of vertical-layer diversity in improving cooperative learning and network adaptability. Finally, Fig. 3(c) compares the system energy efficiency across all benchmark schemes. As can be seen from the figure, the proposed MAPPO-Hetero-Net achieves the highest energy efficiency, outperforming G-Net and UG-Net by 55.81% and 198.49%, respectively, and surpassing traditional non-DRL baselines such as Random and Fixed+Heuristic by 298.73% and 181.35%, respectively. V Conclusion In this paper, we investigated uplink data collection in heterogeneous LoRa networks integrating ground-based WSNs and underground WUSNs. We proposed a Hetero-Net framework that jointly optimizes the spreading factor, transmission power, and 3D placement of UAV-mounted gateways to maximize system energy efficiency. The problem was formulated as a POSG and addressed using a MAPPO-based MARL approach within the CTDE paradigm. The proposed model incorporates quality-of-service constraints and the reliability of both G2A and UG2A channel models. Simulation results show that the proposed Hetero-Net framework significantly improves energy efficiency and learning performance compared to homogeneous network designs and non-DRL baselines. Future work will focus on experimental validation in more complex environments, as well as extending the framework to include ground-vehicle-mounted gateways and multi-layered Non-Terrestrial Network (NTN) architectures. References [1] J. Aczél, B. Forte, and C. T. Ng (1974) Why the Shannon and Hartley entropies are ‘natural’. Advances in applied probability 6 (1), p. 131–146. Cited by: §I-C. [2] R. Airiyoshi, M. Hasegawa, T. Ohtsuki, and A. Li (Mar. 2025) Energy Efficient Transmission Parameters Selection Method Using Reinforcement Learning in Distributed LoRa Networks. In IEEE Wirel. Comm. and Net. Conf. (WCNC), Vol. . Cited by: §I. [3] H. Gong, B. Huang, B. Jia, and H. Dai (2023) Modeling Power Consumptions for Multirotor UAVs. IEEE Trans on Aerospace and Electronic Systems 59 (6), p. 7409–7422. Cited by: §I-B2. [4] M. Jouhari, K. Ibrahimi, J. B. Othman, and E. M. Amhoud (Jun. 2023) Deep Reinforcement Learning-Based Energy Efficiency Optimization for Flying LoRa Gateways. In IEEE Inter. Conf. on Communications, Cited by: §I, §I-B1. [5] M. Jouhari, N. Saeed, M. Alouini, and E. M. Amhoud (2023) A Survey on Scalable LoRaWAN for Massive IoT: Recent Advances, Potentials, and Challenges. IEEE Communications Surveys & Tutorials. Cited by: §I. [6] H. Li et al. (2024) Collaborative Task Offloading and Resource Allocation in Small-Cell MEC: A Multi-Agent PPO-Based Scheme. IEEE Trans. on Mobile Computing. Cited by: §I. [7] K. Lin and T. Hao (2021) Experimental Link Quality Analysis for LoRa-Based Wireless Underground Sensor Networks. IEEE Internet of Things Journal 8 (8), p. 6565–6577. Cited by: §I-A1. [8] K. Lin, M. A. Ullah, H. Alves, K. Mikhaylov, and T. Hao (Feb. 2024) Energy Efficiency Optimization for Subterranean LoRaWAN Using a Reinforcement Learning Approach: A Direct-to-Satellite Scenario. IEEE Wireless Communications Letters 13 (2), p. 308–312. Cited by: §I, §I, §I, §I-A1, §I-A1. [9] V. L. Mironov, LG. Kosolapova, and SV. Fomin (2009) Physically and Mineralogically Based Spectroscopic Dielectric Model for Moist Soils. IEEE Trans. on Geos. and Rem. Sens. 47 (7), p. 2059–2070. Cited by: §I-A1. [10] H. Ren et al. (2019) Achievable Data Rate for URLLC-Enabled UAV Systems With 3-D Channel Model. IEEE Wirel. Comm. Letters 8 (6), p. 1587–1590. Cited by: §I-A2. [11] C. Yu et al. (2022) The surprising effectiveness of PPO in cooperative multi-agent games. Advances in neural information processing systems 35, p. 24611–24624. Cited by: §I, §I. [12] Y. Yu, L. Mroueh, S. Li, and M. Terré (Sep. 2020) Multi-Agent Q-Learning Algorithm for Dynamic Power and Rate Allocation in LoRa Networks. In IEEE 31st Annual International Symposium on Personal, Indoor and Mobile Radio Communications, Cited by: §I. [13] H. Zhang, G. Liu, and T. Jiang (2025) Parameter Configuration Scheme for Optimal Energy Efficiency in LoRa-Based Wireless Underground Sensor Networks. IEEE Trans on Vehicular Technology. Cited by: §I. [14] Y. Zhong et al. (2024) Heterogeneous-Agent Reinforcement Learning. Journal of Machine Learning Research 25 (32), p. 1–67. Cited by: §IV.