Paper deep dive
KC-BFPRL: Knowledge-Guided Multi-UAV Collaboration for Grassland Restoration via Bilevel Formerpointer-Based Reinforcement Learning
Dongbin Jiao, Xianyi Wang, Yuchen Yuan, Weibo Yang, Peng Yang, Peng Zhao, Zhanhuan Shang, Shi Yan
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 8/18/2026, 4:10:29 AM
Summary
The paper proposes KC-BFPRL, a knowledge-guided collaborative bilevel formerpointer reinforcement learning framework for multi-UAV grassland restoration. It addresses the Restoration Area Maximization Problem (RAMP) by decomposing it into global task allocation and local restoration planning (trajectory and area allocation). The framework uses a Transformer-based encoder and Pointer Network decoder trained with an actor-critic method, incorporating ecological rules to solve the RL cold-start problem. Experiments show it outperforms baselines like MAPDP with a 0.00% optimality gap in complex scenarios and faster inference.
Entities (8)
Relation Signals (6)
KC-BFPRL → appliedto → Grassland Restoration
confidence 95% · KC-BFPRL: Knowledge-Guided Multi-UAV Collaboration for Grassland Restoration
KC-BFPRL → solves → RAMP
confidence 95% · We propose a novel knowledge-guided collaborative bilevel formerpointer reinforcement learning framework (KC-BFPRL) to address this complexity [RAMP].
KC-BFPRL → uses → Transformer
confidence 95% · Our specialized architecture pairs featuring a Transformer-based encoder that fuses static environmental features with dynamic UAV states
KC-BFPRL → uses → Pointer Network
confidence 95% · and a Pointer Network decoder trained via a robust actor-critic framework.
KC-BFPRL → basedon → Deep Reinforcement Learning
confidence 90% · this paper proposes a novel framework based on deep reinforcement learning (DRL).
KC-BFPRL → outperforms → MAPDP
confidence 90% · It maintains a 0.00% optimality gap in the most complex scenarios U8-R160 and operates nearly three times faster than MAPDP
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Multi-unmanned aerial vehicle (UAV) systems provide scalable service platforms for large-scale environmental tasks, such as grassland ecosystem restoration. However, coordinating fleet operations requires solving the restoration area maximization problem (RAMP). This non-linear combinatorial optimization challenge is complicated by payload-dependent energy dynamics and heterogeneous ecological degradation. We propose a novel knowledge-guided collaborative bilevel formerpointer reinforcement learning framework (KC-BFPRL) to address this complexity. Using a hierarchical paradigm, KC-BFPRL decomposes RAMP into global task allocation and local restoration planning, with the latter further divided into upper-level trajectory planning and lower-level restoration area allocation. Our specialized architecture pairs featuring a Transformer-based encoder that fuses static environmental features with dynamic UAV states, and a Pointer Network decoder trained via a robust actor-critic framework. By embedding ecological priority rules and heuristic logic, KC-BFPRL achieves a structured warm-start, solving the RL cold-start problem while ensuring strict constraint satisfaction. Extensive experiments demonstrate that KC-BFPRL consistently outperforms state-of-the-art baselines, achieving superior objective values and efficiency. It maintains a $0.00\%$ optimality gap in the most complex scenarios U8-R160 and operates nearly three times faster than MAPDP, validating its robustness, scalability, and real-time applicability for large-scale automated ecological restoration.
Tags
Links
- Source: https://arxiv.org/abs/2608.16326v1
- Canonical: https://arxiv.org/abs/2608.16326v1
Trouble viewing inline? Open PDF directly →
Full Text
78,676 characters extracted from source content.
Expand or collapse full text
KC-BFPRL: Knowledge-Guided Multi-UAV Collaboration for Grassland Restoration via Bilevel Formerpointer-Based Reinforcement Learning Dongbin Jiao12, Xianyi Wang1, Yuchen Yuan1, Weibo Yang5, Peng Yang4, , Peng Zhao1, Zhanhuan Shang6, and Shi Yan1 Affiliation: 1School of Information Science and Engineering, Lanzhou University, Lanzhou, 730000, P. R. China (e-mail: jiaodb, wxianyi2025, yuanych2020, zhaopeng, yanshi@lzu.edu.cn). Affiliation: 2Key Laboratory of Tourism Information Fusion Processing and Data Ownership Protection, Ministry of Culture and Tourism, Lanzhou University, Lanzhou, 730000, P. R. China. Affiliation: 4Department of Statistics and Data Science, Southern University of Science and Technology, Shenzhen 518055, P. R. China (e-mail: yangp@sustech.edu.cn). Affiliation: 5School of Automobile, Chang’an University, Xi’an, 710064, P. R. China (e-mail: wbyang@chd.edu.cn). Affiliation: 6State Key Laboratory of Grassland Agro-Ecosystem, College of Ecology, Lanzhou University, Lanzhou, 730000, P. R. China (e-mail: shangzhh@lzu.edu.cn). Abstract Multi-unmanned aerial vehicle (UAV) systems provide scalable service platforms for large-scale environmental tasks, such as grassland ecosystem restoration. However, coordinating fleet operations requires solving the restoration area maximization problem (RAMP). This non-linear combinatorial optimization challenge is complicated by payload-dependent energy dynamics and heterogeneous ecological degradation. We propose a novel knowledge-guided collaborative bilevel formerpointer reinforcement learning framework (KC-BFPRL) to address this complexity. Using a hierarchical paradigm, KC-BFPRL decomposes RAMP into global task allocation and local restoration planning, with the latter further divided into upper-level trajectory planning and lower-level restoration area allocation. Our specialized architecture pairs featuring a Transformer-based encoder that fuses static environmental features with dynamic UAV states, and a Pointer Network decoder trained via a robust actor-critic framework. By embedding ecological priority rules and heuristic logic, KC-BFPRL achieves a structured warm-start, solving the RL cold-start problem while ensuring strict constraint satisfaction. Extensive experiments demonstrate that KC-BFPRL consistently outperforms state-of-the-art baselines, achieving superior objective values and efficiency. It maintains a 0.00%0.00\% optimality gap in the most complex scenarios U8-R160 and operates nearly three times faster than MAPDP, validating its robustness, scalability, and real-time applicability for large-scale automated ecological restoration. Index Terms: Multi-UAV collaboration, grassland restoration, knowledge-guided learning, deep reinforcement learning (DRL), trajectory planning, restoration area allocation. I Introduction Grassland ecosystems are vital for global ecological stability, yet accelerating environmental degradation demands scalable and efficient restoration interventions [29, 34]. While traditional manual methods are prohibitively labor-intensive, the integration of unmanned aerial vehicles (UAVs) has introduced a transformative paradigm in ecological engineering by offering high operational flexibility and low deployment costs. Despite these advantages, current UAV frameworks are predominantly designed for passive monitoring. Transitioning to active restoration, such as aerial seeding, introduces profound operational complexities. Unlike lightweight monitoring missions, carrying heavy seed payloads drastically depletes battery reserves and constrains aerodynamic maneuverability. Consequently, the severe energy and payload limitations of a single UAV render it insufficient for executing large-scale restoration across degraded terrains [13]. To overcome this bottleneck, a strategic shift toward multi-UAV collaborative systems is imperative [26, 40]. While fleet coordination enables parallel execution, orchestrating such a system introduces immense computational complexity, formulated here as the restoration area maximization problem (RAMP). Distinct from standard vehicle routing problem (VRP) formulations, RAMP represents a highly coupled, non-linear combinatorial optimization challenge characterized by heterogeneous degradation levels and stringent resource constraints [13]. The system dynamics exhibit strong nonlinearity and intricate coupling among decision variables: different regions require distinct seeding densities, while a UAV’s energy expenditure is non-linear, decreasing progressively with payload discharge. As a result, RAMP entails a tightly coupled joint optimization of task allocation and path planning, with the primary objective of maximizing the total restored areas before battery depletion. Existing approaches to solving such complex collaborative problems generally fall into two categories: heuristic optimization and deterministic methods [18]. Traditional meta-heuristics (e.g., Genetic Algorithms (GA), Particle Swarm Optimization (PSO)) are computationally lightweight but demand substantial domain expertise, lack cross-scale generalization, and frequently stagnate in inferior local optima [32]. Conversely, deterministic methods struggle with the high dimensionality and dynamic state spaces of multi-agent environments, rapidly becoming computationally intractable as fleet and network sizes expand [11]. To address the longstanding trade-off between solution quality and computational efficiency, this paper proposes a novel framework based on deep reinforcement learning (DRL). DRL is particularly suited to this problem due to its capacity to handle high-dimensional state spaces and learn nonlinear decision policies directly from interaction data [19]. Building on this capability, we introduce the knowledge-guided collaborative bilevel formerpointer reinforcement learning framework, termed KC-BFPRL. To effectively fuse static environmental features with dynamic UAV states, KC-BFPRL incorporates a specialized architecture featuring a Transformer-based encoder and a Pointer Network decoder trained via a robust actor-critic mechanism. Departing from inefficient, undirected exploration, KC-BFPRL utilizes ecological priority rules and heuristic scheduling logic as an inductive scaffold [42]. This structured domain knowledge explicitly addresses the cold-start problem, significantly enhancing sample efficiency and policy feasibility [27]. To implement this vision, the grassland restoration process is modeled as a hierarchical Markov decision process (MDP), in which each UAV as an intelligent agent. This design decomposes the complex RAMP into global task allocation and local restoration planning. Specifically, the local planning subproblem is further decoupled into two tightly interdependent levels: upper-level trajectory planning, where UAVs optimize visitation order to minimize energy waste; and lower-level restoration area allocation, which dynamically determines the optimal seeding quantity at each node to balance ecological impact against residual payload and energy reserves. By integrating these levels, KC-BFPRL provides a principled warm-start mechanism that enables UAVs to efficiently acquire complex cooperative behaviors while simultaneously optimizing individual energy states, ultimately achieving enhanced restoration coverage with rapid, real-time inference capability. The main contributions are summarized as follows: (1) We mathematically model large-scale multi-UAV grassland restoration as RAMP, a complex, non-linear optimization problem coupling routing, task allocation, and payload-dependent energy dynamics. (2) We propose KC-BFPRL, a knowledge-guided bilevel reinforcement learning framework that integrates ecological priorities and heuristic logic into a hierarchical decision-as-a-service architecture. (3) Experiments on 10800 instances demonstrate that The KC-BFPRL shows KC-BFPRL outperforms baselines, maintaining a 0.00%0.00\% optimality gap in complex cases and achieving nearly three times faster inference than MAPDP with robust generalization. The remainder of this paper is organized as follows. Section I reviews the related literature. Section I details the system model and problem formulation. Section IV introduces the proposed KC-BFPRL framework. Section V presents the experimental evaluation, and Section VI concludes this work. I Related Work This section reviews the methodological shift from traditional heuristics to knowledge-guided, data-driven paradigms for solving RAMP. By examining the intersection of precision agriculture, combinatorial optimization, and MARL, identifying critical research gaps that motivate our KC-BFPRL framework. I-A UAV-based Ecological Restoration and Agriculture. The application of UAVs in precision agriculture has evolved significantly over the past decade. Initially, UAVs served primarily as mobile platforms for remote sensing and environmental monitoring [1, 23, 3]. early studies predominantly leveraged hyperspectral imaging to classify vegetation health and assess degradation levels [38, 22]. However, recent advances in payload capacity have transformed these platforms from passive observers to active actuators capable of complex tasks like aerial seeding [21]. Unlike traditional agricultural spraying that assumes uniform coverage [15, 24], grassland restoration involves highly heterogeneous environments where seeding requirements vary strictly by local degradation levels [13]. Most existing studies focus on simple coverage metrics and fail to account for the complex coupling between a UAV’s limited payload, energy constraints, and the spatially varying urgency of restoration demands. I-B Multi-UAV Task Allocation and Path Planning. To address the challenges of routing and task allocation for multi-UAV system, early studies rely heavily on mathematical programming and heuristic algorithms. Exact methods, such as mixed-integer linear programming (MILP), can yield global optimal solutions but are constrained by NP-hardness, making them computationally intractable for large-scale scenarios involving dozens of tasks [14]. As a result, heuristic and meta-heuristic approaches have emerged as the prevailing solution paradigm. Algorithms including GA, PSO, and ant colony optimization (ACO) have been extensively applied to the VRP and its variants [37, 6, 41, 39]. A representative example is CHAPBILM [13], a coupled heuristic framework designed explicitly for bi-level optimization in grassland restoration. Although these heuristic approaches provide stable and interpretable solutions, they incur substantial computational overhead. As the scale of degraded regions expands, their iterative search processes exhibit near-exponential growth in execution time. Furthermore, these methods typically require complete re-optimization from scratch whenever environmental parameters shift. This inherent inflexibility precludes their deployment in time-critical, large-scale scenarios that demand rapid, real-time decision-making and dynamic replanning. I-C DRL for Multi-Agent Collaboration. To reduce the computational bottlenecks of heuristic optimization and deterministic methods, recent research has increasingly shifted toward DRL. Within this paradigm, the joint problem of routing and allocation is formulated as a neural combinatorial optimization (NCO) challenge [36]. Pioneering works, such as Pointer Networks [33] and the attention model [17], have demonstrated that neural architectures can learn to construct near-optimal solutions for the Traveling Salesman Problem (TSP) and VRP almost instantaneously following offline training [12]. In the multi-UAV domain, MARL frameworks, such as MADDPG and QMIX, have been adopted to enable decentralized cooperation among agents [25]. Recent advances include methods like MAPDP, which exploits context embeddings to decompose large-scale problems [43], and CAMP, which leverages attention mechanisms to facilitate inter-agent communication [10]. However, despite their rapid inference speeds, these purely learning-based approaches face profound challenges in strictly constrained restoration scenarios. Lacking prior domain knowledge, agents depend entirely on undirected exploration, resulting in slow convergence and a pronounced cold-start problem. Moreover, their centralized training paradigms often collapse as the number of agents and nodes increases, a fundamental limitation known as the curse of dimensionality. Most importantly, the inherent “black box” nature of pure DPL makes it difficult to rigorously enforce hard operational constraints, such as energy limits and payload capacities, frequently resulting in infeasible restoration planning. I-D Knowledge-Guided and Hybrid Learning Approaches. The inherent limitations of purely heuristic and strictly learning-based approaches have catalyzed a growing interest in knowledge-guided paradigms. These methods aim to embed domain-specific rules or expert priors directly into the learning process to guide state-space exploration and ensure solution feasibility [9, 28]. Existing hybrid approaches typically depend on heuristics to generate demonstration data for imitation learning or utilize RL to select high-level heuristic operators [5]. However, these existing approaches rarely address the tightly coupled, non-linear constraints characteristic of ecological restoration within a unified, end-to-end architecture. To bridge these gaps, this paper proposes KC-BFPRL, a knowledge-guided collaborative bilevel formerpointer RL framework for large-scale grassland restoration. Unlike pure MARL methods (e.g., MAPDP), KC-BFPRL explicitly integrates ecological priority and heuristic scheduling logic into the policy learning process, providing a structured warm-start that effectively mitigates the cold-start problem. In contrast to traditional heuristic approaches (e.g., CHAPBILM), KC-BFPRL shifts the computational burden to the offline training phase, enabling rapid real-time inference during deployment. By synergistically combining domain-structured optimality with the adaptability and scalability of DRL, KC-BFPRL offers a robust and efficient solution for constrained, large-scale grassland restoration planning. I System Model and Problem Formulation This section details the multi-UAV collaborative grassland restoration model, UAV energy consumption dynamics, and the RAMP formulation. I-A Multi-UAV Collaborative Grassland Restoration Model As depicted in Fig. 1, we considers a scenario where a set U of homogeneous UAVs equipped with BeiDou Navigation Satellite System (BDS) modules and seed dispensers. These UAVs fly at a fixed altitude with a constant flight speed. The mission is coordinated from a central Base Station (BS), which handles scheduling, maintenance, and data processing. array[]l [width]Multi-UAV-Collaborative.pdf\\ array Fig. 1: An example of multi-restored regions by multi-UAV collaborative. The collaborative restoration process is modeled as a complete directed weighted graph G=(V,A)G=(V,A). The vertex set V=v0,v1,…,vNV=\v_0,v_1,…,v_N\ includes the BS v0v_0 (the depot) and the degraded regions Va=V∖v0V_a=V \v_0\. The arc set A=aij=(vi,vj)|vi,vj∈V,i≠jA=\a_ij=(v_i,v_j)|v_i,v_j∈ V,i≠ j\ represents the flight paths with a Euclidean distance dijd_ij. Each region viv_i has a degradation severity level li∈(0,1)l_i∈(0,1), where a higher score indicates more severe degradation (represented by lighter colors in Fig. 1). Following international ecological restoration standards [8], regions with li<0.3l_i<0.3 can self-recovery, while those with li>0.8l_i>0.8 are beyond effective UAV intervention. Consequently, we focuses on the critical interval li∈[0.3,0.8]l_i∈[0.3,0.8], where UAV seeding significantly accelerates recovery and reduces costs [13]. Spatially, each region viv_i is discretized into cic_i unit circles. A UAV hovers over a circle to sow a seed quantity determined by lil_i. All UAVs depart from the BS with maximum energy EmaxE_max and seed payload Q, and must complete their tasks and return to the BS before energy depletion. I-B UAV Energy Consumption Model Following [13], a UAV’s mission energy consumption comprises seeding EsE^s, aerial photography EapE^ap, and flight EfE^f components. As established in [7], flight energy at a constant altitude and speed is directly proportional to the total payload weight. I-B1 Energy Consumption for Seeding The total energy required to dispense seeds across all restored regions is formulated as: Es=∑i=1N∑j≠iNσieixij, E^s=Σ^N_i=1Σ^N_j≠ i _ie_ix_ij, (1) where σi _i is the number of restored unit circles in region viv_i, and xijx_ij indicates whether a UAV travels from viv_i to vjv_j. The unit seeding energy is ei=ηqie_i=η q_i, where η>0η>0 is a coefficient and qi=(1+li)γq_i=(1+l_i)^γ is the seed weight, with γ being an environmental-specific grassland parameter. I-B2 Energy Consumption for Aerial Photography The total energy consumed by the onboard hyperspectral camera to acquire data at all restored regions is: Eap=eap∑i=1N∑j≠iNσixij, E^ap=e^apΣ^N_i=1Σ^N_j≠ i _ix_ij, (2) where eape^ap is the data acquisition energy per unit circle. I-B3 Energy Consumption for Flight The flight energy depends on travel distance and varying payload weight: Ef=∑i=1N∑j≠iNeijfdijxij, E^f=Σ^N_i=1Σ^N_j≠ ie^f_ijd_ijx_ij, (3) where eijfe^f_ij is the energy consumption rate per unit distance along edge (i,j)(i,j), defined as [7]: eijf=P(q¯ij)=(M+q¯ij)32g32ρςh, e^f_ij=P( q_ij)=(M+ q_ij) 32 g^32ρ h, (4) where q¯ij q_ij is the seed payload from viv_i to vjv_j, and M=W+mM=W+m is the UAV’s tare weight (frame weight W and battery weight m). The constants g, ρ, ς , and h denote gravitational acceleration, air density, rotor disc area, and the number of rotors, respectively. I-C Restoration Area Maximization Model While collaborative multi-UAV systems significantly enhance restoration coverage and efficiency through coordinated task allocation, individual UAV remains confronted with stringent individual constraints of limited flight endurance and payload capacity. These persistent energy limitations preclude the complete restoration of all degraded sites in a single mission. Consequently, it is necessary to develop resource-aware optimization strategies that prioritize ecologically critical regions to maximize the restoration impact within available energy budgets. To quantify this impact, we define the optimization objective function C as follows. C=∑i=1N[1+(li−0.3)]σi, C= _i=1^N [1+(l_i-0.3) ] _i, (5) where the term li−0.3l_i-0.3 serves as an ecological weight factor, explicitly prioritizing regions with higher degradation severity levels lil_i to maximize the environmental benefit of the intervention. I-D Mathematical Model The objective of the multi-UAV collaborative scheduling is to maximize the weighted sum of restored areas while strictly adhering to energy and payload constraints. The problem is formulated as a mixed-integer nonlinear programming (MINLP) model: maxxuijσui _ subarraycx_uij\\[-1.0pt] _ui subarray ∑u∈U∑i=1N∑j≠iN(1+li−0.3)σuixuij _u∈ U _i=1^N _ subarraycj≠ i subarray^N(1+l_i-0.3)\, _ui\,x_uij (6a) s.t. ∑i=1N∑j≠iN(σuieui+eapσui)xuij _i=1^N _ subarraycj≠ i subarray^N( _uie_ui+e^ap _ui)x_uij +∑i=0N∑j≠iNeuijfdijxuij≤Emax,∀u∈U, + _i=0^N _ subarraycj≠ i subarray^Ne_uij^fd_ijx_uij≤ E_ , ∀ u∈ U, (6b) ∑u∈U∑i=1i≠jNσuiquixuij≤Q,∀j∈Va _u∈ U _ subarrayci=1\\ i≠ j subarray^N _uiq_uix_uij≤ Q,\;∀ j∈ V_a (6c) ∑j=0j≠iNq¯uji−∑j=0j≠iNq¯uij=σuiqui,∀i∈Va,∀u∈U _ subarraycj=0\\ j≠ i subarray^N q_uji- _ subarraycj=0\\ j≠ i subarray^N q_uij= _uiq_ui,\;∀ i∈ V_a,∀ u∈ U (6d) q¯uij≤Qxuij,∀(i,j)∈A,∀u∈U q_uij≤ Qx_uij,\;∀(i,j)∈ A,∀ u∈ U (6e) ∑u∈U∑j=0j≠iNxuji=∑u∈U∑j=0j≠iNxuij=1,∀i∈Va _u∈ U _ subarraycj=0\\ j≠ i subarray^Nx_uji= _u∈ U _ subarraycj=0\\ j≠ i subarray^Nx_uij=1,\;∀ i∈ V_a (6f) ∑j=1Nxu0j=∑j=1Nxuj0=1,∀u∈U _j=1^Nx_u0j= _j=1^Nx_uj0=1,\;∀ u∈ U (6g) xuij∈0,1,∀(i,j),∀u∈U x_uij∈\0,1\,\;∀(i,j),∀ u∈ U (6h) 1≤σui≤cui,σui∈ℕ+,∀i∈Va,∀u∈U 1≤ _ui≤ c_ui, _ui ^+,∀ i∈ V_a,\ ∀ u∈ U (6i) q¯uij≥0,∀(i,j),∀u∈U. q_uij≥ 0,\;∀(i,j),∀ u∈ U. (6j) Here, the binary decision variable xuijx_uij equals 1 if UAV u traverses arc (vi,vj)(v_i,v_j). Constraints (6b) ensure that the total energy consumed by both operations and flight does not exceed the capacity EmaxE_max for each UAV. Constraints (6c) require that the total seed weight Q carried by each UAV must be fully dispensed before returning to the base station. Constraints (6d) govern the payload dynamics: the seed weight carried by each UAV decreases by exactly the amount required at each restored region, while simultaneously eliminating illegal subtours. Constraints (6e) guarantee that the seed demand at each restoration region vjv_j does not exceed the remaining payload capacity of the servicing UAV. Constraints (6f) ensures that each UAV visits each restoration region at most once and departs after completing the seeding operation. Constraints (6g) requires that each UAV route begins and terminates at the base station. Constraint (6h) enforces binary integrality on the decision variables. Constraint (6i) limits the restoration work at each region to not exceed its maximum capacity. Constraint (6j) imposes nonnegativity restrictions on all relevant variables. The optimization problem (6) is a complex, NP-hard VRP variant with variable demands and nonlinear costs. Its direct solution is intractable due to three intrinsic challenges: (1) coupled decision variables, as restoration region size, seed demand, and UAV trajectories are strictly interdependent; (2) dynamic problem structure, where payload-dependent energy consumption fluctuates as service sequences evolve; and (3) high computational complexity, making exact optimization methods infeasible for large-scale instances. IV Methodology This section outlines the problem decomposition and challenges of the RAMP for multi-UAV collaboration grassland restoration. It then introduces the BFPRL for single-UAV trajectory planning and local restoration decisions, and finally details the KC-BFPRL for solving the large-scale RAMP. array[]l [width]Framework.pdf\\ array Fig. 2: Knowledge-guided multi-UAV collaborative framework. IV-A Problem Decomposition As illustrated in 2, the large-scale problem is decomposed into two interrelated subproblems: multi-UAV task allocation and single-UAV restoration planning. The former assigns specific regions to individual UAVs to balance fleet workloads and maximize resource efficiency. The latter optimizes the trajectory and operational coverage of each UAV, maximizing the total restored area while minimizing energy consumption. These two subproblems are solved through a hierarchical and collaborative mechanism. The task allocation layer provides the structural framework that guides subsequent restoration planning. Reciprocally, the outcomes of the planning layer dynamically inform and update the task allocation, establishing a real-time feedback loop that ensures the mission is executed cooperatively. Furthermore, the single-UAV restoration planning subproblem is further decomposed into a tightly coupled bilevel structure: upper-level trajectory planning and lower-level restoration area allocation. This bilevel coupling is detailed in Section I-D and is consistent with structures discussed in existing literature [13]. IV-B Single-UAV Trajectory Planning and Restoration Area Allocation IV-B1 Modeling RAMP via BFPRL Modeled as an intelligent agent starting from depot v0v_0, each UAV employs a stochastic policy πθ _θ to generate a trajectory τ=(vti,ati)t=0Tτ=\(v^i_t,a^i_t)\^T_t=0, where vtiv^i_t and atia^i_t denote the visited node and restored area at step t, respectively. To maximize the total restored areas under operational constraints, we formulate this process as a four-component MDP and train πθ _θ using the REINFORCE algorithm with a Greedy Rollout Baseline [2] to maximize the expected objective in Eq. (5). State At each step t, the MDP state is defined as t=(t,qtrem,Etrem,t,t)s_t=(v_t,q_t^rem,E_t^rem,m_t, ξ_t), where qtremq_t^rem and EtremE_t^rem denote the remaining payload and energy, t∈0,1Nm_t∈\0,1\^N is the visited-node mask vector, and t ξ_t tracks the remaining restorable unit circles per region. Each node ti=(xi,yi,li)v^i_t=(x_i,y_i,l_i) contains its static geographical coordinates (xi,yi)(x_i,y_i) and degradation level lil_i. To avoid ambiguity, action variables are not part of the state, and they are denoted separately in the Action paragraph. In the BFPRL model, the UAV’s state at decision step i is denoted as <is_<i, where i∈[1,n+1]∩ℕ+i∈[1,n+1] ^+. The initial state is <1=(v0,Q,Emax,0,0)s_<1=(v_0,Q,E_ ,m_0, ξ_0), where only the depot is marked as visited in 0m_0. Termination occurs when the UAV returns to the depot or no feasible actions remain. Action In BFPRL, the action at step i expands the partial solution <is_<i by selecting ai=(ji,δi)a_i=(j_i, _i), where jij_i is the next selected node and δi _i is the seeding amount (restored circles) executed at jij_i. To ensure operational viability, a feasibility mask ℳi(j,δ)∈0,1M_i(j,δ)∈\0,1\ restricts sampling to unvisited nodes that satisfy all payload and energy constraints, including a safe return to the depot. This stepwise formulation enables the joint optimization of trajectory planning and area allocation. Transition State transitions are deterministic, i.e., (s<i+1∣s<i,si)=1P (s_<i+1 s_<i,s_i )=1. This means that selecting action ai=(ji,δi)a_i=(j_i, _i) in state <is_<i transitions the system to the next state <i+1s_<i+1 with probability 1. Crucially, the transition also includes strict resource updates: remaining energy becomes Erem(i+1)=Erem(i)−Ec(i)E^rem(i+1)=E^rem(i)-E^c(i), where the total energy consumption at step i is equal to Ef(i)+Es(i)+Eap(i)E^f(i)+E^s(i)+E^ap(i). Meanwhile, the payload becomes qi+1rem=qirem−δiqjiq_i+1^rem=q_i^rem- _iq_j_i. The visitation mask and remaining restorable areas are updated via i+1[ji]=1m_i+1[j_i]=1 and ξi+1(ji)=max0,ξi(ji)−δi _i+1(j_i)= \0, _i(j_i)- _i\. Infeasible actions are masked to −∞-∞ prior to sampling. If an infeasible action is still selected due to numerical instability, the episode is immediately terminated with a penalty and a forced return to the depot. Reward Guided by this principle and the specific characteristics of the RAMP, we design a multi-component reward function R(π∣V)R(π V). This function balances the dual objectives of restoration area maximization and operational energy constraints, defined as follows: R(π∣V)=αr⋅∑i=1n(1+(li−0.3))σi−αp⋅pe,R(π V)= _r· _i=1^n (1+(l_i-0.3) ) _i- _p· p_e, (7) where αr _r and αp _p are balancing coefficients. The penalty pep_e is triggered if the energy constraint is violated: pe=∑i=0n−1di,i+1+dn,0,Erest<00,Erest≥0,p_e= cases _i=0^n-1d_i,i+1+d_n,0,&E_rest<0\\[5.69054pt] 0,&E_rest≥ 0, cases (8) where di,i+1d_i,i+1 denotes the flight distance between consecutive restored regions viv_i and vi+1v_i+1. If the energy constraint is violated, the penalty equals the total path length. This design encourages the model to prioritize shorter, feasible paths during early training stages while optimizing restoration allocation. Policy The restoration process is modeled as a sequence of decisions, where the action at step i is ai=(ji,δi)a_i=(j_i, _i) and the policy is factorized via the chain rule: p(π∣s)=∏i=1np(ji∣s<i)p(δi∣ji,s<i),p(π s)= _i=1^np (j_i s_<i )\,p ( _i j_i,s_<i ), (9) where p(ji∣s<i)p (j_i s_<i ) denotes the probability of selecting the next node, and p(δi∣ji,s<i)p ( _i j_i,s_<i ) denotes the conditional probability of selecting the restoration level after node jij_i is chosen. This factorization explicitly captures the sequential dependency of the joint action on the historical trajectory of visited states and actions. IV-B2 BFPRL Architecture Design To learn the stochastic policy π in Eq. (9), we parameterize it as a neural network πθ _θ, where θ∈Θθ∈ denotes the complete set of trainable parameters across the encoder and decoder modules, as shown in Fig. 3. array[]l [width]BFPRL.pdf\\ array Fig. 3: Architecture of BFPRL. State Feature Extraction The model MθM_θ processes both static and dynamic state features. Static features encompass intrinsic environmental attributes, such as the spatial coordinates and degradation levels of target regions. In contrast, dynamic features capture the time-varying operational status of the UAV, such as remaining energy and payload weight. Optimizing Θ enables the model to perform high-dimensional feature embedding and sequential decision modeling based on these extracted state representations. Improved Encoder-Decoder Architecture Building on [4], we propose a tailored encoder-decoder architecture for the RAMP. As illustrated in Fig. 3, the model adopts a Transformer-based encoder [31] to extracts global contextual representations from input sequences. Subsequently, an Pointer Network [33] is integrated as an autoregressive decoder to sequentially generate environment-dependent restoration decisions. This enables the joint optimization of dynamic UAV trajectories and adaptive area allocation. Encoder The encoder integrates heterogeneous inputs, including static environmental features (coordinates, degradation) and dynamic UAV states (energy, payload), to generate a global contextual representation. It consists of nlayersn_layers Transformer blocks utilizing Batch Normalization (BN) instead of standard LayerNorm to stabilize training. First, input features in∈ℝn×dinH_in ^n× d_in are linearly embedded into dimension ded_e: l=0=inin∈ℝn×de,H^l=0=H_inW_in ^n× d_e, (10) where inW_in is the input embedding weight matrix. Each subsequent layer l∈0,…,nlayers−1l∈\0,…,n_layers-1\ applies multi-head attention (MHA) and a feed-forward network (FFN) with residual connections and BN: rcl+1 ^l+1_rc =BN(MHAl+1(l)+l), =BN (MHA^l+1(H^l)+H^l ), (11) l+1 ^l+1 =BN(ReLU(rcl+11l+1)2l+1+rcl+1), =BN (ReLU (H^l+1_rcW_1^l+1 )W_2^l+1+H^l+1_rc ), (12) where MHAl+1(⋅)MHA^l+1(·) denotes the attention mechanism at layer l+1l+1, rcl+1H^l+1_rc is the intermediate residual connection, 1l+1W_1^l+1 and 2l+1W_2^l+1 are the weight matrices of the FFN, and BN(⋅)BN(·) denotes batch normalization. The final output (nlayers)H^(n_layers) provides the comprehensive global representation required by the decoder. Decoder The autoregressive decoder constructs the UAV restoration plan via stepwise decoding. At step i, a Gated Recurrent Unit (GRU) maintains the sequential hidden state: i,igru=GRU(i0,),i=0GRU(i−1,i−1gru),i>0,o_i,H_i^gru= casesGRU(H_i^0,0),&i=0\\ GRU(H_i^i-1,H_i-1^gru),&i>0, cases (13) where igruH_i^gru is the GRU hidden state, io_i is the GRU output, i0H_i^0 is the initial input, and i−1H_i^i-1 is the current input feature at step i. This recurrent mechanism ensures that the decision at the current step is informed by the complete historical trajectory constructed so far. Conditioned on the encoder output and the current dynamic state, an attention mechanism computes the decoder context representation at i as follows: idec=Attention((nlayers),dynamici,i).h^dec_i=Attention (H^(n_layers),dynamic_i,o_i ). (14) Following the computation of the contextual representation, the pointer decoder evaluates the selection probability for each candidate node. The raw selection score μij _ij for candidate node j is computed as: μij=ζ⋅tanh((qidec)⊤(kjenc)),j∉πi−∞,otherwise, _ij= casesζ· ((W_qh^dec_i) (W_kh^enc_j) ),&j∉ _i\\ -∞,&otherwise, cases (15) where qW_q and kW_k are query and key weight matrices, and ζ is a scaling factor (typically set to 10). If πi _i is the set of already visited nodes, its score is set to −∞-∞, which ensures that the node cannot be selected again. The node-selection probability is then obtained by pθ(ji=j∣s<i)=Softmax(μi)j.p_θ(j_i=j s_<i)=Softmax( _i)_j. (16) After the node jij_i is selected, an area-allocation head predicts the restoration level from a feasible set Di(ji)D_i(j_i). This set is dynamically constrained by the remaining payload, energy, and residual restorable areas: νiδ(ji)=δ⊤ReLU(δ[idec;jienc;dynamici]),δ∈Di(ji), _iδ^(j_i)=w_δ ReLU (W_δ [h^dec_i;h^enc_j_i;dynamic_i ] ), δ∈ D_i(j_i), (17) pθ(δi=δ∣ji,s<i)=Softmax(νi(ji))δ.p_θ( _i=δ j_i,s_<i)=Softmax ( _i^(j_i) )_δ. (18) Accordingly, the joint action distribution factorizes as pθ(ai∣s<i)=pθ(ji∣s<i)pθ(δi∣ji,s<i).p_θ(a_i s_<i)=p_θ(j_i s_<i)\,p_θ( _i j_i,s_<i). (19) During forward propagation, the model Mθ(x)M_θ(x) takes the input problem instance x and outputs both the trajectory cost C and the aggregated log-probability: (C,logp)←Mθ(x),(C, p)← M_θ(x), (20) where the aggregated log-probability across all L steps is given by: logp=∑i=1L[logpθ(ji∣s<i)+logpθ(δi∣ji,s<i)]. p= _i=1^L [ p_θ(j_i s_<i)+ p_θ( _i j_i,s_<i) ]. (21) This trajectory-level output facilitates the subsequent gradient-based optimization of the policy parameters θ. IV-B3 Model Training and Optimization To train the single-UAV restoration model, we employ an actor-critic algorithm [30] combined with the Adam optimizer [16] to update the policy network πθ _θ. As an extension of the REINFORCE algorithm [35], this dual-network approach effectively stabilizes training and accelerates convergence. Actor Network Optimization The actor network approximates the policy probability function p(π|s)p(π|s) for the restoration task. Let θ denote its trainable parameters, the optimization objective is to maximize the expected return under the policy induced by θ: J(θ∣s)=π∼pθ(⋅∣s)[R(π∣s)].J(θ s)=E_π p_θ(· s)[R(π s)]. (22) Following [35], the policy gradient is derived using the advantage function A(π∣s)A(π s): ∇θJ(θ∣s) _θJ(θ s) =π∼pθ(⋅∣s)[(R(π∣s)−b(s))∇θlnpθ(π∣s)] =E_π p_θ(· s) [ (R(π s)-b(s) ) _θ p_θ(π s) ] (23) =π∼pθ(⋅∣s)[A(π∣s)∇θlnpθ(π∣s)], =E_π p_θ(· s) [A(π s) _θ p_θ(π s) ], where R(π∣s)R(π s) denotes the reward obtained during the grassland restoration process, and pθ(π∣s)p_θ(π s) is the probability that the actor network selects a restoration region and its corresponding restoration area size at each autoregressive decision step. The term b(s)b(s) is the baseline value estimate provided by the critic network. The advantage function A(π∣s)=R(π∣s)−b(s)A(π s)=R(π s)-b(s) evaluates the relative performance of policy π against the expected value of state s. A positive advantage reinforces the selected actions via the term ∇θlnpθ(π∣s) _θ p_θ(π s), whereas a negative advantage reduces suboptimal actions. For a training batch of size B with L decision steps, the gradient is approximated as: ∇θJ(θ∣s)≈1B∑i=1B∑j=1L[A(πi,j∣si,j)∇θlnpθ(πi,j∣si,j)]. _θJ(θ s)≈ 1B _i=1^B _j=1^L [A( _i,j s_i,j) _θ p_θ( _i,j s_i,j) ]. (24) Critic Network Optimization Let θc _c denote the parameters of the critic network, the objective of the critic is to minimize the discrepancy between the predicted baseline value bθc(Vi)b_ _c(V_i) and the experienced return R(π∣Vi)R(π V_i). Accordingly, the loss function can be approximated by the mean squared error (MSE): ℒ(θc∣s)≈1B∑i=1B‖bθc(Vi)−R(π∣Vi)‖22.L( _c s)≈ 1B _i=1^B \|b_ _c(V_i)-R(π V_i) \|_2^2. (25) Unified Training Procedure 1 Input : Initial model parameters θ, training epochs NepochN_epoch, batch size B, learning rate α, gradient clipping threshold gmaxg_ , validation set size |Dval||D_val|. Output : Optimal policy parameters θ∗θ^*, trained actor network ℳθ∗M_θ^*, critic baseline network ℬ∗B^*. 2 // Initialization ℳθ←InitializeActor(θ)M_θ (θ); ℬmodel←InitializeCritic()B_model (); 3 ←Adam(ℳθ,ℬmodel,α)O (\M_θ,B_model\,α) ; // Init Optimizer Dval←GenerateInstances(|Dval|)D_val (|D_val|) ; // Validation Set Rbest←−∞R_best←-∞; θ∗←θ^*←θ; 4 5 // Main Training Loop for epoch ←1← 1 to NepochN_epoch do 6 Dtrain←GenerateInstances()D_train () ; // Training Set 7 foreach Batch ∈Split(Dtrain,B) (D_train,B) do 8 (C,logp)←ℳθ(Batch)(C, p) _θ(Batch) ; // Forward pass bv←ℬmodel(Batch)b_v _model(Batch) ; // Baseline estimate 9 // Compute Losses (Eq. 26) ℒR←1B∑i=1B(Ci−bv,i)⋅logpiL_R← 1B _i=1^B(C_i-b_v,i)· p_i; 10 ℒb←BaselineLoss(bv,C)L_b (b_v,C); 11 ℒ←ℒR+ℒbL _R+L_b ; // Total Loss 12 // Backpropagation ∇θ←∇θℒ _θ← _θL; 13 ‖∇θ‖2←min(‖∇θ‖2,gmax)\| _θ\|_2← (\| _θ\|_2,g_ ) ; // Gradient Clipping θ←(θ,∇θ)θ (θ, _θ) ; // Update Parameters end foreach 14 15 // Validation Phase ℳθ.SetMode(GreedyDecoding)M_θ.SetMode(GreedyDecoding); 16 Cval←EvaluatePolicy(ℳθ,Dval)C_val (M_θ,D_val); 17 r¯←−1|Dval|∑i=1|Dval|Cval,i r←- 1|D_val| _i=1^|D_val|C_val,i ; // Avg Reward 18 if r¯>Rbest r>R_best then 19 θ∗←θ^*←θ; Rbest←r¯R_best← r ; // Save Best Model end if 20 end for 21 return θ∗,ℳθ∗,ℬ∗θ^*,M_θ^*,B^*; 22 Algorithm 1 BFPRL Training for Single-UAV Trajectory Planning and Restoration Area Allocation As detailed in Algorithm 1, the actor and critic networks are jointly optimized using Adam with a scheduled learning rate α. The total loss ℒ=ℒR+ℒbL=L_R+L_b combines the critic’s baseline loss ℒbL_b and the actor’s policy gradient loss: ℒR=1B∑i=1B(Ci−bv,i)⋅logpi,L_R= 1B _i=1^B(C_i-b_v,i)· p_i, (26) where CiC_i is the actual cost (negative reward) and bv,ib_v,i is the critic-estimated baseline. To prevent gradient explosion, gradients are clipped via ‖∇θ‖2←min(‖∇θ‖2,gmax)\| _θ\|_2← (\| _θ\|_2,g_ ) prior to parameter updates. After each epoch, the model is evaluated on a validation set DvalD_val via greedy decoding. We retain the optimal parameters θ∗θ^* that maximize the average validation reward: θ∗=argmaxθ(−1|Dval|∑i=1|Dval|Cval,i)θ^*= _θ (- 1|D_val| _i=1^|D_val|C_val,i ) (27) This architecture ensures stable gradient estimation and rapid convergence to an optimal single-UAV restoration policy. IV-C Knowledge-Guided Multi-UAV Collaboration Framework To address the challenges of large-scale degraded grassland restoration, we propose KC-BFPRL, a knowledge-guided multi-UAV collaboration framework. Built upon a bilevel formerpointer reinforcement learning approach, KC-BFPRL adopts the hierarchical collaborative architecture illustrated in Fig. 2. Following an “external collaboration, internal intelligence” paradigm, a ground control center (GCC) manages global cooperative scheduling and dynamic map updates. Concurrently, individual UAVs utilize the Transformer-Pointer based BFPRL model for local sequential decision-making to optimize flight trajectories and restoration allocations. This architecture strategically couples global scheduling data with local perception, effectively balancing computational efficiency and global restoration performance. IV-C1 Collaboration Scheduling Mechanism To facilitate collaboration, the global target set Va=V∖v0V_a=V \v_0\ is partitioned into m mutually exclusive subsets, one for each UAV u∈U=1,…,mu∈ U=\1,…,m\, satisfying m≪nm n. Each subset Vu=vu1,vu2,…,vuhV_u=\v^1_u,v^2_u,…,v^h_u\ assigns h regions to UAV u, where each region is defined by a tuple vui=(idui,pui,lui,aui)v_u^i=(id_u^i,p_u^i,l_u^i,a_u^i) representing its identifier, spatial coordinates, degradation level, and restorable area size, respectively. The operational state of UAV u is defined as Su=(u,Eurem,Vuvis,Aurep)S_u= (p_u,E_u^rem,V^vis_u,A^rep_u ), where up_u is the current position, EuremE^rem_u is the remaining energy, VuvisV^vis_u tracks the sequence of visited regions, and AurepA^rep_u is the total accumulated restored areas. To synchronize interactions between the UAVs and the GCC, a semaphore variable SPuSP_u is introduced to estimate the time cost for the next task: SPu=d(u,vnexti)v+ρ⋅airepr,SP_u= d(p_u,v_next^i)v+ρ· a^rep_ir, (28) where d(⋅)d(·) is the Euclidean distance, v denotes the UAV’s flight speed, r represents the restoration rate (areas per unit time), ρ is a weighting factor, and airepa_i^rep is the target region size at the next node vinextv^next_i. IV-C2 Multi-UAV Information Sharing Strategy To optimize task distribution via dynamic information exchange, a matching cost function K(u,vi)K(u,v_i) evaluates the suitability of assigning region viv_i to UAV u. Balancing spatial proximity, task urgency, and workload, it is formulated as: K(u,vi)=α⋅d(u,i)+β⋅1li+ϵ⋅1ai,K(u,v_i)=α· d(p_u,p_i)+β· 1l_i+ε· 1a_i, (29) where α, β, ϵε are tunable weight coefficients for distance, degradation level, and restoration area size, respectively. A lower K value indicates a higher assignment priority. To achieve global optimization, this cooperative scheduling process iteratively executes the following five key steps. Step 1: Initial Partitioning. The global target region VaV_a is partitioned via spatial clustering (e.g., K-means), allocating an initial subset Vu0V_u^0 to each UAV u. Step 2: Baseline Trajectory Optimization. Based on Vu0V_u^0, each UAV computes an energy-feasible trajectory to maximize its baseline restoration gain Au1A_u1: (Pu1,Au1)=argmax∑vi∈P∈(Mu)airep,(P_u1,A_u1)= _P (M_u) _v_i∈ Pa_i^rep, (30) where (Mu)P(M_u) denotes the set of energy-feasible paths and airepa_i^rep is the number of restoration areas in region viv_i. Step 3: Global Task Reallocation. The GCC tentatively reassigns tasks based on the real-time fleet states, allocating each region to the UAV that minimizes the matching cost K: Vutmp=vi∈Va|u=argminu′∈UK(u′,vi).V_u^tmp= \v_i∈ V_a |u= _u ∈ UK(u ,v_i) \. (31) Step 4: Candidate Trajectory Generation. Based on the tentative partition VutmpV_u^tmp, each UAV replans its trajectory to determine the potential restoration gain, denoted as Au2A_u2. Step 5: Global Decision. The system evaluates the efficacy of the reallocation by comparing aggregate restoration gains of baseline (Step 2) and candidate (Step 4) plans. The tentative allocation is adopted only if it improves global performance: Vu=Vutmp,if ∑u∈UAu2≥∑u∈UAu1Vu,otherwise.V_u= casesV_u^tmp,&if _u∈ UA_u2≥ _u∈ UA_u1\\ V_u,&otherwise. cases (32) This iterative information-sharing mechanism facilitates dynamic task reallocation, maximizing the total restored area under strict energy constraints. IV-C3 Decision-Making and Execution UAVs execute restoration decisions sequentially based on calculated semaphore priorities. The UAV with the minimum semaphore is selected first: u∗=argminu∈USPu.u^*= _u∈ USP_u. (33) Upon completing its current task at vicurrv^curr_i, UAV u∗u^* selects the next target region within its assigned subset by minimizing the local matching cost: vinext=argminvi∈Vu∗K(u∗,vi).v^next_i= _v_i∈ V_u^*K(u^*,v_i). (34) The state of u∗u^* is then updated as: Su∗=(vinext,Eu∗rem−Eflight−Erepair,OPENVu∗vis∪vicurr,Au∗rep+airep), splitS_u^*=& (v_i^next,\ E_u^*^rem-E_flight-E_repair, .\\ & .V_u^*^vis∪\v_i^curr\,\ A_u^*^rep+a_i^rep ), split (35) The process (summarized in Algorithm 2) repeats until all global tasks are completed or energy constraints force a return to the depot. 1 Input : Global parameter set P, initial restoration map partitions Vu0u=1m\V_u^0\_u=1^m, UAV states Su0u=1m\S_u^0\_u=1^m, pre-trained RL model parameters ℳθM_θ. Output : UAV trajectories Puu=1m\P_u\_u=1^m, cumulative restored areas Auu=1m\A_u\_u=1^m, remaining energy Euu=1m\E_u\_u=1^m. 2 // Initialization Initialize Vu←Vu0,Pu←∅,Su←Su0,∀u∈UV_u← V_u^0,P_u← ,S_u← S_u^0,∀ u∈ U; ℳglobal←⋃VuM_global← V_u; 3 4 // Main Scheduling Loop while ℳglobal≠∅M_global≠ do 5 // Phase I: Distributed Planning & Global Reallocation Compute baseline: (Eu(1),Au(1))←PlanTrajectory(ℳθ,Vu)∀u∈U(E_u^(1),A_u^(1)) (M_θ,V_u)\ ∀ u∈ U; 6 Update global map: ℳglobal←UpdateAllocation(Vu,Au(1))M_global (\V_u,A_u^(1)\); 7 Compute candidate: Vutmp←GetTasks(ℳglobal)V_u^tmp (M_global); (Eu(2),Au(2))←PlanTrajectory(ℳθ,Vutmp)∀u(E_u^(2),A_u^(2)) (M_θ,V_u^tmp)\ ∀ u; 8 9 if ∑Au(2)≥∑Au(1)Σ A_u^(2)≥Σ A_u^(1) then Update maps: Vu←Vutmp\V_u\←\V_u^tmp\; 10 11 // Phase I: Semaphore-based Execution Select active UAV: u∗←argminuSPuu^*← _uSP_u; Select target: v∗←argminv∈Vu∗K(u∗,v)v^*← _v∈ V_u^*K(u^*,v); 12 13 if v∗≠NULLv^* then 14 ExecuteRestoration(u∗,v∗u^*,v^*); Vu∗←Vu∗∖v∗V_u^*← V_u^* \v^*\; 15 Update state Su∗S_u^* (Energy, Position, Area) per Eq. (35); 16 Update semaphore SPu∗SP_u^* based on new state; 17 end if 18 end while 19 20 // Termination foreach UAV u∈Uu∈ U do ReturnToDepot(u); 21 ; 22 return uu=1U,Auu=1U,Euu=1U\P_u\_u=1^U,\A_u\_u=1^U,\E_u\_u=1^U. Algorithm 2 Knowledge-Guided Multi-UAV Collaboration Scheduling Algorithm for RAMP V Experimental Evaluation and Analysis This section presents comprehensive experiments validating the effectiveness and generalization of the proposed KC-BFPRL framework for the RAMP. V-A Experiment Settings All experiments were conducted on a workstation equipped with an AMD Ryzen Threadripper 3970X CPU, NVIDIA RTX 3090 Ti GPU, and 192 GB RAM, utilizing Python 3.9 and PyTorch 1.12.1 (CUDA 11.3) on Ubuntu 22.04. V-A1 Instance Generation To evaluate the KC-BFPRL framework’s generalization across diverse spatial scales and constraints, we generated 10800 test instances spanning 108 parameter configurations. These configurations combine three fleet sizes M∈4,6,8M∈\4,6,8\, six problem scales N∈60,80,…,160N∈\60,80,…,160\ target regions, spanning areas from 500×500m2500× 500\,m^2 to 1000×1000m21000× 1000\,m^2, and six workload levels, where the number of unit circles δ per degraded region is uniformly sampled from the set 10,15,…,35\10,15,…,35\ in increments of 55 across all scenarios. Configurations are denoted as UM-RN, e.g., U4-R60. For each instance, the BS is fixed at (0,0)(0,0). Target coordinates are uniformly sampled in [0,1]2[0,1]^2 and scaled to the specific spatial area, with degradation levels li∼U(0,1)l_i U(0,1). To accommodate larger domains, the initial UAV energy EmaxE_max is proportionally scaled from 1.00×1071.00× 10^7 J to 7.59×1077.59× 10^7 J. To ensure statistical reliability, we evaluate 100 independent random instances per configuration. V-A2 Compared Algorithms We evaluate KC-BFPRL against three state-of-the-art baselines: (1) CHAPBILM [13]: A specialized heuristic for UAV grassland restoration combining population-based incremental learning with a maximum-residual-energy local search. (2) MAPDP [43]: A multi-agent routing algorithm utilizing paired context embedding and a collaborative advantage actor-critic (A2C) approach. (3) CAMP [10]: An attention-based multi-agent model featuring inter-agent communication and a centralized-training distributed-execution (CTDE) framework. Since the original MAPDP and CAMP were not originally designed for the RAMP, we adapted their input layers and decoder masking mechanisms to accommodate our problem constraints while preserving their core architectures. Detailed parameter configurations are provided in Appendix A. V-A3 Parameter Settings Following [7], the UAV operational parameters are set as follows: the frame mass is initialized to M=1.5M=1.5 kg, gravitational acceleration to g=9.8g=9.8 m/s2, air density ρ=1.024ρ=1.024 kg/m3, rotor disc area ς=0.2 =0.2 m2, and the number of rotors to h=6h=6. Regarding the energy consumption dynamics in Section I-B, the seeding energy coefficient is set to η=5×104η=5× 10^4, the environmental parameter to γ=1.5γ=1.5, and the aerial photography energy consumption to eap=2×104e^ap=2× 10^4 J. V-B Performance Evaluation Following established evaluation protocols [13, 20], we evaluate KC-BFPRL framework across six key performance metrics: average objective value (Obj.), average optimality gap (Gap), average energy efficiency (ratio of task energy to total energy), average total trajectory length, average total restoration areas (Areas), and average computation time (Time). All data are rounded to the shown precision (with Areas rounded down), and best results are in bold. V-B1 Scalability and Performance Comparison To validate its adaptability and scalability, the KC-BFPRL framework is evaluated across diverse problem scales. The task complexity is systematically tested using target node counts ranging from n=60n=60 to n=160n=160 in increments of 20, spanning scenarios from small-to-medium scenarios to medium-to-large scales. For each scale, performance is further assessed under varying fleet configurations of 4, 6, and 8 UAVs. First, we evaluate training stability of the model in Fig. 8 of Appendix B, which reveals a trade-off between efficiency and performance. A distinct trade-off between training efficiency and restoration performance is evident. While MAPDP demonstrates the fastest convergence and CAMP converges moderately with limited improvement in restoration area, KC-BFPRL requires the longest training period. However, this extended training enables KC-BFPRL to achieve a significantly superior restoration outcome. As confirmed by the final convergence states in Fig. 4, the KC-BFPRL models exhibit robust stability, indicating the successful acquisition of effective, high-performance restoration policies. Subsequently, to provide intuitive insight into the model’s decision-making, Fig. 9 of Appendix B visualizes the optimal trajectories for random instances with seed 1234. These results confirm that KC-BFPRL generates logical, efficient paths across various scales, effectively coordinating the multi-UAV fleet to maximize restoration coverage. array[]l [width]Training-Loss-Comparison.pdf\\ array Fig. 4: Comparison of training convergence curves for the three learning-based algorithms in scenario U4-R60. V-B2 Comparison With State-of-the-Art Methods (a) U4-R60 (b) U6-R80 (c) U8-R100 (d) U4-R60 (e) U6-R80 (f) U8-R100 (g) U4-R60 (h) U6-R80 (i) U8-R100 (j) U4-R60 (k) U6-R80 (l) U8-R100 Fig. 5: Optimal trajectories generated by the four algorithms for 4, 6, and 8 UAVs across three small-to-medium-scale scenarios: (a)-(c) CAMP, (d)-(f) CHAPBILM, (g)-(i) MAPDP, and (j)-(l) KC-BFPRL. TABLE I: Comparison of Average Performance Across Four Algorithms for 4-UAV Collaborative Grassland Restoration in All Scenarios. Areas Pending Restoration Metric Algorithms CAMP CHAPBILM MAPDP KC-BFPRL 60 Obj. 675.60 798.51 862.34 945.80 Areas 612 742 789 863 Time (s) 11.21 892.49 45.79 15.81 Gap (%) 28.59 15.57 8.83 0.00 80 Obj. 878.00 1065.42 1125.51 1248.27 Areas 798 968 1023 1134 Time (s) 53.24 1287.65 62.44 16.85 Gap (%) 29.65 14.64 9.83 0.00 100 Obj. 1048.18 1306.40 1628.46 1585.39 Areas 952 1187 1478 1442 Time (s) 67.82 1654.37 78.98 21.55 Gap (%) 35.65 19.77 0.00 2.64 120 Obj. 1272.43 1566.26 1688.11 1908.76 Areas 1156 1423 1534 1735 Time (s) 25.89 2021.71 94.59 28.90 Gap (%) 33.33 17.95 11.55 0.00 140 Obj. 1526.65 1858.53 2006.35 2296.12 Areas 1387 1689 1823 2087 Time (s) 98.52 2456.27 111.81 31.22 Gap (%) 33.52 19.05 12.63 0.00 160 Obj. 1798.08 2186.17 2756.82 2698.50 Areas 1634 1987 2504 2453 Time (s) 114.22 2874.56 128.93 36.70 Gap (%) 34.75 20.69 0.00 2.12 TABLE I: Comparison of Average Performance Across Four Algorithms for 6-UAV Collaborative Grassland Restoration in All Scenarios. Areas Pending Restoration Metric Algorithms CAMP CHAPBILM MAPDP KC-BFPRL 60 Obj. 1010.59 1198.18 1306.18 1425.49 Areas 918 1089 1187 1295 Time (s) 16.82 1340.78 68.98 22.74 Gap (%) 29.10 15.95 8.37 0.00 80 Obj. 1317.09 1528.79 1898.60 1834.28 Areas 1197 1389 1726 1666 Time (s) 76.57 1789.31 89.21 24.74 Gap (%) 30.63 19.48 0.00 3.39 100 Obj. 1572.18 1907.65 2116.23 2348.22 Areas 1429 1734 1923 2134 Time (s) 95.88 2287.49 112.81 31.24 Gap (%) 33.05 18.77 9.88 0.00 120 Obj. 1908.63 2298.59 2528.63 2858.27 Areas 1735 2089 2298 2598 Time (s) 38.96 2798.62 136.71 41.54 Gap (%) 33.22 19.59 11.53 0.00 140 Obj. 2292.66 2702.57 3007.60 3398.74 Areas 2084 2456 2734 3089 Time (s) 135.88 3356.91 161.33 45.11 Gap (%) 32.56 20.48 11.51 0.00 160 Obj. 2695.76 3154.12 3518.65 3998.35 Areas 2450 2867 3198 3634 Time (s) 54.21 3897.26 186.84 56.82 Gap (%) 32.57 21.12 12.00 0.00 TABLE I: Comparison of Average Performance Across Four Algorithms for 8-UAV Collaborative Grassland Restoration in All Scenarios. Areas Pending Restoration Metric Algorithms CAMP CHAPBILM MAPDP KC-BFPRL 60 Obj. 1347.36 1428.41 1789.27 1718.95 Areas 1224 1298 1625 1562 Time (s) 69.77 1643.51 82.11 22.64 Gap (%) 24.69 20.16 0.00 3.93 80 Obj. 1756.63 1868.25 2076.01 2298.61 Areas 1596 1698 1887 2089 Time (s) 29.87 2156.76 107.32 32.15 Gap (%) 23.58 18.71 9.68 0.00 100 Obj. 2097.20 2298.57 2604.63 2898.23 Areas 1906 2089 2367 2634 Time (s) 114.52 2734.87 134.90 37.52 Gap (%) 27.64 20.69 10.13 0.00 120 Obj. 2544.18 2776.27 3154.13 3518.46 Areas 2312 2523 2867 3198 Time (s) 42.19 3287.69 162.83 48.71 Gap (%) 27.69 21.09 10.35 0.00 140 Obj. 2990.23 3264.11 3726.34 4178.61 Areas 2718 2967 3387 3798 Time (s) 163.24 3865.97 192.11 53.48 Gap (%) 28.45 21.88 10.83 0.00 160 Obj. 3547.14 3835.99 4386.42 4914.08 Areas 3224 3487 3987 4467 Time (s) 65.19 4453.23 222.46 68.54 Gap (%) 27.82 21.93 10.74 0.00 We evaluate our method against CHAPBILM [13], MAPDP [43], and CAMP [10]. The experimental results for fleet sizes of U4, U6, and U8 are presented in Tables I, I, and I, respectively. The optimality gap is defined as the normalized difference between the average objective value of a given method (Obj.) and the best objective value (ObjbestObj_best) across all methods: Gap=Objbest−ObjObjbest×100%Gap= Obj_best-ObjObj_best× 100\%. Tables I to I show that KC-BFPRL consistently outperforms baselines across all fleet sizes and scales, particularly in large-scale scenarios involving 120 to 160 regions, with the minor exception of U4-R160. In the complex U8-R160 instance, it surpasses MAPDP and CAMP by 12.03%12.03\% in objective value and approximately 38.55%38.55\% in total restored areas, respectively. While MAPDP’s optimality gap widens to 10.74%10.74\% as complexity increases, KC-BFPRL maintains a 0.00%0.00\% gap across all 8-UAV instances. Furthermore, it is nearly three times faster than MAPDP, effectively mitigating the “curse of dimensionality”. Its robustness is highlighted by maintaining a mean optimality gap under 2%2\% as the fleet scales. (a) Energy efficiency (b) Number of restoration areas (c) Trajectory length (d) Objective function value Fig. 6: Comparative performance of the four algorithms on four key metrics. Statistical results in presented Fig. 6 and Table V of Appendix B demonstrate the superiority of KC-BFPRL across four key metrics. It achieves the highest average restoration areas of 1109.73, outperforming MAPDP, CHAPBILM, and CAMP by 11.60%11.60\%, 28.09%28.09\%, and 58.52%58.52\%, respectively. It also yields the most efficient flight paths, with an average length of 4068.40, representing distance reductions of 3.69%3.69\%, 12.83%12.83\%, and 13.56%13.56\% compared to the baselines. Moreover, KC-BFPRL exhibits higher median and minimum values along with lower standard deviations in energy efficiency, resulting in a mean value of 0.538 that underscores its exceptional stability. Finally, KC-BFPRL surpasses its closest competitor, MAPDP, by 8.90%8.90\% in overall objective value, validating its practical engineering value for large-scale, complex ecological restoration tasks. V-B3 Generalization and Robustness Study (a) U4-R120 (b) U6-R120 (c) U8-R120 (d) U4-R140 (e) U6-R140 (f) U8-R140 (g) U4-R160 (h) U6-R160 (i) U8-R160 Fig. 7: Visualization of optimal KC-BFPRL trajectories for 4, 6, and 8 UAVs across three medium-to-large-scale scenarios. We evaluate the model’s generalization by testing a fixed policy across varying problem sizes without fine-tuning. Fig. 7 shows that KC-BFPRL maintains high-quality solutions even in mismatched settings with minimal performance loss. Notably, the pre-trained model often outperforms baselines specifically trained for those scenarios. These results confirm the versatility of our knowledge-guided framework and its potential for broad applicability in complex multi-UAV challenges. V-B4 Discussion Based on the comparative performance analysis, we delineate the specific applicability of each algorithm to different operational contexts. KC-BFPRL balances optimality and efficiency, making it ideal for large-scale, complex missions involving over 100 restoration regions that require real-time decision-making. In contrast, MAPDP is suitable for precision-critical tasks where latency is acceptable, while CHAPBILM is best for smaller, training-free offline scenarios with fewer than 80 regions. KC-BFPRL’s success lies in its knowledge-guided paradigm, which prunes the search space using ecological priority, heuristic logic, and hierarchical coordination. This “warm-start” avoids learning from scratch and mitigates the “curse of dimensionality.” Furthermore, by utilizing transparent heuristics, the framework enhances interpretability and provides a systematic guide for selecting algorithms based on mission scale and resource constraints. VI Conclusion This paper addressed the critical challenge of multi-UAV collaborative grassland restoration by formulating the RAMP and proposing the KC-BFPRL framework. By integrating domain expertise, specifically ecological priority and heuristic scheduling logic, with the adaptive capability of DRL, the proposed approach effectively bridges the gap between centralized coordination and decentralized execution. Extensive experiments demonstrates that KC-BFPRL significantly outperforms state-of-the-art heuristic and learning-based baselines, successfully enabling real-time, high-precision decision-making. Essentially, the knowledge-guided paradigm ensures robust scalability and interpretability, maintaining near-optimal performance even as task complexity increases. Future work will explore problem-decomposition for lower latency and conduct real-world trials under dynamic environmental uncertainties to advance autonomous ecological engineering. Appendix A Hyperparameter Configuration for Multi-Agent Reinforcement Learning Methods All multi-agent reinforcement learning (MARL) baselines were implemented using the source code and hyperparameters provided in their original publications. To ensure an equitable comparison, each baseline model was trained for 10000 epochs, consistent with the training protocol established for KC-BFPRL. Because the native architectures of MAPDP and CAMP cannot be directly applied to RAMP, we adapted their input representations and constraint-handling mechanisms to align with our mathematical formulation. Specifically, we explicitly integrated the critical constraints governing spatially heterogeneous degradation levels and dynamic UAV energy limitations into their decision-making processes. The specific parameter settings are shown in the table IV. TABLE IV: Transformer model training parameters Parameter Description Value nlayersn_layers number of Transformer layers 3 dind_in input dimension 4 (x, y, l, R) ded_e embedding dimension 128 C attention scaling factor 10 NepochN_epoch number of epochs 100 B batch size 256 α learning rate (Adam) 3×10−43× 10^-4 gmaxg_max gradient clipping threshold 0.5 DvalD_val validation set size 256 TepochsT_epochs number of epochs (training) 10000 αp _p penalty term weight 10.0 αr _r recovery term weight 1.0 Appendix B Extended Experimental Results Fig. 8 illustrates the accumulation of total restored areas across training epochs, comparing the learning curves of the three learning-based approaches under the U4-R60 configuration. array[]l [width]Total-Area-Training-Episodes.pdf\\ array Fig. 8: Evolution of total restoration areas across training epochs for the three learning-based algorithms in scenario U4-R60. Fig. 9 visualizes the optimal trajectories generated by the proposed KC-BFPRL framework. The rows correspond to varying fleet sizes (4, 6, and 8 UAVs), while the columns denote three distinct medium-to-large-scale scenarios. This layout provides a clear qualitative assessment of the algorithm’s path-planning performance across different mission scales and spatial constraints. (a) U4-R60 (b) U6-R60 (c) U8-R60 (d) U4-R80 (e) U6-R80 (f) U8-R80 (g) U4-A100 (h) U6-R100 (i) U8-R100 Fig. 9: Visualization of optimal KC-BFPRL trajectories for 4, 6, and 8 UAVs across three small-to-medium-scale scenarios. Table V presents a comprehensive comparison of the average performance metrics achieved by the four algorithms. The reported values are averaged across all test instances, providing a consolidated evaluation of each algorithm’s relative effectiveness across the primary criteria. TABLE V: Comparison of average performance metrics across the four algorithms. Metric Algorithm Mean Std Median Min Max Q75 Restoration Areas CAMP 700 359 625 203 2216 838 CHAPBILM 866 449 725 292 2263 1067 MAPDP 994 487 872 262 2510 1232 KC-BFPRL 1109 544 971 322 2941 1395 Energy Efficiency CAMP 0.457 0.043 0.437 0.411 0.570 0.477 CHAPBILM 0.494 0.052 0.469 0.431 0.582 0.547 MAPDP 0.537 0.047 0.555 0.441 0.598 0.575 KC-BFPRL 0.538 0.035 0.548 0.453 0.597 0.566 Trajectory Length CAMP 4706.47 2677.40 4052.19 853.68 13725.89 6173.27 CHAPBILM 4667.37 2716.91 3669.33 859.86 13507.69 6631.64 MAPDP 4224.46 2561.48 3433.10 949.60 13821.66 5639.78 KC-BFPRL 4068.40 2379.57 3419.49 918.53 12655.74 5509.92 Objective Value CAMP 1024.52 543.90 893.03 233.07 3239.36 1204.10 CHAPBILM 1313.71 712.87 1121.74 334.85 4000.82 1620.81 MAPDP 1515.66 829.19 1296.49 230.85 4568.51 1936.94 KC-BFPRL 1650.54 865.87 1437.29 442.23 5009.81 2073.11 Acknowledgment This work was supported in part by the National Natural Science Foundation of China under Grant 62233003 and 62272210, and in part by the Natural Science Foundation of Gansu Province under Grant 24JRRA430. We thank Mr. B. Liu for his assistance and the anonymous reviewers for their constructive feedback which improved this manuscript. References [1] K. Anderson and K. J. Gaston (2013) Lightweight unmanned aerial vehicles will revolutionize spatial ecology. Frontiers in Ecology and the Environment 11 (3), p. 138–146. Cited by: §I-A. [2] I. Bello, H. Pham, Q. V. Le, M. Norouzi, and S. Bengio (2017) Neural combinatorial optimization with reinforcement learning. In International Conference on Learning Representations (ICLR), Workshop Track, p. 1–5. Cited by: §IV-B1. [3] N. Bono Rossello, R. F. Carpio, A. Gasparri, and E. Garone (2022) Information-driven path planning for UAV with limited autonomy in large-scale field monitoring. IEEE Transactions on Automation Science and Engineering 19 (3), p. 2450–2460. Cited by: §I-A. [4] X. Bresson and T. Laurent (2021) The transformer network for the traveling salesman problem. arXiv preprint arXiv:2103.03012. Cited by: §IV-B2. [5] V. Bui and T. Mai (2023) Imitation improvement learning for large-scale capacitated vehicle routing problems. In Proceedings of the International Conference on Automated Planning and Scheduling, Vol. 33, p. 551–559. Cited by: §I-D. [6] J. Chen, F. Ling, Y. Zhang, T. You, Y. Liu, and X. Du (2022) Coverage path planning of heterogeneous unmanned aerial vehicles based on ant colony system. Swarm and Evolutionary Computation 69, p. 101005. Cited by: §I-B. [7] K. Dorling, J. Heinrichs, G. G. Messier, and S. Magierowski (2017) Vehicle routing problems for drone delivery. IEEE Transactions on Systems, Man, and Cybernetics: Systems 47 (1), p. 70–85. Cited by: §I-B3, §I-B, §V-A3. [8] G. D. Gann, T. McDonald, B. Walder, J. Aronson, C. R. Nelson, J. Jonson, J. G. Hallett, C. Eisenberg, M. R. Guariguata, J. Liu, et al. (2019) International principles and standards for the practice of ecological restoration. Restoration Ecology 27 (S1), p. S1–S46. Cited by: §I-A. [9] X. Gao, J. Si, Y. Wen, M. Li, and H. H. Huang (2020) Knowledge-guided reinforcement learning control for robotic lower limb prosthesis. In 2020 IEEE International Conference on Robotics and Automation (ICRA), p. 754–760. Cited by: §I-D. [10] C. Hua, F. Berto, J. Son, S. Kang, C. Kwon, and J. Park (2025) CAMP: collaborative attention model with profiles for vehicle routing problems. In Proceedings of the 24th International Conference on Autonomous Agents and Multiagent Systems, p. 1015–1024. Cited by: §I-C, §V-A2, §V-B2. [11] X. Huo and M. Liu (2022) Two-facet scalable cooperative optimization of multi-agent systems in the networked environment. IEEE Transactions on Control Systems Technology 30 (6), p. 2317–2332. Cited by: §I. [12] D. Jiao, Z. Chen, X. Wang, J. Shi, S. Liu, and S. Yan (2026) OD-DEAL: dynamic expert-guided adversarial learning with online decomposition for scalable capacitated vehicle routing. arXiv preprint arXiv:2602.00488. Cited by: §I-C. [13] D. Jiao, L. Wang, P. Yang, W. Yang, Y. Peng, Z. Shang, and F. Ren (2024) Unmanned Aerial Vehicle-enabled grassland restoration with energy-sensitive of trajectory design and restoration areas allocation via a cooperative memetic algorithm. Engineering Applications of Artificial Intelligence 133, p. 108084. Cited by: §I, §I, §I-A, §I-B, §I-A, §I-B, §IV-A, §V-A2, §V-B2, §V-B. [14] M. Jones, S. Djahel, and K. Welsh (2023) Path-planning for unmanned aerial vehicles with environment complexity considerations: a survey. ACM Computing Surveys 55 (11), p. 1–39. Cited by: §I-B. [15] J. Kim, S. Kim, C. Ju, and H. I. Son (2019) Unmanned aerial vehicles in agriculture: a review of perspective of platform, control, and applications. IEEE Access 7, p. 105100–105115. Cited by: §I-A. [16] D. P. Kingma and J. Ba (2015) Adam: a method for stochastic optimization. In Proceedings of the 3rd International Conference on Learning Representations (ICLR), San Diego, CA. Cited by: §IV-B3. [17] W. Kool, H. van Hoof, and M. Welling (2019) Attention, learn to solve routing problems!. In International Conference on Learning Representations, Cited by: §I-C. [18] N. Lahrichi, T. G. Crainic, M. Gendreau, W. Rei, G. C. Crişan, and T. Vidal (2015) An integrative cooperative search framework for multi-decision-attribute combinatorial optimization: application to the MDPVRP. European Journal of Operational Research 246 (2), p. 400–412. Cited by: §I. [19] J. Liang, J. Zhao, C. Wang, X. Yang, K. Yue, and W. Li (2025) Enhancing the robustness of UAV search path planning based on deep reinforcement learning for complex disaster scenarios. IEEE Transactions on Vehicular Technology 75 (1), p. 392–404. Cited by: §I. [20] X. Mao, G. Wu, M. Fan, Z. Cao, and W. Pedrycz (2024) DL-DRL: a double-level deep reinforcement learning approach for large-scale task scheduling of multi-UAV. IEEE Transactions on Automation Science and Engineering 22, p. 1028–1044. Cited by: §V-B. [21] R. I. Mukhamediev, K. Yakunin, M. Aubakirov, I. Assanov, Y. Kuchin, A. Symagulov, V. Levashenko, E. Zaitseva, D. Sokolov, and Y. Amirgaliyev (2023) Coverage path planning optimization of heterogeneous UAVs group for precision agriculture. IEEE Access 11, p. 5789–5803. Cited by: §I-A. [22] L. Nepi, G. Quattrini, S. Pesaresi, A. Mancini, and R. Pierdicca (2025) AI-based estimation of forest plant community composition from UAV imagery. Ecological Informatics, p. 103199. Cited by: §I-A. [23] L. P. Osco, J. M. Junior, A. P. M. Ramos, L. A. de Castro Jorge, S. N. Fatholahi, J. de Andrade Silva, E. T. Matsubara, H. Pistori, W. N. Gonçalves, and J. Li (2021) A review on deep learning in UAV remote sensing. International Journal of Applied Earth Observation and Geoinformation 102, p. 102456. Cited by: §I-A. [24] P. Radoglou-Grammatikis, P. Sarigiannidis, T. Lagkas, and I. Moscholios (2020) A compilation of UAV applications for precision agriculture. Computer Networks 172, p. 107148. Cited by: §I-A. [25] T. Rashid, M. Samvelyan, C. S. De Witt, G. Farquhar, J. Foerster, and S. Whiteson (2020) Monotonic value function factorisation for deep multi-agent reinforcement learning. Journal of Machine Learning Research 21 (178), p. 1–51. Cited by: §I-C. [26] Y. Rizk, M. Awad, and E. W. Tunstel (2019) Cooperative heterogeneous multi-robot systems: a survey. ACM Computing Surveys (CSUR) 52 (2), p. 1–31. Cited by: §I. [27] J. Roy, P. Barde, F. Harvey, D. Nowrouzezahrai, and C. Pal (2020) Promoting coordination through policy regularization in multi-agent deep reinforcement learning. Advances in Neural Information Processing Systems 33, p. 15774–15785. Cited by: §I. [28] R. Shahbazian, A. Ciacco, G. Macrina, and F. Guerriero (2025) Knowledge-guided hybrid deep reinforcement learning for the dynamic multi-depot electric vehicle routing problem. Computers & Operations Research, p. 107217. Cited by: §I-D. [29] L. Sommer, B. C. Campos, S. Harvolk-Schöning, T. W. Donath, T. Kleinebecker, and Y. P. Klinger (2023) Grassland restoration with plant material transfer–bridging the knowledge gap between science and practice. Global Ecology and Conservation 47, p. e02638. Cited by: §I. [30] R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour (1999) Policy gradient methods for reinforcement learning with function approximation. Advances in Neural Information Processing Systems 12. Cited by: §IV-B3. [31] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin (2017) Attention is all you need. Advances in Neural Information Processing Systems 30. Cited by: §IV-B2. [32] F. J. C. Verdù, L. Castelli, and L. Bortolussi (2025) Scaling combinatorial optimization neural improvement heuristics with online search and adaptation. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 39, p. 27135–27143. Cited by: §I. [33] O. Vinyals, M. Fortunato, and N. Jaitly (2015) Pointer networks. Advances in Neural Information Processing Systems 28. Cited by: §I-C, §IV-B2. [34] B. G. Waring (2024) Grand challenges in ecosystem restoration. Vol. 11, Frontiers Media SA. Cited by: §I. [35] R. J. Williams (1992) Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine Learning 8 (3), p. 229–256. Cited by: §IV-B3, §IV-B3. [36] X. Wu, D. Wang, L. Wen, Y. Xiao, C. Wu, Y. Wu, C. Yu, D. L. Maskell, and Y. Zhou (2024) Neural combinatorial optimization algorithms for solving vehicle routing problems: a comprehensive survey with perspectives. arXiv preprint arXiv:2406.00415. Cited by: §I-C. [37] J. Xie and J. Chen (2022) Multiregional coverage path planning for multiple energy constrained UAVs. IEEE Transactions on Intelligent Transportation Systems 23 (10), p. 17366–17381. Cited by: §I-B. [38] H. P. Xu, J. Zhang, X. P. Pang, Q. Wang, W. N. Zhang, J. Wang, and Z. G. Guo (2019) Responses of plant productivity and soil nutrient concentrations to different alpine grassland degradation levels. Environmental Monitoring and Assessment 191 (11), p. 678. Cited by: §I-A. [39] R. Xu, Z. Huang, C. Wang, and H. Yan (2025) Evolving collaborative differential evolution for dynamic multi-objective UAV path planning. IEEE Transactions on Vehicular Technology 75 (5), p. 7456–7468. Cited by: §I-B. [40] Z. Yang, Z. Xiao, Z. Han, and X. Xia (2025) Joint path planning and transmission scheduling for multi-UAV data collection. IEEE Transactions on Vehicular Technology 75 (4), p. 6600–6615. Cited by: §I. [41] C. Zhang, J. Guo, F. Wang, B. Chen, C. Fan, L. Yu, and Z. Wang (2025) A dynamic parameters genetic algorithm for collaborative strike task allocation of unmanned aerial vehicle clusters towards heterogeneous targets. Applied Soft Computing 175, p. 113075. Cited by: §I-B. [42] H. Zhang, Q. Li, and X. Yao (2024) Knowledge-guided optimization for complex vehicle routing with 3D loading constraints. In International Conference on Parallel Problem Solving from Nature, p. 133–148. Cited by: §I. [43] Z. Zong, M. Zheng, Y. Li, and D. Jin (2022) MAPDP: cooperative multi-agent reinforcement learning to solve pickup and delivery problems. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 36, p. 9980–9988. Cited by: §I-C, §V-A2, §V-B2.