Paper deep dive
Multi-AUV Ad-hoc network-based Target Tracking: A Value Gradient Guidance Multi-Agent Diffusion Reinforcement Learning Approach
Jiaao Ma, Chuan Lin, Guangjie Han, Shengchao Zhu, Qian Zhu, Ying Liu, Zhenyu Wang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/14/2026, 3:51:26 AM
Summary
This paper proposes VGG-MADiffRL, a value-gradient-guided multi-agent diffusion reinforcement learning algorithm, and MDCA, a hierarchical control architecture, to solve the problem of cooperative target tracking by multi-AUV ad-hoc networks. The approach addresses challenges such as high-dimensional state-action spaces, training instability, and noise sensitivity in dynamic underwater environments by integrating diffusion policies with value gradients to guide action generation, ensuring faster convergence and higher tracking accuracy.
Entities (8)
Relation Signals (9)
VGG-MADiffRL → uses → Diffusion Model
confidence 96% · VGG-MADiffRL builds on diffusion policies and incorporates value gradients
VGG-MADiffRL → ispartof → MDCA
confidence 95% · Within MDCA, the local online training layer is the policy learning framework; VGG-MADiffRL builds on diffusion policies
VGG-MADiffRL → guides → Action Generation
confidence 94% · incorporates value gradients to guide action generation in the reverse denoising process
MDCA → enables → cooperative tracking
confidence 93% · formulating cooperative tracking for multi-AUV ad-hoc networks as an MDP... The proposed MDCA constitutes a three-tier closed-loop control framework
MDCA → consistsof → Global Intelligent Control Layer
confidence 92% · The proposed MDCA constitutes a three-tier closed-loop control framework: a global intelligent control layer...
MDCA → consistsof → Local Online Training Layer
confidence 92% · ...a local online training layer, and a physical action execution layer.
MDCA → consistsof → Physical Action Execution Layer
confidence 92% · ...a local online training layer, and a physical action execution layer.
VGG-MADiffRL → mitigates → Training Oscillations
confidence 90% · It employs twin value networks with joint optimization and soft target updates to mitigate overestimation and training oscillations
Multi-AUV ad-hoc network → faceschallenges → dynamic topology
confidence 89% · requires networked autonomous underwater vehicles (AUVs) to cooperatively track maneuvering targets under constrained acoustic communication, dynamic topology
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Multi-AUV ad-hoc network-based target tracking requires networked autonomous underwater vehicles (AUVs) to cooperatively track maneuvering targets under constrained acoustic communication, dynamic topology, and uncertain ocean disturbances. Although multi-agent reinforcement learning (MARL) enables decentralized coordination through centralized training, existing methods suffer from high-dimensional joint state-action modeling, noise-sensitive policy generation, leading to unstable training and degraded tracking. To address these issues, we propose VGG-MADiffRL, a value-gradient-guided multi-agent diffusion RL algorithm, and MDCA, a diffusion?based hierarchical control architecture. Leveraging underwater mission characteristics, we model sonar detection mechanisms and ocean current disturbances, formulating cooperative tracking for multi-AUV ad-hoc networks as an MDP. The proposed MDCA constitutes a three-tier closed-loop control framework: a global intelligent control layer, a local online training layer, and a physical action execution layer. This structure enables synergistic optimization across task allocation, local decision processes, and execution feedback. Within MDCA, the local online training layer is the policy learning framework; VGG-MADiffRL builds on diffusion policies and incorporates value gradients to guide action generation in the reverse denoising process, steering the generated actions towards higher expected returns. It employs twin value networks with joint optimization and soft target updates to mitigate overestimation and training oscillations, promoting more stable convergence. Experimental results show that VGG-MADiffRL consistently achieves faster convergence, higher tracking accuracy, and smoother training dynamics in cooperative tracking scenarios, validating its effectiveness and practical engineering value in dynamic underwater settings.
Tags
Links
- Source: https://arxiv.org/abs/2608.12436v1
- Canonical: https://arxiv.org/abs/2608.12436v1
Trouble viewing inline? Open PDF directly →
Full Text
73,274 characters extracted from source content.
Expand or collapse full text
Multi-AUV Ad-hoc network-based Target Tracking: A Value Gradient Guidance Multi-Agent Diffusion Reinforcement Learning Approach Jiaao Ma Chuan Lin Guangjie Han Shengchao Zhu Qian Zhu Ying Liu Zhenyu Wang Thanks: Corresponding author: Guangjie Han Thanks: Jiaao Ma, Chuan Lin, Qian Zhu, Ying Liu and Zhenyu Wang are with the Software College, Northeastern University, Shenyang, China. (e-mails: 2727746375@q.com; chuanlin1988@gmail.com; zhuq@swc.neu.edu.cn; liuy@swc.neu.edu.cn;larrywang1019@outlook.com). Thanks: Guangjie Han is with the Key Laboratory of Maritime Intelligent Network Information Technology, Ministry of Education, Hohai University, Changzhou, China (e-mail: hanguangjie@gmail.com). Thanks: Shengchao Zhu is with the College of Computer Science and Software Engineering, Hohai University, Nanjing, 210013, China (e-mail: zhushengchao77@gmail.com). Abstract Multi-AUV ad-hoc network-based target tracking requires networked autonomous underwater vehicles (AUVs) to cooperatively track maneuvering targets under constrained acoustic communication, dynamic topology, and uncertain ocean disturbances. Although multi-agent reinforcement learning (MARL) enables decentralized coordination through centralized training, existing methods suffer from high-dimensional joint state-action modeling, noise-sensitive policy generation, leading to unstable training and degraded tracking. To address these issues, we propose VGG-MADiffRL, a value-gradient-guided multi-agent diffusion RL algorithm, and MDCA, a diffusion-based hierarchical control architecture. Leveraging underwater mission characteristics, we model sonar detection mechanisms and ocean current disturbances, formulating cooperative tracking for multi-AUV ad-hoc networks as an MDP. The proposed MDCA constitutes a three-tier closed-loop control framework: a global intelligent control layer, a local online training layer, and a physical action execution layer. This structure enables synergistic optimization across task allocation, local decision processes, and execution feedback. Within MDCA, the local online training layer is the policy learning framework; VGG-MADiffRL builds on diffusion policies and incorporates value gradients to guide action generation in the reverse denoising process, steering the generated actions towards higher expected returns. It employs twin value networks with joint optimization and soft target updates to mitigate overestimation and training oscillations, promoting more stable convergence. Experimental results show that VGG-MADiffRL consistently achieves faster convergence, higher tracking accuracy, and smoother training dynamics in cooperative tracking scenarios, validating its effectiveness and practical engineering value in dynamic underwater settings. Index Terms: Autonomous underwater vehicle, multi-AUV ad-hoc network system, underwater target tracking, multi-agent reinforcement learning, diffusion model. I Introduction The ocean contains abundant biological and mineral resources and is a critical domain for sustaining human society and advancing deep ocean strategic initiatives [2]. As underwater communication technologies [15], particularly acoustic communication, have advanced, the research focus has shifted from single AUV systems to AUV swarm systems [16]. In such systems, AUVs form autonomously organized, distributed, and intelligent Multi-AUV ad-hoc Networks (MAANs) [6] that enable collaborative execution of complex underwater missions, such as multiple target tracking and encirclement [8, 24], even in GPS-denied or otherwise unavailable environments. However, underwater environments are dynamic and uncertain [7, 14]. Conventional data-transmission-focused architectures treat the network as a passive conduit, making operational consistency difficult given acoustic channel challenges: rapidly shifting topologies and severely constrained bandwidth [10, 3]. Overcoming this bottleneck requires a fundamental change in AUV network management: from passive data reception to active perception and reasoning about the environment [17]. Although Centralized Training with Decentralized Execution (CTDE) has been widely used to address nonstationarity in multi-agent settings [19, 9], its policy learning and optimization struggle with precise continuous cooperative tracking, especially when node relationships in the autonomous network change frequently. In Multi-AUV ad-hoc network cooperative tracking tasks, existing methods based on MARL still face three main challenges: (1) Insufficient modeling of complex continuous action distributions: traditional deterministic policies struggle to capture the strong interdependencies and couplings of continuous joint actions needed for collaborative decision-making under dynamic topologies, often leading to limited expressiveness and suboptimal decisions [13]; (2) Training instability caused by relying on a single optimization objective: when policy updates depend solely on a single reward signal or local supervisory signal, they become highly sensitive to intermittent communication failures, which can cause oscillations in policy learning, slow convergence, and performance fluctuations [27, 26]; (3) Lack of value guidance during sampling: in the reverse denoising process of diffusion models, the absence of explicit value constraints allows sampled actions to drift away from regions of high expected return and introduces ineffective noise, reducing decision consistency, a problem that is especially harmful in low-bandwidth underwater environments [4]. This work addresses training instability, restricted policy representation, and suboptimal action sampling in cooperative tracking with multi-AUV ad-hoc networks. We build on the generative modeling paradigm of diffusion processes. Using a joint optimization scheme that combines loss functions informed by value estimates with policy gradients, we propose the Value Gradient Guided Multi-Agent Diffusion Reinforcement Learning (VGG-MADiffRL) algorithm. By replacing deterministic actors with diffusion policies and steering the reverse denoising trajectory via value gradients, this framework allows multi-AUV networks to achieve robust policy optimization, faster convergence, and precise cooperative actions under shifting topologies. Its core components are as follows. (1) A diffusion policy architecture that improves the modeling of complex, interdependent cooperative actions. (2) A dual-objective framework that combines value signals with policy gradients to stabilize training when the topology changes. (3) A sampling mechanism that uses value gradients to refine denoising trajectories, prioritizing actions with high expected returns. 1. Framework for policy learning via diffusion models: VGG-MADiffRL replaces the conventional deterministic actor with a diffusion model to construct a generative policy architecture closely integrated with value estimation. This enables stable and robust policy optimization in multi-AUV ad-hoc networks, thereby enhancing the policy capability to model complex, continuous, and highly interdependent action distributions inherent in dynamic cooperative scenarios. 2. Combined optimization mechanism for actor networks: VGG-MADiffRL incorporates a dual objective optimization scheme that jointly minimizes a loss function informed by value estimates alongside a policy gradient loss. The component guided by value signals leverages global value estimates to constrain the diffusion policy update direction, effectively mitigating policy oscillations caused by single objective optimization under volatile network topologies. 3. Diffusion sampling mechanism guided by value gradients: VGG-MADiffRL integrates value signals into the reverse denoising process to iteratively refine the denoising trajectory, steering action sampling toward regions of high expected return from the outset. This mechanism reduces the influence of suboptimal or irrelevant actions and enhances training stability in underwater networks with limited bandwidth and autonomously organized topologies. The remainder of this paper is organized as follows. Section I reviews the related work. Section I displays the formulation of the problem and the system preliminaries. Section IV proposes the MDCA framework. Section V details the implementation of the multi-AUV cooperative tracking algorithm based on VGG-MADiffRL. Section VI presents experimental evaluations and results. Finally, Section VII concludes the paper and outlines future research directions. I Related Works In this section, we mainly review the latest research works related to the subject, specifically divided into the following two areas: 1) AUV target tracking algorithms based on reinforcement learning/multi-agent reinforcement learning; 2) reinforcement learning methods based on diffusion models. I-A AUV Tracking Algorithm Based on RL In [11], ACL-SAC, a target tracking framework for AUVs that combines an attention-based convolutional LSTM with Soft Actor-Critic (SAC), was proposed. It fuses multi-sensor data through attention-weighted temporal feature extraction and optimizes stochastic policies to balance entropy maximization and reward accumulation. A composite reward function and prioritized experience replay improve training efficiency, enabling robust tracking under ocean current disturbances, sound speed uncertainty, and sparse observations. In [20], DDPG-SAC was introduced for high-precision path following of underactuated AUVs under ocean currents. The method decouples surge and heading control, uses an enhanced Line-of-Sight (LOS) law to compensate for drift, and applies a low-pass filter to suppress control chattering. By redesigning state/action spaces and the reward function, it handles nonlinear dynamics, time-varying hydrodynamics, and environmental disturbances, achieving strong generalization and stability in simulations. In [5], AMAML was developed for trajectory tracking under unknown time-varying dynamics. It decomposes the problem into fixed-dynamics subtasks, models tracking as a Markov decision process using LOS guidance, and embeds an attention module to capture hidden dynamic features. Integrating proximal policy optimization with maximum-entropy meta-learning enables fast adaptation across tasks, overcoming poor generalization and model dependency in conventional RL. Building on these single-agent advances, research has shifted toward multi-agent coordination, where diffusion models are increasingly integrated with multi-agent reinforcement learning to address complex cooperative decision-making in dynamic underwater environments. I-B Diffusion Models-Based MARL In [30], MADiff introduces a diffusion-based offline multi-agent reinforcement learning framework that models inter-agent coordination via latent-space attention and unifies decentralized execution with centralized training and teammate modeling. It uses classifier-free guidance to generate high-return trajectories and recovers actions through an inverse dynamics model, effectively addressing extrapolation errors and limited expressivity in offline settings. In [18], HGCD extends this idea to heterogeneous teams by integrating a heterogeneous graph attention mechanism into the diffusion process, enabling adaptation to unseen team compositions. Combined with offline meta-RL, it achieves policy generalization across teams while supporting decentralized deployment and tackling challenges in data diversity, compositional generalization, and real-world applicability. In [25], DMADRL applies diffusion models to online decision-making in semantic vehicular edge computing. It formulates denoising as an MDP to jointly optimize semantic task offloading and resource allocation. Using a composite reward (semantic fidelity, priority, energy) and Gumbel-Softmax reparameterization, it handles mixed discrete-continuous action spaces and dynamic communication constraints, improving system utility and latency. Despite these advances, existing diffusion-based MARL methods are either offline (MADiff, HGCD) or target discrete-continuous hybrid domains like vehicular networks (DMADRL), making them unsuitable for online, purely continuous, fixed-formation multi-AUV tracking in dynamic underwater environments. To bridge this gap, this paper proposes VGG-MADiffRL and the MDCA architecture, which integrate value-gradient-guided diffusion sampling, dual-objective policy optimization, and hierarchical coordination to enhance training stability, action quality, and robustness in complex cooperative underwater tracking tasks. I Preliminary Materials To accurately emulate the obstacle avoidance and target tracking behaviors of multi-AUV ad-hoc networks in dynamic underwater environments, we incorporate ocean current disturbances and inter-node communication constraints into the modeling process, thus capturing realistic underwater network dynamics. Based on this model, the cooperative operation problem of the multi-AUV system is formulated as a Markov decision process (MDP). I-A Ocean Circulation Modeling Given the uncertain and dynamic underwater environment, sonar is employed for precise relative positioning between AUVs and targets. The AUV emits acoustic waves to scan its surroundings, and a sector coverage receiver array captures echoes from multiple directions, enabling target localization via intensity analysis. The target detection process is modeled using the active sonar equation in Eq. (1): EM=SL−2TL+TS−(NL−DI)−DT,EM=SL-2TL+TS-(NL-DI)-DT, (1) where EMEM is the excess margin, SLSL is the source level, TLTL is the transmission loss, TSTS is the target strength, NLNL is the ambient noise level, DIDI is the directivity index, and DTDT is the detection threshold (in dB). The Navier–Stokes equation is a fundamental governing equation for fluid motion. It accurately characterizes the dynamic behavior of fluids and can be used to analyze and calculate the hydrodynamic forces acting on underwater vehicles in marine environments. The mathematical form of the equation is given in Eq. (2): ρ(∂u∂t+u⋅∇u)=−∇p+μ∇2u+F,ρ ( ∂ t+u· )=-∇ p+μ∇^2u+F, (2) where ρ denotes the density of the fluid, uu is the velocity field, ∂u∂t ∂ t represents the temporal variation of the velocity, u⋅∇uu· is the convective term, ∇p∇ p is the pressure gradient, μ denotes the dynamic viscosity, ∇2u∇^2u is the diffusion term, and FF is the external force term. I-B Kalman Filter-Based State Estimation We define the AUV state by its three-dimensional position and velocity, forming a state-space model for Kalman filtering. The state vector is given by Eq. (3): xk=[xk,yk,zk,vx,k,vy,k,vz,k]⊤,x_k= [x_k,\;y_k,\;z_k,\;v_x,k,\;v_y,k,\;v_z,k ] , (3) where xkx_k, yky_k, zkz_k are the position components at step k, and vx,kv_x,k, vy,kv_y,k, vz,kv_z,k the corresponding velocities. The discrete-time dynamics with control inputs and process noise are described by the state transition model in Eq. (4): xk+1=Axk+Buk+wk,x_k+1=Ax_k+Bu_k+w_k, (4) where A is the state transition matrix, B the control input matrix, uku_k the control input, and wkw_k the process noise. Measurements are related to the true state through the observation model with additive noise, as given in Eq. (5): zk=Hxk+vk,z_k=Hx_k+v_k, (5) where zkz_k is the measurement vector, H the observation matrix, and vkv_k the observation noise. A Kalman filter is used to suppress underwater noise and improve the accuracy of the state estimate. In the prediction step, the prior estimate and its covariance are computed from the previous posterior according to Eq. (6): xk|k−1=Axk−1|k−1+Buk−1,Pk|k−1=APk−1|k−1A⊤+Q, \ aligned x_k|k-1&=Ax_k-1|k-1+Bu_k-1,\\ P_k|k-1&=AP_k-1|k-1A +Q, aligned . (6) where xk|k−1x_k|k-1 and Pk|k−1P_k|k-1 are the predicted state and covariance, and Q is the process noise covariance. In the measurement update, the Kalman gain KkK_k is computed from the prior error covariance and the observation noise covariance, balancing the relative confidence in the model prediction and the measurement, as given in Eq. (7): Kk=Pk|k−1H⊤(HPk|k−1H⊤+R)−1,K_k=P_k|k-1H (HP_k|k-1H +R )^-1, (7) where Pk|k−1P_k|k-1 is the prior estimation error covariance matrix, H the observation matrix, and R the observation noise covariance matrix. The prior state estimate is then updated by the observation innovation, which combines the model prediction with the real-time measurement to form the posterior estimate. This step is described by Eq. (8): xk|k=xk|k−1+Kk(zk−Hxk|k−1),x_k|k=x_k|k-1+K_k (z_k-Hx_k|k-1 ), (8) where xk|k−1x_k|k-1 and xk|kx_k|k are the prior and posterior state estimate vectors, zkz_k is the observation vector at step k, and the term zk−Hxk|k−1z_k-Hx_k|k-1 is the observation innovation (residual) that is mapped by KkK_k into the state correction. Finally, the posterior error covariance matrix is updated to reflect the residual uncertainty after incorporating the measurement and to serve as the starting point for the next prediction cycle, as shown in Eq. (9): Pk|k=(I−KkH)Pk|k−1,P_k|k= (I-K_kH )P_k|k-1, (9) where Pk|kP_k|k and Pk|k−1P_k|k-1 are the posterior and prior estimation error covariance matrices, I is the identity matrix, and KkHK_kH represents the combined gain and observation mapping. I-C Markov Process Modeling In the underwater cooperative decision-making process of a multi-AUV system, the interaction with the environment is formalized as a Markov Decision Process (MDP). The MDP is defined by the tuple in Eq. (10): ℳ=(,,,ℛ,γ)M=(S,A,P,R,γ) (10) where S is the state space, A the action space, P the state transition probability, ℛR the reward function, and γ the discount factor. The state space =s1,s2,…,snS=\s_1,s_2,…,s_n\ contains the states of all AUVs. The state of the i-th AUV, sis_i, combines its ego state ηi _i, an environmental perception ϕi _i, and an observation vector oio_i. The ego state includes position pi∈ℝ3p_i ^3 and velocity vi∈ℝ3v_i ^3. The observation vector oi=(κi,σi,λi)o_i=( _i, _i, _i) captures the relative position information of targets, neighboring AUVs, and environmental landmarks, respectively. Its dimension depends on the number of targets NκN_κ, the number of neighbors NσN_σ, and the number of landmarks NλN_λ. The action space =a1,a2,…,anA=\a_1,a_2,…,a_n\ represents the continuous actions of the multi-AUV system. For AUV i, the action vector ai∈ℝ3a_i ^3 provides control inputs along the x, y, and z axes, i.e., ai=[ax,i,ay,i,az,i]⊤a_i=[a_x,i,a_y,i,a_z,i] . Actions are generated by a diffusion policy. To meet execution constraints, the policy outputs undergo range clipping and scaling in the environment execution layer before being converted into executable control commands. The state transition probability P captures the dynamic evolution of system states and is built on a physical kinematic model. The discount factor γ∈[0,1)γ∈[0,1) balances the weight of future returns. Together, these components define the long-term cumulative return used in policy evaluation. The reward function ℛR accounts for several factors: it explicitly incorporates tracking accuracy, collision avoidance, and environmental constraints to improve cooperative multi-AUV tracking performance. The precise formulation is given in Section V. IV Hierarchical Multi-Agent Collaborative Control Architecture Based on Value Gradient-Guided Diffusion Strategy MARL In this section, we present the proposed multi-agent diffusion-based collaborative architecture (MDCA), a CTDE reinforcement learning framework specifically designed to address the highly dynamic topology and communication-constrained nature of multi-AUV ad-hoc networks. Building on this architecture, we introduce the value gradient-guided multi-agent diffusion reinforcement learning algorithm (VGG-MADiffRL). IV-A Overview of MDCA This subsection presents a hierarchical collaborative control architecture for multi-agent systems under the centralized training with decentralized execution framework. By decomposing global coordination objectives into local operational domains, the proposed design strengthens cooperative consistency within each cluster while decoupling interactions between distinct regions. Fig. 1: Multi-Agent Diffusion-Based Collaborative Architecture As illustrated in Fig. 1, the collaborative control architecture for multi-AUV ad-hoc networks comprises three layers: the global intelligent control layer, the local online training layer, and the physical action execution layer. This design aligns well with ad-hoc network characteristics: decentralization, self-organization, dynamic topology, and distributed collaboration. The functional principles and implementation details of these layers are described in the following subsections. Global Intelligent Control Layer: The global intelligent control layer, centered on the Unmanned Surface Vessel-based cooperative gateway (USV-CG), serves as the global coordination entry point of the multi-AUV ad-hoc network. It receives global task commands from the satellite. Through the northbound interface with the underwater controller, this layer performs information parsing, forwarding, and scheduling via a hierarchical block mechanism, enabling efficient interaction with the lower-level local online training layer. The global intelligent control layer also receives the ad-hoc network topology, node states, and task progress reported by the local online training layer in real time, dynamically maintaining the global network topology and jointly performing global task decomposition and policy scheduling based on the mission scenario, link quality, and node resources. This design ensures the multi-AUV ad-hoc network can still stably execute cooperative tracking tasks under topology fluctuations. Local Online Training Layer: The local online training layer formulates cooperative decisions for dynamically formed AUV clusters in designated operational domains. It maintains a reliable data link with the global intelligent control layer via the northbound interface to acquire and interpret global mission directives, then validate commands and encapsulate local protocols. Leveraging current node distribution, link connectivity, and available resource margins within its assigned domain, this layer routes mission instructions to designated AUV units through the southbound interface, enabling distributed cooperative execution. Concurrently, it acquires navigation states, channel quality indicators, and execution deviations from the physical actuation layer to stabilize the local network topology and optimize policies online. The layer transmits aggregated local states and mission outcomes to the global tier, forming a continuous feedback loop supplying precise operational data for strategic planning. Physical Action Execution Layer: As the terminal layer, the physical action execution tier comprises the AUV swarm for mission execution and perception. Via the southbound interface, it receives trajectory planning and cooperative constraint directives from the local online training layer. Each AUV generates control actions using an Actor diffusion policy: the forward process injects noise to characterize environmental and channel uncertainties, while the reverse process reconstructs feasible control sequences guided by critic value gradients. These actions satisfy cooperative constraints and enable high-precision target tracking. Operational data (navigation states and topological variations) is uploaded to an experience replay buffer. This feedback drives online policy updates in the local training layer, ensuring robust cooperative execution under dynamic conditions. In summary, the proposed hierarchical architecture operates under the centralized training with decentralized execution (CTDE) framework and is suitable for Multi-AUV ad-hoc networks. Decomposing global coordination into localized domains, the design strengthens cooperative consistency within each cluster while decoupling interactions between distinct regions. This structure enhances target tracking performance and system robustness under low bandwidth and highly dynamic underwater conditions. IV-B Proposed Multi-Agent Reinforcement Learning Algorithm To address training instability and convergence difficulties in multi-AUV cooperative continuous control tasks caused by non-stationarity in multi-agent environments, Q-value estimation bias, and distributional mismatches in experience replay, this paper proposes a value gradient-guided multi-agent diffusion reinforcement learning algorithm (VGG-MADiffRL). Fig. 2: Forward and Critic-Guided Reverse Diffusion Processes of VGG-MADiffRL IV-B1 Diffusion Architecture Guided by Value Gradients Under the constraints of multi-AUV ad-hoc networks, the forward and reverse diffusion processes of the VGG-MADiffRL algorithm are illustrated in Fig. 2. The framework explicitly accounts for key underwater characteristics of ad-hoc networks—namely, limited bandwidth, highly dynamic topology, and high communication latency. During the forward diffusion phase, progressive noise injection is applied to construct state-action sequences that emulate the dual uncertainties arising from both the complex marine environment and volatile ad-hoc network conditions. In the reverse diffusion phase, a multilayer perceptron (MLP) network performs iterative denoising guided by value gradients, enabling accurate reconstruction of the original action sequence that satisfies cooperative tracking requirements—all under distributed execution with minimal communication overhead. Forward Diffusion Process: The forward diffusion process defines how noise is gradually added to the initial action, enabling the marginal distribution and sampling form of the noisy action at any diffusion step to be computed in closed form. The noise schedule coefficient and cumulative signal retention are defined in Eq. (11): αt=1−βt,α¯t=∏i=1tαi, _t=1- _t, α_t= _i=1^t _i, (11) where βt _t is the noise variance coefficient in step t, αt _t denotes the signal retention ratio per step, and α¯t α_t denotes the cumulative signal retention ratio from the initial timestep to step t. Given the initial action sample, the corresponding marginal distribution is described in Eq. (12): q(xt∣x0)=(xt,α¯tx0,(1−α¯t)I),q(x_t _0)=N\! (x_t;\, α_t\,x_0,\,(1- α_t)I ), (12) where N denotes a Gaussian distribution, the mean term α¯tx0 α_t\,x_0 represents the scaled initial action, and the covariance term (1−α¯t)I(1- α_t)I represents the accumulated noise covariance over time. The sampling form based on the reparameterization trick is given in Eq. (13): xt=α¯tx0+1−α¯tϵ,ϵ∼(,I),x_t= α_t\,x_0+ 1- α_t\, ε, ε (0,I), (13) where xtx_t is the noisy action sample at step t, and ϵ ε is Gaussian noise drawn from the standard normal distribution. This formulation enables direct sampling from the initial action to any diffusion step through reparameterization. Reverse Diffusion Process Guided by Value Gradients: The reverse sampling process progressively reconstructs the initial action from noisy actions through denoising, while further improving action quality with value-function-guided gradients. The reconstruction estimate of the clean action is computed in Eq. (14): x^0=1α¯t(xt−1−α¯tϵθ(xt,t,st)), x_0= 1 α_t (x_t- 1- α_t\, ε_θ(x_t,t,s_t) ), (14) where x^0 x_0 is the denoised action estimate reconstructed from the reverse process, and ϵθ(xt,t,st) ε_θ(x_t,t,s_t) is the output of the noise prediction network at timestep t, which is used to recover the initial action estimate from the noisy action xtx_t. The value-function-guided correction is given in Eq. (15): x^0=x^0+λ∇x^0Qϕ(s,x^0), x_0= x_0+λ\, _ x_0Q_φ(s, x_0), (15) where x^0 x_0 is the action estimate after value gradient guidance, Qϕ(s,x^0)Q_φ(s, x_0) is the Critic value function, ∇x^0Qϕ _ x_0Q_φ denotes the gradient of the value function with respect to the action, pointing toward the direction that most rapidly increases the expected return, and λ is the guidance strength coefficient that controls the influence of the value gradient. The value gradient correction in Eq. (15) applies only during sampling. Gradients through this correction term are truncated, preventing direct parameter updates for the diffusion policy. Instead, this guidance is captured by the joint Actor loss in Eq. (23), where ℒQ-guideL_Q-guide steers the value-guided action estimate x^0 x_0 toward high-value regions, while ℒpgL_pg enhances expected returns through differentiable sampled actions. This design circumvents explicit higher-order gradient computation from the value signal during Actor optimization, improving training stability. The reverse diffusion sampling update is completed by Eq. (16): xt−1=α¯t−1x^0+1−α¯t−1−σt2ϵθ(xt,t,st)+σt,x_t-1= α_t-1\, x_0+ 1- α_t-1- _t^2\, ε_θ(x_t,t,s_t)+ _tz, (16) where xt−1x_t-1 is the action sample at step t−1t-1 in the reverse process, the first term α¯t−1x^0 α_t-1\, x_0 is the guided action estimate, the second term is the predicted noise component, and the third term σt _tz is the injected random noise, where z follows the standard normal distribution. This equation realizes the denoising transition from timestep t to timestep t−1t-1. The noise standard deviation of the reverse process is determined by Eq. (17): σt=η1−α¯t−11−α¯t1−α¯tα¯t−1, _t=η 1- α_t-11- α_t 1- α_t α_t-1, (17) where σt _t is the noise standard deviation used in reverse sampling at step t, and η is the randomness control coefficient (DDPM parameter). When η=0η=0, the sampling is deterministic; when η=1η=1, the sampling is fully stochastic. This parameter balances the diversity and stability of the generation process. Value Gradient Estimation: The value gradient is defined in Eq. (18) to quantify the sensitivity of the Critic value function to action variations: ∇a0Qϕ(s,a0)=[∂Qϕ∂a(1)⋯∂Qϕ∂a(d)]⊤|a=a0, _a_0Q_φ(s,a_0)= . bmatrix ∂ Q_φ∂ a^(1)&·s& ∂ Q_φ∂ a^(d) bmatrix |_a=a_0, (18) where ∂Qϕ∂a(i) ∂ Q_φ∂ a^(i) is the partial derivative of the value function with respect to the i-th action component, and (⋅)⊤(·) denotes vector transpose. This gradient vector points toward the direction of the fastest increase in the value function. IV-B2 Joint Optimization Loss of VGG-MADiffRL In the design of the loss function, we explicitly account for both algorithmic stability and the generative characteristics of the diffusion model, decomposing the overall optimization objective into two components: the dual critic loss and the diffusion-based actor loss. Dual Critic Loss: The critic loss is constructed based on the temporal difference (TD) error. First, the target Q-value for the i-th mini-batch sample is computed as given in Eq. (19), and then the mean squared error is used to measure the deviation between the online critics and the target Q-value, as shown in Eq. (20): yi=ri+γ(1−di)minQ1,ϕi′(s′,a′),Q2,ϕi′(s′,a′),y_i=r_i+γ(1-d_i) \Q _1, _i(s ,a ),\,Q _2, _i(s ,a ) \, (19) ℒcritic(i)=‖Q1,ϕi(s,a)−yi‖2+‖Q2,ϕi(s,a)−yi‖2,L_critic^(i)= \|Q_1, _i(s,a)-y_i \|^2+ \|Q_2, _i(s,a)-y_i \|^2, (20) where s, s′s denote the current and next global states input to the critics; a is the current joint action of all agents; a′a is the next joint action generated by the target actor networks; rir_i and did_i are the immediate reward and termination flag of agent i; γ is the reward discount factor; Q1,ϕiQ_1, _i, Q2,ϕiQ_2, _i are the dual online critic networks for agent i; and Q1,ϕi′Q _1, _i, Q2,ϕi′Q _2, _i are the corresponding dual target critic networks. Under the CTDE framework, the centralized critics receive the global state s and joint action a, while the diffusion-based actors condition on each agent’s local observation oio_i to generate actions. Diffusion Actor Loss: The Q-guided loss term of the diffusion-based actor guides action generation using the Q-values output by the critics, with its mathematical form given in Eq. (21): ℒQ-guide=−λq⋅1N∑i=1NminQ1(s,a0,i),Q2(s,a0,i),L_Q-guide=- _q· 1N _i=1^N \Q_1(s,a_0,i),\,Q_2(s,a_0,i) \, (21) where λq _q is the Q-guidance coefficient, and a0,ia_0,i is the denoised action estimate used in the Q-guidance term. The policy gradient loss directly maximizes the action value evaluated by the critics, with its mathematical expression given in Eq. (22): ℒpg=−1N∑i=1NminQ1(s,ai),Q2(s,ai),L_pg=- 1N _i=1^N \Q_1(s,a_i),\,Q_2(s,a_i) \, (22) where aia_i is the differentiable sampled action used in the policy gradient term. The total actor loss combines both components: ℒactor=ℒQ-guide+ℒpg,L_actor=L_Q-guide+L_pg, (23) where ℒactorL_actor is the final actor loss, ℒQ-guideL_Q-guide is the Q-guided loss defined in Eq. (21), and ℒpgL_pg is the policy gradient loss defined in Eq. (22). V Proposed Multi-AUV Cooperative Target Tracking Scheme This paper considers cooperative target tracking in multi-AUV ad-hoc networks. Using the proposed VGG-MADiffRL algorithm, we present the key tracking strategy for ad-hoc network constraints and describe the algorithm’s complete execution pipeline. V-A Proposed Reward Function for Multi-AUV Tracking We present a composite reward function for multi-AUV ad-hoc networks. In the RL framework, cooperative tracking under ad-hoc network constraints is a Markov decision process (MDP) maximizing cumulative reward, guiding AUVs toward stable coordination in bandwidth-limited, dynamic, and communication-impaired underwater environments. We design three reward components for tracking fidelity, inter-agent safety, and environmental constraints, ensuring efficiency and long-term policy stability in complex ad-hoc network scenarios. The total reward for the i-th AUV at the current timestep is given by Eq. (24): Ri=α⋅rpos+β⋅rcol+δ⋅rland,R_i=α· r_pos+β· r_col+δ· r_land, (24) where RiR_i denotes the instantaneous scalar reward received by the i-th AUV; α, β, and δ are positive weighting coefficients that balance the influence of the target position reward (rposr_pos), the collision penalty (rcolr_col), and the landmark constraint term (rlandr_land) on policy learning. To emulate realistic underwater tracking environments, the target position reward is defined as given in Eq. (25): rpos=−dt,dt>dt,min,wdt−(w+1)dt,min,dt≤dt,min,r_pos= cases-d_t,&d_t>d_t, ,\\ wd_t-(w+1)d_t, ,&d_t≤ d_t, , cases (25) where dtd_t denotes the distance between the agent and the target, dt,mind_t, is the threshold distance defining the target’s proximity zone, and w is a reward modulation parameter within this zone. To prevent collisions among nodes during dense cooperative maneuvers in the ad-hoc network, a safety constraint penalty is designed based on relative inter-agent distances. The collision penalty is formulated as shown in Eq. (26): rcol=−λ1(do,min−do)2,do<do,min,−min(λ2,λ3do),do≥do,min,r_col= cases- _1(d_o, -d_o)^2,&d_o<d_o, ,\\ - ( _2, _3d_o),&d_o≥ d_o, , cases (26) where dod_o represents the relative distance between two AUVs, do,mind_o, is the minimum allowable safe distance, λ1 _1 controls the penalty intensity for close-range proximity, λ2 _2 sets the upper bound for long-range penalties, and λ3 _3 is the linear slope coefficient governing the penalty decay. To smoothly model landmark region constraints within the reward function, a Sigmoid function is introduced, as expressed in Eq. (27): σ(x)=11+e−x,σ(x)= 11+e^-x, (27) where σ(x)σ(x) serves as a smooth activation function that transforms hard-threshold penalty relationships into a continuously differentiable form, thereby enhancing training stability. Building upon Eq. (27), the obstacle avoidance penalty is defined as: rland=−λl∑k=13σ(τk−dl,ksk),r_land=- _l _k=1^3σ ( _k-d_l,ks_k ), (28) where λl _l is the landmark penalty weight, dl,kd_l,k denotes the distance between the agent and the k-th landmark, τk _k represents the constraint threshold for the corresponding landmark, and sks_k is a smoothing adjustment parameter. In summary, the proposed composite reward function comprises three core components: target tracking accuracy, swarm safety constraints, and underwater obstacle avoidance penalties. This formulation comprehensively captures the primary objectives and operational constraints of multi-AUV ad-hoc networks in dynamic underwater scenarios. By closely mirroring real-world underwater task dynamics, this reward mechanism significantly improves simulation fidelity and effectively enhances the learning efficiency, convergence stability, and cooperative robustness of reinforcement learning models under bandwidth-limited and topologically volatile network conditions. V-B Proposed Tracking Algorithm Based on VGG-MADiffRL Algorithm 1 Proposed Multi-AUV Cooperative Target Tracking Algorithm Based on VGG-MADiffRL 1: Number of episodes NepN_ep, episode length LepL_ep, batch size B, update interval IupdateI_update, minimal buffer size SminS_ , soft update coefficient τ, guidance interval IguideI_guide 2: Trained diffusion policies and twin critics for multi-AUV cooperative tracking 3: Initialize environment env and VGG-MADiffRL model 4: Initialize replay buffer D 5: Set global step counter ttotal←0t_total← 0 6: for e=1e=1 to NepN_ep do 7: Reset environment and obtain initial state s 8: for t=1t=1 to LepL_ep do 9: Set critic-guidance flag gt←(tmodIguide=0)g_t (t I_guide=0) 10: for each agent i=1i=1 to N do 11: if gt=1g_t=1 then 12: Generate action ia_i by guided diffusion sampling (Eqs. (14)–(18)) 13: else 14: Generate action ia_i by diffusion policy without guidance 15: Execute joint action =(1,…,N)a=(a_1,…,a_N) in env 16: Observe next state ′s , reward r, and done flag d 17: Store transition (,,,′,,0)(s,a,r,s ,d,0) in D 18: Update current state ←′s 19: ttotal←ttotal+1t_total← t_total+1 20: if ||≥Smin|D|≥ S_ and ttotalmodIupdate=0t_total I_update=0 then 21: Sample minibatch ℬB from D 22: for each agent i=1i=1 to N do 23: Construct target joint action ′a from target diffusion policies 24: Compute target value yiy_i using Eq. (19) 25: Update twin critics by minimizing ℒcritic(i)L_critic^(i) in Eq. (20) 26: Update diffusion actor by minimizing ℒactor(i)L_actor^(i) in Eq. (23) 27: where ℒQ-guide(i)L_Q-guide^(i) and ℒpg(i)L_pg^(i) are defined in Eqs. (21)–(22) 28: Soft update Actor and Critic target networks for all agents by θtarget←(1−τ)θtarget+τθ _target←(1-τ) _target+τθ. 29: if =Trued=True then 30: break In this section, the proposed VGG-MADiffRL-based tracking algorithm is formalized in Algorithm 1. The complete workflow for ad-hoc networks is decomposed into three steps, as illustrated in Fig. 3. Step 1: Simulation Environment Construction. The underwater simulation framework for multi-AUV ad-hoc networks integrates 3D environmental modeling, acoustic channel simulation, and sonar-based state representation, capturing hydrodynamic and topological uncertainties within an MDP and providing a robust validation platform. Step 2: Diffusion Policy Architecture Design. A diffusion-based policy learning framework is deployed where the forward process injects progressive noise to emulate environmental and network fluctuations. During the reverse process, Critic value gradients guide iterative denoising to reconstruct high-quality cooperative action sequences, ensuring precise distributed decision-making under bandwidth constraints. Step 3: Closed-Loop Cooperative Execution. In distributed execution, each AUV acts autonomously on local observations and generates actions via the trained diffusion policy. Interaction trajectories are stored in the experience replay buffer. Periodically, buffered samples are used to update the diffusion policy and dual Critic networks through soft target updates. This closed-loop paradigm enables the multi-AUV system to sustain adaptive coordination and robust target tracking in dynamic underwater environments. The proposed VGG-MADiffRL-based tracking algorithm appears in Algorithm 1, integrating guided diffusion sampling, twin-critic learning, and soft target updates into a unified multi-AUV loop. Algorithm 1 initializes the environment, replay buffer, and global timestep counter ttotalt_total (Lines 1–3) and iterates over episodes (Line 4). Each episode resets environment (Line 5) and proceeds over timesteps (Line 6). At each timestep, tmodIguidet I_guide determines guidance flag gtg_t (Line 7), and each agent generates actions by either value-gradient-guided diffusion sampling (Eqs. (14)–(18), Lines 9–10) or unguided diffusion policy sampling (Lines 11–12). The joint action is executed in the environment, and the next state, reward, and done flag are observed (Lines 13–14). The transition (,,,′,,0)(s,a,r,s ,d,0) is stored in replay buffer D, and the current state and global timestep counter are updated (Lines 15–17). When ||≥Smin|D|≥ S_ and ttotalmodIupdate=0t_total I_update=0 (Line 18), a minibatch is sampled from D (Line 19), and each agent is updated (Line 20): target joint actions constructed from target diffusion policies (Line 21), target values computed via Eq. (19) (Line 22), twin critics optimized via Eq. (20) (Line 23), and diffusion actors optimized via Eq. (23), with ℒQ-guideL_Q-guide and ℒpgL_pg defined in Eqs. (21)–(22) (Lines 24–25). Each update cycle concludes with soft updates of Actor and Critic target networks via θtarget←(1−τ)θtarget+τθ _target←(1-τ) _target+τθ (Line 26). The episode terminates early if the done flag is true (Lines 27–28). This procedure enables stable, sample-efficient cooperative tracking by coupling value-guided diffusion action generation with twin-critic-based policy optimization. Fig. 3: Workflow of the Generative Multi-AUV MARL Algorithm VI Evaluations This section presents a comprehensive experimental evaluation of the proposed algorithm. The cooperative tracking performance of the multi-AUV ad-hoc network is analyzed across multiple metrics, including average reward, convergence stability, tracking accuracy, and robustness. Comparative experiments with state-of-the-art multi-agent reinforcement learning algorithms are conducted to validate the effectiveness of VGG-MADiffRL. VI-A Simulation Setup All experiments are conducted on a computational platform equipped with an AMD Ryzen 9 8940HX processor, RTX 5060 GPU, and 16 GB RAM. All code is implemented in Python 3.10. The simulation environment is built on OceanGym [23], a benchmark environment for underwater embodied agents. We adopt its modular agent-environment interface for standardized multi-agent underwater interaction and extend it with custom hydrodynamic effects and acoustic communication constraints to model the physical dynamics and network conditions of multi-AUV ad-hoc networks. In the evaluations, the target moves at a predefined constant velocity, while AUVs are initially distributed in a circular ring-shaped region approximately 3.5 to 5 km from the target. To comprehensively evaluate algorithm performance under varying network scales, four distinct tracking scenarios are employed in the assessment: 10 AUVs tracking 3 targets, 8 AUVs tracking 3 targets, 6 AUVs tracking 2 targets, and 4 AUVs tracking 2 targets. All the parameters in the evaluations are detailed in Table I. VI-B Results and Discussion We compare VGG-MADiffRL against two groups of MARL methods. The first group consists of five MARL algorithms for continuous control: MASAC [22], MAPPO [26], MAAC [28], MATD3 [1], and MADDPG [12]. The second group consists of two MARL methods designed for underwater AUV scenarios, DSBM [21] and MA-A3C [29]. Our approach is evaluated mainly from the following aspects: 1) convergence speed; 2) tracking accuracy; 3) mean tracking error; 4) error standard deviation; 5) diffusion steps required for strategy generation in our framework; 6) ablation studies validating each component’s contribution to system performance, and 7) system availability under dynamic ad-hoc network conditions. TABLE I: Simulation Parameters Parameter Description Value NAN_A Number of AUVs [4,6,8,10] NTN_T Number of Targets [2,3] lrl_r Learning rate 1e−31e^-3 NEN_E Training rounds 4000 NhN_h Hidden layer neurons 256 γ Discount factor 0.95 τ Network update coefficient 1e−21e^-2 dminϕd_ ^φ Minimum tracking distance 80 m dminκd_ ^κ Minimum AUV distance 80 m LeL_e Episode length 400 BsB_s Buffer size 100,000 UiU_i Update interval 400 MsM_s Minimal buffer size 4000 B Batch size 256 ρ Fluid density 1000 kg/m3 μ Fluid viscosity 10−310^-3 Pa⋅·s D Damping factor 0.25 Δt t Simulation time step 0.1 s (a) Scenario of 4 AUVs Tracking 2 Targets (b) Scenario of 6 AUVs Tracking 2 Targets (c) Scenario of 8 AUVs Tracking 3 Targets (d) Scenario of 10 AUVs Tracking 3 Targets Fig. 4: Convergence Speed Evaluation TABLE I: Tracking Accuracy Comparison Algorithms 4 AUVs Tracking 2 Targets 6 AUVs Tracking 2 Targets 8 AUVs Tracking 3 Targets 10 AUVs Tracking 3 Targets Mean±SDMean± SD Mean±SDMean± SD Mean±SDMean± SD Mean±SDMean± SD VGG-MADiffRL 74.63%± 0.19% 72.73%± 0.04% 77.40%± 0.05% 61.52%± 0.28% DSBM 55.83%± 0.47% 43.83%± 2.78% 38.58%± 1.86% 35.21%± 1.50% MA-A3C 33.42%± 0.61% 31.68%± 1.89% 25.74%± 0.20% 32.29%± 0.66% MAPPO 60.27%± 0.13% 52.82%± 0.41% 67.75%± 0.11% 48.76%± 0.02% MASAC 64.22%± 0.05% 62.97%± 0.02% 45.11%± 1.59% 10.32%± 1.00% MAAC 61.19%± 0.26% 45.07%± 0.17% 41.35%± 1.13% 40.64%± 0.73% MATD3 64.96%± 0.62% 35.89%± 0.48% 48.84%± 0.15% 49.38%± 0.26% MADDPG 48.91%± 1.41% 26.25%± 0.27% 50.46%± 0.26% 47.61%± 0.34% TABLE I: Mean Tracking Error Comparison Algorithms 4 AUVs Tracking 2 Targets 6 AUVs Tracking 2 Targets 8 AUVs Tracking 3 Targets 10 AUVs Tracking 3 Targets Mean±SDMean± SD Mean±SDMean± SD Mean±SDMean± SD Mean±SDMean± SD VGG-MADiffRL 0.1329± 0.0001 0.1159± 0.0001 0.1099± 0.0001 0.1524± 0.0001 DSBM 0.1629± 0.0009 0.1929± 0.0210 0.1852± 0.0048 0.1966± 0.0006 MA-A3C 0.3463± 0.0028 0.2053± 0.0019 0.3718± 0.0017 0.2243± 0.0006 MAPPO 0.1439± 0.0002 0.1497± 0.0004 0.1312± 0.0004 0.2114± 0.0001 MASAC 0.1433± 0.0002 0.1462± 0.0004 0.1807± 0.0011 0.3086± 0.0000 MAAC 0.1379± 0.0003 0.1962± 0.0008 0.2321± 0.0005 0.2792± 0.0004 MATD3 0.1472± 0.0020 0.2541± 0.0046 0.1671± 0.0004 0.1782± 0.0004 MADDPG 0.1678± 0.0009 0.2010± 0.0005 0.1908± 0.0005 0.2246± 0.0015 TABLE IV: Tracking Error Standard Deviation Comparison Algorithms 4 AUVs Tracking 2 Targets 6 AUVs Tracking 2 Targets 8 AUVs Tracking 3 Targets 10 AUVs Tracking 3 Targets Mean±SDMean± SD Mean±SDMean± SD Mean±SDMean± SD Mean±SDMean± SD VGG-MADiffRL 0.1755± 0.0000 0.1818± 0.0000 0.1748± 0.0002 0.1938± 0.0001 DSBM 0.1917± 0.0003 0.1837± 0.0080 0.1952± 0.0030 0.1924± 0.0005 MA-A3C 0.2679± 0.0014 0.1823± 0.0034 0.2423± 0.0009 0.2033± 0.0007 MAPPO 0.1898± 0.0001 0.1816± 0.0002 0.1975± 0.0002 0.2369± 0.0001 MASAC 0.2122± 0.0003 0.2091± 0.0002 0.1886± 0.0009 0.2098± 0.0007 MAAC 0.1806± 0.0000 0.1952± 0.0008 0.2626± 0.0006 0.3483± 0.0001 MATD3 0.2008± 0.0013 0.2976± 0.0006 0.1909± 0.0001 0.2239± 0.0001 MADDPG 0.1970± 0.0006 0.1799± 0.0010 0.2281± 0.0002 0.2478± 0.0005 VI-B1 System Convergence Speed To evaluate the training convergence of VGG-MADiffRL, we compare its convergence performance against the two categories of methods described above across four multi-AUV ad hoc network scenarios. The convergence curves are shown in Figs. 4(a)–4(d). Among the general-purpose MARL baselines, MAPPO achieves the strongest convergence, benefiting from its centralized training with Generalized Advantage Estimation (GAE) and a stochastic Actor policy that balances global coordination with policy diversity. MASAC, MAAC, MATD3, and MADDPG converge more slowly, particularly in larger-scale scenarios (8 AUVs Tracking 3 Targets and 10 AUVs Tracking 3 Targets), where increased agent interactions amplify multi-agent non-stationarity. Their deterministic or entropy-regularized policies cannot adequately model the interdependent action distributions needed for coordinated tracking under dynamic ad hoc topologies. Among the underwater-specific methods, DSBM and MA-A3C train stably but converge to lower returns than VGG-MADiffRL. DSBM uses dynamic-switching attention for multi-target tracking; it performs moderately but plateaus early because its discrete switching mechanism restricts the expressiveness of continuous cooperative actions. MA-A3C adopts a hierarchical software-defined architecture with advantage-attention actor-critic and advantage resampling; it converges steadily but reaches a lower final return because its deterministic policy gradient limits action expressiveness and its reward-weighted attention compression discards fine-grained coordination signals needed under fast-changing ad hoc topologies. VGG-MADiffRL converges faster than all compared methods across all scenarios. During early training, the critic value-guided mechanism drives rapid policy improvement: the diffusion policy generates actions via differentiable sampling, and a joint loss combining policy gradient objectives with Q-guidance terms from global dual-Q network outputs steers updates toward high-value regions. Batch updates that start after the replay buffer reaches a minimum size suppress small-sample bias and improve sample reuse. In later training, the algorithm remains smooth and stable. The dual-Q target networks take the minimum of two independent estimates to reduce overestimation, while soft updates avoid abrupt parameter shifts. Gradient clipping limits update magnitudes, and the inherent action continuity of diffusion-based generation helps avoid the oscillations seen in several baseline methods. VI-B2 Tracking Accuracy In multi-AUV ad-hoc network cooperative tracking tasks, tracking accuracy serves as a core metric for evaluating training effectiveness and policy optimization. To validate the effectiveness of the proposed algorithm, experiments are configured with AUVs initially deployed at a distance of 4.5 km from the target, and a tracking error threshold of 0.8 km is adopted for performance assessment. As summarized in Table I, the comparative results demonstrate that VGG-MADiffRL achieves the highest tracking accuracy under this scenario, significantly outperforming existing baseline methods. These results fully verify the robustness and reliability of the proposed algorithm in achieving high-precision, sustained, and stable tracking within complex, dynamic underwater ad-hoc networks. VI-B3 Mean Tracking Error (MTE) In multi-AUV ad-hoc network cooperative tracking, the mean tracking error quantifies overall temporal tracking deviation and serves as a core metric for evaluating policy accuracy and stability. Under a unified experimental setup, MTE is defined as the sample mean of Euclidean distances between each agent and its target, across all timesteps and agents per evaluation episode. This formulation comprehensively reflects the error level throughout execution, rather than solely on terminal states. Table I shows that VGG-MADiffRL achieves the lowest MTE value among all compared methods, demonstrating a more pronounced error advantage. These results indicate the proposed method enables higher-precision and more robust continuous cooperative tracking in complex dynamic underwater ad-hoc networks. VI-B4 Error Standard Deviation (Error Std) To characterize the stability and fluctuation of tracking errors in multi-AUV ad-hoc network cooperative tracking, this study computes the standard deviation of distance error samples between each agent and its target across all timesteps and agents per evaluation episode, defined as the Error Std metric. This indicator quantifies error dispersion: a smaller value signifies more temporally consistent tracking performance with reduced fluctuations. Table IV shows that VGG-MADiffRL achieves the lowest Error Std among all compared methods, demonstrating that the proposed algorithm not only maintains a low mean tracking error but also exhibits superior error suppression and more stable dynamic tracking performance in complex underwater ad-hoc networks. In Table IV, standard deviations below the reported decimal precision are denoted as ±0.0000± 0.0000. Fig. 5: Convergence Speed Across Different Numbers of Diffusion Time Steps in VGG-MADiffRL (a) Ablation analysis in 4 AUVs Tracking 2 Targets (b) Ablation analysis in 8 AUVs Tracking 3 Targets Fig. 6: Ablation Evaluation (a) Mid-Phase of 8 AUVs Tracking 3 Targets (b) Final Phase of 8 AUVs Tracking 3 Targets (c) Mid-Phase of 10 AUVs Tracking 3 Targets (d) Final Phase of 10 AUVs Tracking 3 Targets Fig. 7: Availability Evaluation VI-B5 Number of Diffusion Time Steps The number of diffusion steps, the count of reverse sampling iterations in the diffusion-based policy generation process, is a critical hyperparameter affecting strategy accuracy and computational overhead. A larger number of diffusion steps enables more thorough denoising in the reverse process, theoretically yielding higher-quality action distributions and more refined policy representations, yet concurrently increases inference overhead and latency. Fig. 5 illustrates a scenario of 4 AUVs in an ad-hoc network Tracking 2 Targets, where comparative evaluations under different diffusion step settings are conducted under unified experimental conditions to analyze their trade-off effects on control precision and efficiency. Experimental results demonstrate that the adopted diffusion step configuration achieves a better balance between tracking effectiveness and computational cost, while maintaining satisfactory cooperative tracking performance in multi-AUV ad-hoc networks. VI-B6 Ablation Evaluation To systematically evaluate the contribution of each key module in the proposed method for multi-AUV ad-hoc networks, an ablation study is conducted. The following comparative variants are configured: (1) removal of the value-gradient-guided reverse diffusion mechanism (excluding value function guidance during action sampling); and (2) replacement of the diffusion policy module with a conventional deterministic policy network to examine the individual impact of diffusion modeling on performance. All other training configurations remain identical to ensure a fair comparison. Fig. 6 shows the complete method consistently outperforms both ablated variants in convergence stability and cumulative return. These results demonstrate that both the value-gradient guidance mechanism and the diffusion-based policy module play critical roles in enhancing cooperative tracking performance, thereby validating the effectiveness of the proposed architectural design. VI-B7 Availability Evaluation To show the training convergence and cooperative tracking performance of the proposed method, we build a high-fidelity underwater simulation environment using the 3D modeling and physics engine of Unity. Fig. 7 visualizes the cooperative tracking process and environment configuration, showing the mid and late stages of 10 AUVs tracking three targets and 8 AUVs tracking three targets. Yellow moving entities represent AUVs, glowing spheres represent dynamic targets, spirals represent obstacles, and colored trajectory lines mark the historical path of each AUV. The simulation reproduces complex underwater dynamics (acoustic communication constraints, ocean current disturbances, and time-varying network topologies) and provides a reliable platform for validating the stability and effectiveness of the algorithm. All AUV nodes, targets, and obstacles are modeled and rendered in real time, enabling direct assessment of the algorithm’s operation in dynamic ad-hoc networks. VII Conclusion This paper investigated cooperative target tracking in multi-AUV ad-hoc networks and proposed the MDCA hierarchical control architecture together with the VGG-MADiffRL algorithm. MDCA decomposes global coordination into three layers (global intelligent control, local online training, and physical action execution), enabling synergistic optimization under the CTDE paradigm. VGG-MADiffRL introduces three key innovations: a diffusion-based policy that replaces the deterministic Actor to model complex continuous action distributions; a dual-objective joint optimization mechanism that combines Q-guided and policy gradient losses to stabilize training under volatile topologies; and a value-gradient-guided reverse sampling mechanism that steers the denoising process toward high-return action regions, reducing ineffective sampling and improving policy robustness. Extensive experiments across four multi-AUV tracking scenarios demonstrate that VGG-MADiffRL consistently outperforms seven state-of-the-art MARL algorithms in convergence speed, tracking accuracy, mean tracking error, and error stability. Ablation studies confirm that both the value-gradient guidance and the diffusion policy module contribute substantially to overall performance. Several directions warrant further work: (1) optimizing underwater obstacle avoidance to reduce potential AUV damage; (2) balancing energy consumption among AUVs to extend system endurance; and (3) designing robust control frameworks that explicitly account for unstable underwater acoustic communication. References [1] J. Ackermann, V. Gabler, T. Osa, and M. Sugiyama (2019) Reducing overestimation bias in multi-agent domains using double centralized critics. External Links: 1910.01465 Cited by: §VI-B. [2] D. M. Bailey and C. R. Hopkins (2023) Sustainable use of ocean resources. Marine Policy 154, p. 105672. External Links: ISSN 0308-597X, Document Cited by: §I. [3] J. Huang, X. Ye, Y. Wang, and L. Fu (2025) Leveraging propagation delays: a delay-aware multiagent reinforcement learning mac protocol for underwater acoustic networks. IEEE Internet of Things Journal 12 (20), p. 42076–42089. External Links: Document Cited by: §I. [4] M. Janner, Y. Du, J. Tenenbaum, and S. Levine (2022) Planning with diffusion for flexible behavior synthesis. In Proceedings of the 39th International Conference on Machine Learning, K. Chaudhuri, S. Jegelka, L. Song, C. Szepesvari, G. Niu, and S. Sabato (Eds.), Proceedings of Machine Learning Research, Vol. 162, p. 9902–9915. Cited by: §I. [5] P. Jiang, S. Song, and G. Huang (2022) Attention-based meta-reinforcement learning for tracking control of auv with time-varying dynamics. IEEE Transactions on Neural Networks and Learning Systems 33 (11), p. 6388–6401. External Links: Document Cited by: §I-A. [6] S. Jiang (2019) On securing underwater acoustic networks: a survey. IEEE Communications Surveys & Tutorials 21 (1), p. 729–752. External Links: Document Cited by: §I. [7] Y. Jiang, K. Zhang, M. Zhao, and H. Qin (2024) Adaptive meta-reinforcement learning for auvs 3d guidance and control under unknown ocean currents. Ocean Engineering 309, p. 118498. External Links: ISSN 0029-8018, Document Cited by: §I. [8] J. Li and Q. Chen (2026) A multi-auv adaptive collaborative target coverage algorithm for unknown environment. Ad Hoc Networks 180, p. 104033. External Links: ISSN 1570-8705, Document Cited by: §I. [9] L. Li, R. An, Z. Guo, and J. Gao (2025) Multi-auv cooperative search for moving targets based on multi-agent reinforcement learning. Journal of Marine Science and Engineering 13 (11). External Links: ISSN 2077-1312, Document Cited by: §I. [10] Z. Li, J. Du, C. Jiang, W. Mi, and Y. Ren (2024) HA-marl: heuristic and apf assisted multi-agent reinforcement learning for wireless data sharing in auv swarms. In ICC 2024 - IEEE International Conference on Communications, Vol. , p. 5401–5406. External Links: Document Cited by: §I. [11] D. Liang, J. Chu, Y. Cui, D. Liang, and Y. Feng (2024) Underwater dynamic tracking control of auv based on complex environment simulation and acl-sac deep reinforcement learning. IEEE Transactions on Intelligent Vehicles (), p. 1–16. External Links: Document Cited by: §I-A. [12] R. Lowe, Y. Wu, A. Tamar, J. Harb, P. Abbeel, and I. Mordatch (2017) Multi-agent actor-critic for mixed cooperative-competitive environments. In Proceedings of the 31st International Conference on Neural Information Processing Systems, NIPS’17, Red Hook, NY, USA, p. 6382–6393. External Links: ISBN 9781510860964 Cited by: §VI-B. [13] S. Luo, Y. Li, S. Liu, X. Zhang, Y. Shao, and C. Wu (2024) Multi-agent continuous control with generative flow networks. Neural Networks 174, p. 106243. External Links: ISSN 0893-6080, Document Cited by: §I. [14] A. Luvisutto, A. Celani, F. Renda, C. Stefanini, and G. De Masi (2025) Enhancing collaboration in uncertain environment: multi-agent reinforcement learning for underwater monitoring. Expert Systems with Applications 277, p. 127256. External Links: ISSN 0957-4174, Document Cited by: §I. [15] S. A. H. Mohsan, Y. Li, M. Sadiq, J. Liang, and M. A. Khan (2023) Recent advances, future trends, applications and challenges of internet of underwater things (iout): a comprehensive review. Journal of Marine Science and Engineering 11 (1). External Links: ISSN 2077-1312, Document Cited by: §I. [16] L. Paull, S. Saeedi, M. Seto, and H. Li (2014) AUV navigation and localization: a review. IEEE Journal of Oceanic Engineering 39 (1), p. 131–149. External Links: Document Cited by: §I. [17] H. Peng, K. Jiang, D. Yuan, Z. Zeng, and Z. Wu (2026) Bio-inspired hierarchical multi-agent reinforcement learning for auv swarm energy weakest-link mitigation. Ocean Engineering 343, p. 123385. External Links: ISSN 0029-8018, Document Cited by: §I. [18] L. Pimentel, S. Ye, J. E. G. Pagan, and M. Gombolay (2025) Diverse heterogeneous graph conditioned diffusion for multi-agent teaming. In Proceedings of the 24th International Conference on Autonomous Agents and Multiagent Systems, AAMAS ’25, Richland, SC, p. 2714–2716. External Links: ISBN 9798400714269 Cited by: §I-B. [19] T. Rashid, M. Samvelyan, C. S. De Witt, G. Farquhar, J. Foerster, and S. Whiteson (2020) Monotonic value function factorisation for deep multi-agent reinforcement learning. J. Mach. Learn. Res. 21 (1). External Links: ISSN 1532-4435 Cited by: §I. [20] C. Wang, J. Du, J. Wang, and Y. Ren (2021) AUV path following control using deep reinforcement learning under the influence of ocean currents. In Proceedings of the 2021 5th International Conference on Digital Signal Processing, ICDSP ’21, New York, NY, USA, p. 225–231. External Links: ISBN 9781450389365, Document Cited by: §I-A. [21] S. Wang, C. Lin, G. Han, S. Zhu, Z. Li, Z. Wang, and Y. Ma (2025) Multi-auv cooperative underwater multi-target tracking based on dynamic-switching-enabled multi-agent reinforcement learning. IEEE Transactions on Mobile Computing 24 (5), p. 4296–4311. External Links: Document Cited by: §VI-B. [22] X. Wu, X. Li, J. Li, P. C. Ching, V. C. M. Leung, and H. V. Poor (2021) Caching transient content for iot sensing: multi-agent soft actor-critic. IEEE Transactions on Communications 69 (9), p. 5886–5901. External Links: Document Cited by: §VI-B. [23] Y. Xue, M. Mao, X. Ru, Y. Zhu, B. Ren, S. Qiao, M. Wang, S. Deng, X. An, N. Zhang, Y. Chen, and H. Chen (2025) OceanGym: a benchmark environment for underwater embodied agents. External Links: 2509.26536, Link Cited by: §VI-A. [24] X. Yan, W. Luo, J. Jia, D. Jiang, and T. Zhang (2025) Static consensus analysis of multi-auv systems with impulsive protocol and time delays. Ocean Engineering 331, p. 121370. External Links: ISSN 0029-8018, Document Cited by: §I. [25] Y. Yang, W. Ma, W. Sun, J. He, Y. Fu, C. Yuen, and Y. Zhang (2025) Diffusion-based multi-agent reinforcement learning for semantic vehicular edge computing. IEEE Transactions on Services Computing 18 (6), p. 3668–3681. External Links: Document Cited by: §I-B. [26] C. Yu, A. Velu, E. Vinitsky, J. Gao, Y. Wang, A. Bayen, and Y. WU (2022) The surprising effectiveness of ppo in cooperative multi-agent games. In Advances in Neural Information Processing Systems, S. Koyejo, S. Mohamed, A. Agarwal, D. Belgrave, K. Cho, and A. Oh (Eds.), Vol. 35, p. 24611–24624. Cited by: §I, §VI-B. [27] K. Zhang, Z. Yang, and T. Başar (2021) Multi-agent reinforcement learning: a selective overview of theories and algorithms. In Handbook of Reinforcement Learning and Control, K. G. Vamvoudakis, Y. Wan, F. L. Lewis, and D. Cansever (Eds.), p. 321–384. External Links: ISBN 978-3-030-60990-0, Document Cited by: §I. [28] J. Zhao, T. Zhu, S. Xiao, Z. Gao, and H. Sun (2022) Actor-critic for multi-agent reinforcement learning with self-attention. International Journal of Pattern Recognition and Artificial Intelligence 36 (09), p. 2252014. External Links: Document Cited by: §VI-B. [29] S. Zhu, G. Han, C. Lin, and Q. Tao (2024) Underwater target tracking based on hierarchical software-defined multi-auv reinforcement learning: a multi-auv advantage-attention actor-critic approach. IEEE Transactions on Mobile Computing 23 (12), p. 13639–13653. External Links: Document Cited by: §VI-B. [30] Z. Zhu, M. Liu, L. Mao, B. Kang, M. Xu, Y. Yu, S. Ermon, and W. Zhang (2024) MADiff: offline multi-agent learning with diffusion models. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, Cited by: §I-B. Jiaao Ma is currently pursuing a Bachelor’s degree at the Software College, Northeastern University, Shenyang, China. His research interests include reinforcement learning, diffusion models, and supervised learning. Chuan Lin [S’17, M’20] is currently an associate professor with the Software College, Northeastern University, Shenyang, China. He received the B.S. degree in Computer Science and Technology from Liaoning University, Shenyang, China in 2011, the M.S. degree in Computer Science and Technology from Northeastern University, Shenyang, China in 2013, and the Ph.D. degree in computer architecture in 2018. From Nov. 2018 to Nov. 2020, he was a Postdoctoral Researcher with the School of Software, Dalian University of Technology, Dalian, China. His research interests include UWSNs, industrial IoT, software-defined networking. Guangjie Han (Fellow, IEEE) is currently a Professor with the Department of Internet of Things Engineering, Hohai University, Changzhou, China. He received his Ph.D. degree from Northeastern University, Shenyang, China, in 2004. In February 2008, he finished his work as a Postdoctoral Researcher with the Department of Computer Science, Chonnam National University, Gwangju, Korea. From October 2010 to October 2011, he was a Visiting Research Scholar with Osaka University, Suita, Japan. From January 2017 to February 2017, he was a Visiting Professor with City University of Hong Kong, China. From July 2017 to July 2020, he was a Distinguished Professor with Dalian University of Technology, China. His current research interests include Internet of Things, Industrial Internet, Machine Learning and Artificial Intelligence, Mobile Computing, Security and Privacy. Dr. Han has over 500 peer-reviewed journal and conference papers, in addition to 160 granted and pending patents. Currently, his H-index is 82 and i10-index is 400 in Google Citation (Google Scholar). The total citation count of his papers raises above 25000 times. Dr. Han is a Fellow of the UK Institution of Engineering and Technology (FIET). He has served on the Editorial Boards of up to 10 international journals, including the IEEE TII, IEEE TCCN, IEEE TVT, IEEE TNSM, IEEE Systems, etc. He has guest-edited several special issues in IEEE Journals and Magazines, including the IEEE JSAC, IEEE Communications, IEEE Wireless Communications, Computer Networks, etc. Dr. Han has also served as chair of organizing and technical committees in many international conferences. He has been awarded 2020 IEEE Systems Journal Annual Best Paper Award and the 2017-2019 IEEE ACCESS Outstanding Associate Editor Award. He is a Fellow of IEEE. Shengchao Zhu (Student member, IEEE) received his B.S. degree in Internet of Things Engineering from Hohai University, Changzhou, China, in 2023. He is currently pursuing the Ph.D. degree with the Department of Computer Science and Technology at Hohai University, Nanjing, China. His current research interests include swarm intelligence, swarm ocean, Multi-Agent Reinforcement Learning. Qian Zhu is an associate professor with the Software College, Northeastern University, Shenyang, China. She received the B.S. degree in Information and Computing Science (2006), the M.S. degree in Operation Science and Control Theory (2008), and the Ph.D. degree in Communication and Information System (2018), all from Northeastern University, Shenyang, China. Her research interests include artificial intelligence optimization algorithms, Unmanned Aerial Vehicle (UAV) technology and software development for applications. Ying Liu received the B.S., M.S., and Ph.D. degrees from Northeastern University, Shenyang, China, in 2003, 2006, and 2012, respectively, all in computer science. She is currently an Associate Professor with the College of Software, Northeastern University. She has published over 50 articles, and refereed conference papers. Her current research interests include Service Computing and Edge Computing.