1. Introduction
AUVs are unmanned mobile platforms that can navigate underwater and perform sensing, inspection, search, and data collection tasks without real-time human control. They are important tools for marine exploration and underwater missions [
1]. Driven by the development of applications such as ocean resource exploration [
2], underwater rescue [
3], marine environmental monitoring [
4], underwater sensor data collection [
5], and coverage optimization in underwater sensor networks [
6], AUV-based target search and cooperative operation have received increasing attention. Limited by its sensing range and operational efficiency, a single-AUV system can hardly satisfy the practical requirements of multi-target search in complex underwater environments. By contrast, through distributed cooperation, an AUV swarm can expand search coverage, improve target detection, and reduce the impact of individual platform failures on the overall mission. Therefore, AUV swarms have become an important technical approach for addressing underwater multi-target search problems [
7].
In recent years, extensive research has been devoted to AUV target search. From the perspective of algorithmic implementation, existing approaches can be broadly divided into Bio-inspired Neural Networks (BNNs), reinforcement learning (RL), heuristic optimization algorithms, and other methods.
BNN methods have been widely studied in AUV target search. Cao et al. [
8,
9,
10] applied the Glasius BNN (GBNN) to 3D underwater environments, where dynamic propagation among neurons was used to realize target search and obstacle avoidance, and further combined it with a self-organizing map to optimize task allocation. Song et al. [
11] verified the real-time obstacle-avoidance capability of the GBNN in two-dimensional environments with dynamic obstacles and achieved multi-target search. Li et al. [
12] reformulated the GBNN update process as a convolution operation and improved global information propagation through mean pooling and multi-step convolution. Sun et al. [
13] combined fuzzy control with a BNN to improve trajectory smoothness and navigation safety. These methods do not require offline training. Instead, their parameters are preset and can adapt online to environmental changes. However, further improvements are still needed in complex marine environment modeling and computational complexity control.
RL methods have provided new adaptive decision-making tools for AUV target search. Wang et al. [
14] incorporated spatiotemporal information into the multi-agent Deep Deterministic Policy Gradient algorithm and constructed an exploration map to reduce repeated search. Cai et al. [
15] combined fuzzy clustering with an RL based on policy iteration and achieved efficient cooperative search over large areas based on regional importance. Liu et al. [
16] proposed autonomous collaborative search learning, which improved cooperative decision-making for multiple AUVs through an information fusion mechanism. Hou et al. [
17] combined the Age of Information mechanism with multi-agent reinforcement learning to investigate a heterogeneous unmanned-system cooperative optimization method for AUV target search in complex underwater environments. Masmitja et al. [
18] proposed an underwater target localization method that combines Deep RL with Least Squares and Particle Filter estimation to achieve range-only target localization using AUVs and Autonomous Surface Vehicles. However, existing RL methods still face several limitations, including low sample efficiency, poor training stability, limited generalization ability, and insufficient consideration of underwater communication constraints.
Heuristic optimization algorithms have also been widely applied to AUV target search. Yan et al. [
19] proposed an improved Ant Colony Optimization algorithm with fuzzy logic and a dynamic pheromone evaporation mechanism, which improved the efficiency of AUV search path planning. Ni et al. [
20] developed a three-stage cooperative search strategy based on an improved Dolphin Swarm Algorithm, which enhanced the search efficiency and environmental adaptability of multiple AUVs in unknown 3D environments. Li et al. [
21] addressed target search in a heterogeneous AUV–USV system and proposed a local dynamic predictive control framework combined with the Lévy flight method to improve space exploration capability and system robustness. Under local prior information, Li et al. [
22] proposed a two-layer search strategy based on the Grey Wolf-enhanced Remora Optimization Algorithm, which improved the efficiency and stability of cooperative search by multiple AUVs. You et al. [
23] combined Voronoi diagrams with a biological competition model to allocate task regions for multiple AUVs and further proposed an improved Particle Swarm Optimization method for search path planning within subregions. Although existing heuristic methods perform well in task allocation and path optimization, they still lack a unified framework for task allocation and safe obstacle avoidance in multi-target search by AUV swarms in unknown three-dimensional obstacle environments.
In addition, some researchers have explored AUV target search from other algorithmic perspectives. For complete-coverage path planning, Yao et al. [
24] proposed a hierarchical search method that combines region decomposition with adaptive elliptical spiral coverage for efficient search. Zhang et al. [
25] adopted K-means+ for dynamic cooperative partitioning and used the Optimal Allocation Minimum Spanning Tree to generate smooth backtrack-free paths. For region partitioning and task allocation, Ling et al. [
26] performed multi-AUV task region division based on Minimum Spanning Tree clustering. Wang et al. [
27] formulated multi-AUV task allocation as a capacitated vehicle routing problem and used the SCIP solver to obtain energy-constrained solutions. Ling et al. [
28] further improved multi-AUV cooperative search efficiency by combining Minimum-Spanning-Tree-based target clustering with path optimization. For adaptive search in unknown environments, Li et al. [
29] used forward-looking sonar information to guide multiple AUVs to search unfinished subregions and proposed a multi-AUV adaptive predictive search algorithm for unknown 3D environments, achieving target search and localization. In terms of heterogeneous cooperation and encirclement, Xue et al. [
30] proposed a USV–AAV heterogeneous cooperative search algorithm based on time-series constraints, while Li et al. [
31] adopted distributed dynamic predictive control to realize cooperative search and encirclement. Although these methods have advanced AUV target search, they still face limitations in terms of their environmental adaptability, communication constraints, and cooperative capability in 3D scenarios.
Based on the above review, existing studies still have several limitations. First, some methods rely on global information, prior environmental knowledge, or centralized coordination, which may be difficult to obtain in unknown underwater environments with limited communication. Second, many studies focus on a single aspect of the problem, such as task allocation, path optimization, coverage planning, or obstacle avoidance, while the integration of cooperative multi-target search and safe navigation in a unified distributed framework remains insufficient. Third, the cooperative response of AUV swarms to multiple detected targets has not been adequately addressed, especially under local communication and three-dimensional motion constraints.
To address the limitations identified above, particularly the reliance on global information and the insufficient consideration of multi-target response under constrained communication conditions, this paper investigates a distributed collaborative search framework for AUV swarms operating in unknown 3D underwater environments with obstacles. The proposed framework enables AUVs to coordinate search activities using local sensing and communication information, while dynamically adjusting their search behaviors according to target-response information. Specifically, AUVs can switch between independent exploration and target-oriented cooperation, allowing vehicles to conduct cooperative and precise target searches while continuing to explore the environment. This design aims to support effective distributed multi-target search while satisfying navigation safety requirements in complex underwater environments. The main contributions of this paper are summarized as follows:
- ⮚
A distributed collaborative search framework is established for multi-target search by AUV swarms in unknown 3D underwater obstacle environments under limited communication conditions.
- ⮚
An adaptive dual-state search mechanism is designed, where the roaming state enables rapid exploration of the task area to detect targets, and the cooperative state focuses on precise target search. Driven by a target response function, AUVs seamlessly transition between these two states to ensure efficient multi-target search.
- ⮚
A motion-constrained 3D search strategy is developed by integrating virtual-force-based swarm interaction, PSO-based search updating, and tangent-plane obstacle avoidance to jointly address search efficiency, coordination, and navigation safety.
- ⮚
Simulation studies in unknown 3D obstacle environments demonstrate the effectiveness of the proposed framework in terms of adaptive state switching, multi-target search, and obstacle avoidance under limited communication conditions.
The remaining content of this paper is organized as follows.
Section 2 introduces the multi-target search model. In
Section 3, the search method is proposed. In
Section 4, the method is verified by simulation. The conclusions are given in
Section 5.
2. Modeling of Seafloor Target Search in Unknown Environments
The search task is performed in a 1000 m × 1000 m × 200 m underwater environment, with a horizontal area of 1000 m × 1000 m and a vertical range of 200 m. The seafloor elevation difference is less than 200 m and the seabed slope changes continuously, implying that no cliffs are present. The task entities include AUVs, targets, and obstacles. The AUV swarm is denoted by
where
denotes the
-th AUV. The target set is denoted by
By discretizing the seafloor, the obstacle set is obtained as
At time
, the positions of
,
, and
are denoted by
,
, and
, respectively, and the velocity of
is denoted by
. The Euclidean distances among entities are defined as follows: the distance between two AUVs is
the distance between an AUV and a target is
and the distance between an AUV and an obstacle is
Within the mission environment, the target search task can be described as follows. Given a target-reaching threshold
, if there exists an AUV such that
then target
is regarded as successfully found. The search mission is complete when all targets have been found. In practice,
is calibrated based on the effective range of onboard target-verification sensors (e.g., optical cameras or other close-range observation devices).
The AUVs participating in the search are homogeneous. The velocity of
, denoted by
satisfies
. Each AUV can be located at any obstacle-free position within the mission environment. Let
,
, and
denote the maximum communication range, obstacle sensing range, and target detection range, respectively. The communication between AUVs is limited by the maximum communication range
. Two AUVs can communicate only when
In addition, an AUV can detect a target when and can perceive an obstacle when . In practice, is determined by the capability of acoustic modems, while and are calibrated based on the specifications of onboard detection and obstacle-avoidance sensors.
The targets are assumed to be stationary and located within a central 800 m × 800 m × 200 m region of the mission environment, with a vertical distance of 0–20 m from the seafloor. During the search process, each AUV can continuously detect surrounding target signals through onboard sensors. The target signal is related to
through an environment-disturbed function, which is referred to as the target response function and is given in Equation (9):
where
is the target signal detected by
at time
,
represents the stochastic ambient noise,
denotes the environmental attenuation coefficient with
, and
denotes the constant signal power of the target. The term
denotes the actual AUV–target distance in the simulation environment, which is introduced only for target signal modeling and cannot be directly obtained by the AUV. This formulation replicates the physical characteristics of underwater sensory signals decaying over distance and suffering from environmental interference, thereby providing a realistic simulation environment for algorithm validation [
32].
The seafloor surface is discretized into stationary obstacle points with a spacing of 1 m, and the position of each obstacle point is expressed as .
When operating in the three-dimensional mission environment, an AUV can obtain its own position and velocity through onboard sensors, communicate with other AUVs within its communication range, and detect obstacle information within its sensing range. Accordingly, the state of the AUV is defined as follows:
where
,
and
denote the coordinates of the
-th AUV at time
in the global Cartesian coordinate system
, respectively.
denotes the velocity of the
-th AUV at time
;
denotes the yaw angle, and
denotes the pitch angle.
3. Proposed Method
3.1. Response-Threshold-Based Multi-Target Task Allocation
To improve the coordination and efficiency of target search by an AUV swarm, three states are defined for the AUVs: roaming search, cooperative search, and a task-completed state.
As shown in
Figure 1, when no target information is available, an AUV remains in the roaming search state, in which the AUVs repel and disperse from one another at the maximum speed to rapidly explore the task environment. Once a target signal is detected, a Response-Threshold-based Multi-Target Task Allocation (RT-MTTA) model is employed to form sub-alliances. The AUVs in the same sub-alliance then search for the corresponding target, and their states switch to the cooperative search state. When the distance between an AUV and a target is smaller than a given threshold, i.e.,
, the target is regarded as found. The AUV then dissolves the sub-alliance, broadcasts that the target has been found, and continues searching for other targets. The detailed execution of the RT-MTTA is presented in the remainder of this subsection.
During the search process, task allocation mainly depends on how an AUV autonomously organizes, selects between different targets, and decides whether to switch tasks before completing the selected one. Specifically, the AUV first detects the target response values within its sensing range through onboard sensors. If multiple target signals are detected, the response probability of the AUV to each target is calculated using a response probability evaluation model, after which a roulette-wheel selection algorithm is employed to determine the search target. The response probability evaluation is given as follows:
where
denotes the signal of target
detected by AUV
at time
. If the AUV detects signals from
targets, then
denotes the probability that
responds to the stimulus of target
. The roulette-wheel selection is formulated as follows:
When , AUV selects target for cooperative search, where . This mechanism executes probabilistic selection based on target signal intensities, avoiding a greedy search for localized individual targets and thereby maintaining search diversity.
An AUV can obtain target information in two ways. The AUV that directly detects target signals through onboard sensors is referred to as a Type-I AUV. The AUV that fails to detect target signals but indirectly acquires target information through communication with other AUVs is referred to as a Type-II AUV. If the target information obtained by Type-I and Type-II AUVs corresponds to the same target, they can jointly participate in the corresponding search task. When multiple AUVs are assigned to the same target, they form a sub-alliance for coordinated search.
In the 3D mission environment, uneven AUV resource allocation may occur during sub-alliance formation, such as multiple AUVs being assigned to the same target while only a few AUVs are assigned to another target. To address this issue, a closed-loop regulation mechanism is introduced. After the initial sub-alliance task allocation, the resource allocation of each sub-alliance is re-evaluated. If the number of members in a sub-alliance reaches the upper limit , only the top AUVs are retained, while the remaining AUVs withdraw from the alliance and either select another target or switch to the roaming search. If the number of members in a sub-alliance is below , the alliance can recruit nearby AUVs to perform their task.
The member selection priorities are defined as follows. Type-I AUVs have higher priority than Type-II AUVs. If the priority is the same, the AUV with the stronger target signal is preferred. If the number of Type-I AUVs is smaller than
, the remaining members are selected from Type-II AUVs by giving priority to those closer to their communicating AUV. If multiple Type-II AUVs have the same distance to their communicating AUV, the one with the stronger signal is preferred. The detailed rules are listed in
Table 1.
The AUV swarm must not only avoid all obstacles but also search for all targets. To accomplish the target-search tasks efficiently, AUVs can form sub-alliances based on detected target signals and communication with neighboring AUVs, so that multiple AUVs jointly search for the same target. In
Table 1, an example of member priority is given for sub-alliance
. In this case, AUV
detects the signal of target
and then is classified as a Type-I AUV. Since the number of Type-I AUVs is smaller than
,
recruits neighboring AUVs within its communication range. After receiving the recruitment message from
,
,
,
,
, and
join the search task for target
, thereby forming sub-alliance
. According to the member-selection priority rules, the priority order in sub-alliance
is
, and these members jointly participate in the search task for target
.
3.2. Roaming Search Based on the Virtual Force Model
When no target signal is detected, the AUVs conduct roaming search to explore the mission area more efficiently. A virtual force model is adopted, in which repulsive forces are generated when the inter-AUV distance is smaller than to promote rapid dispersion of the AUV swarm. For computational simplicity, each AUV is assumed to be affected only by the two nearest neighboring AUVs.
Assume that the nearest neighboring AUV of
is
. Their positions at time
are given by
and
, respectively, with
. The repulsive force exerted by
on
is calculated by Equation (13), and its direction is from the former to the latter:
where
denotes the repulsive force exerted by
on
at time
.
denotes the obstacle-avoidance enhancement distance, and
is a coefficient for path optimization.
If the two nearest neighboring AUVs of
both satisfy
the virtual forces acting on
are as illustrated in
Figure 2. In the figure,
, and
. The virtual forces
and
are both calculated according to Equation (13). The resultant virtual force acting on
is given by the vector sum
For an AUV in the roaming search state, its velocity is directed along the resultant virtual force. Accordingly, the desired velocity at the next step is given by Equation (15).
3.3. Cooperative Search Based on Motion-Constrained 3D Particle Swarm Optimization
An AUV swarm is a typical distributed system. Since a mapping exists between an AUV swarm and a particle swarm, Particle Swarm Optimization can be introduced into the AUV cooperative search decision. By further considering the motion constraints and communication limitations of AUVs, a motion-constrained particle swarm model is constructed to compute the desired velocity
at the next step, as described below:
where
and
denote the position vector and velocity vector of AUV
at step
, respectively;
is the velocity obtained directly by particle swarm iteration; and
is the desired velocity vector of AUV
at the next step. The parameter
is introduced to account for the inertia of AUV motion, while
represents the discretization of continuous time and can also be regarded as a step size.
and
are the individual cognitive coefficient and social cognitive coefficient, respectively;
and
are random variables in the interval (0, 1); and
is the inertia weight.
denotes the best position experienced by AUV
since joining the current sub-alliance, and
denotes the best position visited by the alliance up to step
. The number of particles corresponds to the number of AUVs within the current sub-alliance.
3.4. Three-Dimensional Tangent-Plane Obstacle Avoidance
Target search for near-bottom objects in a 3D seafloor environment is similar to target search in a 2D environment. The seafloor surface can be regarded as a curved extension of a 2D search plane. If an AUV moves along the seafloor surface while searching, near-bottom target search can be achieved. Since the seafloor is a curved surface, the AUV is assumed to move along its local tangent plane.
The seafloor within the obstacle-sensing range of an AUV is discretized into obstacle points. An example is shown in
Figure 3, where
Figure 3a presents the detected seafloor terrain and
Figure 3b its discretized representation. The discrete points are treated as the obstacle set, and the Euclidean distance between the AUV and an obstacle is given by Equation (17).
For obstacle avoidance, the AUV considers the nearest obstacle point and two neighboring obstacle points that are not collinear with it. In Equation (18),
denotes the obstacle point nearest to AUV
.
Based on
, two additional obstacle points,
and
, are selected to satisfy Equation (19). Since
,
, and
are not collinear, they determine a plane
. To avoid obstacles while searching for near-bottom targets, the AUV moves parallel to this plane.
The plane
is translated to a new plane
passing through
. A local coordinate system
is then established with
as the origin, under which
,
,
, and
are expressed as in Equation (20).
In the local coordinate system
, the equation of the translated plane
can be assumed as
Then, the normal vector of the plane is given by
, and two vectors in the plane are
From the above equations, the normal vector can be obtained as
When
, the AUV starts obstacle avoidance; when
, the AUV enters an enhanced obstacle-avoidance mode, where
. Let the obstacle-avoidance velocity generated by the 3D tangent-plane obstacle-avoidance algorithm be
and let
. The obstacle-avoidance process under the algorithm is divided into two cases.
- (1)
- (2)
In the coordinate system , the normal vector of plane pointing toward AUV , denoted by , can be obtained according to Equation (27).
Plane
repels AUV
with virtual force
.
The values of and vary with the AUV state. Specifically, and in the roaming search state, while and in the cooperative search state.
3.5. AUV Velocity and Position Update
Once the desired velocity is obtained, the AUV starts obstacle avoidance. The next-step velocity is then determined according to the search state, and the desired velocity . The velocity update is divided into six cases:
- (1)
AUV is in the roaming search state, and .
According to Equations (27) and (28),
is obtained, and
is then given by Equation (29).
- (2)
AUV is in the roaming search state, and .
According to Equation (26),
is obtained, and
is then given by Equation (30).
- (3)
AUV is in the roaming search state, and .
is given by Equation (31).
- (4)
AUV is in the cooperative search state, and .
According to Equations (27) and (28),
is obtained, and
is then given by Equation (32).
- (5)
AUV is in the cooperative search state, and .
According to Equation (26), is obtained, and is then given by Equation (32).
- (6)
AUV is in the cooperative search state, and .
is given by Equation (31).
After is obtained, the AUV position is updated accordingly.
Based on the above analysis, the flowchart and pseudocode of the multi-target search algorithm are shown in
Figure 4 and Algorithm 1, respectively. The convergence of the overall search algorithm primarily relies on the convergence properties of the PSO algorithm [
33], while being secondarily supported by the rapid and complete area exploration capabilities of the roaming search phase. Once target signals are detected, some AUVs transition from roaming search to the convergent cooperative phase, ensuring stable target search.
| Algorithm 1 Multi-Target Search by AUV Swarms |
| 1 | Initialization |
| 2 | while not all targets found do |
| 3 | All AUVs detect target signals |
| 4 | Execute Response-Threshold-based Multi-Target Task Allocation |
| 5 | for each AUV do |
| 6 | if AUV performs collaborative search in sub-alliances then |
| 7 | Calculate according to Equation (16) |
| 8 |
else |
| 9 | Calculate according to Equation (15) |
| 10 |
end if |
| 11 | Calculate by 3D tangent-plane obstacle avoidance |
| 12 | Update AUV position by Equation (33) |
| 13 | if
then |
| 14 | Declare target as found |
| 15 |
end if |
| 16 |
end for |
| 17 | end while |
4. Simulation Results and Analysis
The simulation parameters were set with reference to the literature and practical search requirements, as listed in
Table 2 [
32,
33,
34,
35].
Taking
as an example, the multi-target search process is illustrated. In
Figure 5a, the seafloor terrain for target search is presented.
Figure 5b shows the top view of the search area at
where the AUVs are initially distributed in the
region at the lower left corner, while the targets are located in the central
region.
In
Figure 5c,
detects the signal from target
and AUVs
,
,
,
, and
are recruited into the sub-alliance.
In
Figure 5d, from
to
,
,
,
,
,
,
,
,
,
, and
participate in the sub-alliance to cooperative search for
at different times, with some AUVs joining or leaving during the process. Eventually,
finds
at
.
In
Figure 5e,
is found through cooperative search at
. Up to this time, eight targets,
,
,
,
,
,
,
, and
, have been found.
In
Figure 5f, from
to
, 11 AUVs, namely
,
,
,
,
,
,
,
,
,
, and
, are involved in the cooperative search for
at different times, with dynamic joining and leaving during the process. Before that, multiple AUVs had already participated in the cooperative search for other targets.
In
Figure 5g, all 10 targets are found by the AUV swarm at
.
Figure 5h presents the path of all AUVs during the search process, from which it can be seen that the AUVs successfully avoided the seafloor.
The number of AUVs was set to
and 60, while the number of targets was fixed at
. For each setting, 30 repeated trials were conducted, and the results are summarized in
Table 3. It can be seen that, as the number of AUVs increases, the target search task is completed more quickly, while the energy consumption also increases. These results indicate that an AUV swarm can substantially improve the efficiency of multi-target search.
5. Conclusions
This paper addressed the problem of multi-target search by an AUV swarm in unknown 3D underwater environments with limited communication. A dual-state search mechanism, including cooperative and roaming, was established, and an AUV-swarm-based search method was developed by integrating a virtual force model, motion-constrained 3D Particle Swarm Optimization, and a 3D tangent-plane obstacle-avoidance strategy. The proposed method enables AUVs to switch adaptively between roaming search and cooperative search, form sub-alliances for different targets, and maintain safe motion near the seafloor.
Simulation results verified that the proposed method can effectively accomplish multi-target search in unknown 3D underwater obstacle environments. The results also showed that increasing the swarm size can reduce the search time, although with higher energy consumption. Overall, the proposed method provides a feasible solution for AUV-swarm-based multi-target search with communication constraints.
However, this study primarily focused on multi-target search by AUV swarms in unknown 3D underwater obstacle environments under limited communication conditions, while factors such as heterogeneous swarms, communication uncertainties, and ocean current disturbances were not fully considered. Future work will primarily focus on searching for moving targets, extending to heterogeneous swarms, addressing time-varying ocean currents and communication uncertainties, and conducting real-world experimental validation.