Next Article in Journal
Integration Challenges of Turbine-Powered UAVs: Thermal, Structural, Acoustic, and Operational Perspectives
Previous Article in Journal
Performance of an Efficient Hybrid Dilated–Long Short-Term Memory with Residual Learning for High-Fidelity Electrocardiogram Denoising Signal
Previous Article in Special Issue
An Efficient Odor Source Localization Method for Wheeled Mobile Robots in Indoor Ventilated Environments
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

Low-Altitude Multi-UAV Trajectory Planning in Dynamic Urban Environments Using Dynamic-Aware ACO and MPC-GWO

1
School of Mechanical and Electrical Engineering, North University of China, Taiyuan 030051, China
2
Shanxi Key Laboratory of Machine Vision and Virtual Reality, Taiyuan 030051, China
3
School of Mechanical Engineering, North University of China, Taiyuan 030051, China
4
School of Aerospace Engineering, North University of China, Taiyuan 030051, China
*
Author to whom correspondence should be addressed.
Technologies 2026, 14(8), 454; https://doi.org/10.3390/technologies14080454
Submission received: 1 June 2026 / Revised: 15 July 2026 / Accepted: 21 July 2026 / Published: 23 July 2026
(This article belongs to the Special Issue Advances in the Unmanned System: Control and Autonomous Applications)

Abstract

Low-altitude urban environments pose significant challenges to multi-UAV trajectory planning because of dense buildings, constrained airspace, dynamic obstacles, inter-UAV conflicts, and terminal-area congestion. This study proposes a hierarchical three-dimensional cooperative trajectory-planning framework integrating dynamic-risk-aware Ant Colony Optimization (ACO) with cooperative Model Predictive Control–Gray Wolf Optimizer (MPC-GWO). Environmental costs and predicted dynamic-obstacle risks are incorporated into the ACO global search to generate risk-aware reference trajectories, while a sliding-window GWO improves trajectory smoothness and execution feasibility. During online execution, cooperative MPC-GWO combines dynamic-obstacle prediction, inter-UAV separation constraints, reconfigurable formation switching, and goal-neighborhood safety control to achieve adaptive obstacle avoidance, cooperative replanning, and orderly terminal arrival. Thirty-run Monte Carlo simulations show that the proposed method achieves a success rate of 93.3% ± 25.4% and the highest composite score of 96.20 ± 5.30, with zero dynamic-obstacle and inter-UAV collisions. Ablation experiments verify the effectiveness of the dynamic prediction, formation reconfiguration, and terminal safety-control mechanisms. The average online replanning time remains below 0.5 s, demonstrating satisfactory safety, coordination, adaptability, and real-time performance in small- to medium-scale simulated urban scenarios.

1. Introduction

With the continuous development of information-driven and intelligent warfare, unmanned aerial vehicles (UAVs) have been widely employed in reconnaissance and surveillance, target designation, fire-guidance support, and damage assessment because of their high maneuverability, flexible deployment, and low operational risk. In particular, in scenarios such as urban operations, counterterrorism operations, and localized conflicts, complex mission chains can be accomplished by multi-UAV systems through coordinated cooperation, and their operational effectiveness is largely determined by the quality of trajectory-planning results. Compared with open environments, urban environments are characterized by dense buildings, constrained airspace, severe occlusions, and frequent dynamic threats, which expose UAVs to complex three-dimensional obstacle constraints and increased flight risks during low-altitude flight. Therefore, cooperative trajectory planning for multiple UAVs in complex urban environments has become one of the key technologies in unmanned combat system research [1,2,3].
Extensive research has been conducted on cooperative trajectory planning for multiple UAVs. Recent reviews have systematically summarized UAV path-planning algorithms, multi-UAV planning frameworks, and key technologies in UAV motion planning, suggesting that cooperative trajectory planning under complex constraints has emerged as an important research direction. For multi-UAV cooperative trajectory planning, Li et al. [4] proposed an improved Ant Colony Optimization algorithm for multi-UAV route planning, thereby enhancing path-search quality and demonstrating the applicability of ACO to multi-UAV scenarios. Pehlivanoglu et al. [5] developed an efficient path-planning approach for autonomous multi-UAV systems in target-coverage problems, highlighting the broad need for cooperative multi-agent planning in complex mission scenarios. Yi [6] proposed a reinforcement-learning-driven continuous ACO method for multi-UAV path planning in complex environments, which further improved the adaptability of swarm-intelligence-based planning methods. Lu et al. [7] proposed a conflict-free three-dimensional path-planning method for multiple UAVs based on jump point search and incremental updating, providing useful insights into three-dimensional cooperative planning and conflict resolution. In addition, representative studies on UAV coverage path planning, three-dimensional path optimization, and learning-based multi-UAV planning have further enriched the modeling approaches and solution strategies used for UAV path planning in urban and multi-constrained environments [8,9,10].
In recent years, multi-agent reinforcement learning (MARL) has been increasingly used for UAV swarm cooperation, formation control, and collision avoidance. Regarding dynamic-obstacle avoidance and adaptation to complex environments, Chang et al. [11] integrated an improved dynamic window approach with optimal reciprocal collision avoidance for autonomous obstacle avoidance by multiple UAVs, demonstrating the effectiveness of combining local motion planning with reciprocal collision-avoidance mechanisms. Yan et al. [12] proposed a deep-reinforcement-learning-based real-time path-planning method for UAVs in dynamic environments, improving the adaptability of online planning under moving-obstacle constraints. Wang et al. [13] developed a two-stage reinforcement-learning approach for multi-UAV collision avoidance under imperfect sensing, highlighting the importance of robust cooperative avoidance when perception information is incomplete or uncertain. AlMahamid and Grolinger [14] systematically reviewed reinforcement-learning-based autonomous UAV navigation methods, providing a reference for the modeling of adaptive navigation and obstacle-avoidance strategies in complex environments. Xue and Chen [15] proposed a multi-agent deep-reinforcement-learning method for UAV navigation in unknown complex environments, demonstrating the potential of distributed learning policies for cooperative decision-making and collision avoidance. Yan et al. [16] further investigated fixed-wing UAV flocking in continuous spaces using deep reinforcement learning, offering useful insights into learning-based formation coordination and collision-avoidance policy design. Lin et al. [17] proposed a sampling-based path-planning method for UAV collision avoidance, offering an additional solution for collision avoidance in dynamic airspace. Although these studies have made meaningful progress in path generation and dynamic-obstacle avoidance, these learning-based methods usually require large-scale training data, carefully designed reward functions, and additional generalization validation. In contrast, the proposed method focuses on an interpretable optimization-control framework that integrates dynamic-risk-aware global planning, local smoothing, and online replanning without offline policy training.
Moreover, the cooperative flight of multiple UAVs involves not only inter-UAV collision avoidance but also formation maintenance, formation reconfiguration, and cooperative control. Ma et al. [18] proposed a multi-UAV formation obstacle-avoidance method by combining improved simulated annealing with an adaptive artificial potential field, demonstrating that adaptive formation adjustment can be effective in obstacle-rich environments. Liu et al. [19] achieved flexible multi-UAV formation control by integrating deep-reinforcement-learning and affine transformations, thereby improving the adaptability of formation transformation under complex conditions. Li et al. [20] investigated multi-UAV obstacle avoidance and formation control in unknown environments, highlighting the importance of balancing obstacle avoidance with formation maintenance. Liu et al. [21] studied flocking navigation and obstacle avoidance for multi-UAV systems using a hierarchical weighting Vicsek model, providing useful insights into swarm coordination and distributed obstacle avoidance. Choi et al. [22] developed a bearing-based distributed control method for UAV formation tracking and obstacle avoidance, further supporting the integration of formation-tracking and collision-avoidance capabilities. These studies indicate that formation coordination has become a critical direction for improving mission efficiency and organizational capability in multi-UAV systems. However, in highly dynamic urban environments, the balance among formation maintenance, local formation relaxation, terminal-stage formation release, and safe arrival remains insufficiently investigated.
Despite these advances, several limitations remain in cooperative multi-UAV applications in highly dynamic urban environments. First, although existing studies have investigated UAV path planning in three-dimensional environments, dynamic-obstacle avoidance, and adaptive navigation, dynamic obstacles are still often treated through local avoidance, current-state response, or limited short-term prediction. As a result, it remains difficult to fully characterize complex motion behaviors, nonlinear obstacle evolution, and dynamic occupancy in terminal regions [23]. Second, although ACO, GWO, reinforcement-learning-based GWO, and other intelligent optimization methods have been widely applied to UAV and UCAV path planning, most existing studies mainly focus on improving global search capability, convergence performance, or path smoothness, while the coupling among dynamic-risk perception, multi-UAV cooperative constraints, and terminal-area conflict resolution remains insufficiently addressed [24]. Third, MPC-based rolling optimization provides an effective framework for constrained online decision-making, and dynamic adaptive path-planning methods can improve the adaptability of UAVs in three-dimensional environments; however, when dynamic perception, formation reconfiguration, inter-UAV separation, and terminal safe arrival are considered simultaneously, the computational burden of online replanning may increase, making it difficult to balance real-time performance and engineering practicability [25,26,27]. Therefore, it is meaningful to develop an interpretable multi-level cooperative trajectory-planning framework that integrates dynamic-aware global optimization, local online adjustment, safe terminal arrival, and reconfigurable formation coordination for small-scale dynamic urban scenarios.
To address these limitations, a three-dimensional trajectory-planning method for multiple UAVs is proposed by integrating a dynamic-aware ACO algorithm with cooperative MPC-GWO. First, a three-dimensional urban environment model is constructed, and environmental cost and dynamic-risk assessment terms are incorporated into the ACO-based global planning stage to enhance the ability of the initial trajectories to proactively avoid high-risk regions. Subsequently, a sliding-window GWO is employed to smooth and locally optimize the trajectories, thereby improving trajectory continuity and execution feasibility. Finally, during online execution, cooperative MPC-GWO, dynamic-obstacle prediction, reconfigurable formation, and goal-neighborhood safety-control mechanisms are integrated to achieve online obstacle avoidance, inter-UAV separation, formation switching, and orderly arrival in highly dynamic environments. The main contributions of this study are summarized as follows:
First, a dynamic-risk-aware ACO-based global planning strategy is developed within the proposed multi-level framework. In this strategy, urban environmental cost, obstacle risk, and trajectory length are incorporated into the path-search process to improve the safety and adaptability of initial trajectories in dynamic urban environments.
Second, a sliding-window GWO-based local optimization mechanism is employed to smooth and refine local trajectories while preserving the global trajectory structure, thereby improving trajectory continuity and UAV execution feasibility.
Third, a cooperative MPC-GWO-based online replanning strategy is designed to enhance local trajectory adjustment during execution. Dynamic-obstacle prediction, inter-UAV safety-distance constraints, and mission-synchronization requirements are embedded into a receding-horizon optimization framework to improve cooperative obstacle-avoidance capability under dynamic disturbances.
Finally, a reconfigurable formation and goal-neighborhood safety-control mechanism is introduced for constrained passage and terminal arrival. According to channel width, local risk, and goal-area congestion, the UAV team adaptively switches formation structures and terminal-arrival strategies, thereby improving passage capability in narrow spaces and reducing terminal conflicts.
The remainder of this paper is organized as follows. Section 2 formulates the three-dimensional trajectory-planning problem for multiple UAVs, including urban environment modeling, dynamic-obstacle modeling, and multi-UAV safety constraints. Section 3 presents the proposed trajectory-planning method based on dynamic-aware ACO and cooperative MPC-GWO. Section 4 validates the effectiveness of the proposed method through comparative and ablation experiments. Section 5 discusses experimental results, method applicability, and limitations. Section 6 concludes this study and outlines future research directions.

2. Materials and Methods

This study focuses on cooperative trajectory planning for multiple UAVs in low-altitude urban environments with static buildings, dynamic obstacles, and inter-UAV safety constraints. The objective is to generate safe, smooth, and cooperatively consistent trajectories while considering environmental risk, dynamic-obstacle avoidance, formation coordination, and terminal-arrival safety.

2.1. Problem Description

Urban environments are characterized by densely distributed buildings, significant height variations, and constrained airspace, and UAV flights may also be affected by dynamic obstacles and inter-UAV conflicts. Therefore, trajectory planning should not only satisfy the safety requirements of individual UAVs but also account for cooperative constraints and mission efficiency among multiple UAVs.
Assume that the system consists of  M  UAVs, and the trajectory of the  i -th UAV is represented by a sequence of discrete trajectory points as follows:
P i = p i 1 , p i 2 , , p i N i
where  p i k = x i k , y i k , z i k  denotes the three-dimensional position of the  i -th UAV at the  k -th trajectory point, and  N i  denotes the number of nodes contained in its trajectory.
The trajectory of the  i -th UAV is represented as a sequence of discrete three-dimensional waypoints. To reduce the computational burden of a full 3D search, the proposed method first performs path search on a two-dimensional traversable grid and then assigns flight altitude according to local building height and safety clearance. In addition, three formation modes, namely triangular formation, column formation, and free-decoupled mode, are introduced to adapt to open areas, narrow corridors, and high-risk or terminal regions, respectively.

2.2. Environmental Modeling, Constraints, and Objective Function

In this study, the planning region is represented by an urban-building-height field. The planning region is uniformly rasterized and discretized into a two-dimensional grid map. The urban-building-height field is denoted by  H x , y . This height distribution is further quantified through the environmental cost function. To ensure safe low-altitude flight, the altitude of each UAV at any trajectory waypoint is required to exceed the local building height by a prescribed safety margin, which can be expressed as:
z i k = max ( x , y ) N ( x i k , y i k ) H ( x , y ) + h safe
where  h safe  is the minimum safe flight clearance. This height-field-based representation preserves low-altitude safety while reducing the computational complexity of direct 3D trajectory search.
To comprehensively evaluate trajectory quality, a multi-objective optimization function is constructed by considering environmental cost, path length, smoothness, dynamic risk, inter-UAV separation, mission synchronization, formation keeping, and infeasibility penalties:
J i = ω 1 J env + ω 2 J len + ω 3 J smooth + ω 4 J risk + ω 5 J sep + ω 6 J sync + ω 7 J form + ω 8 J inf
where  ω i  denote the weighting coefficients. The corresponding cost terms account for building-height risk, path length, trajectory smoothness, dynamic-obstacle interference, inter-UAV conflicts, arrival-time consistency, formation deviation, and infeasible-solution penalties.
Because different cost terms have different physical meanings and numerical scales, all objective terms are normalized before weighted aggregation. The normalized cost term is defined as:
J ¯ k = J k J k min J k max J k min + ϵ
where  J k  represents a specific cost term,  J k min  and  J k max  denote the minimum and maximum values observed in preliminary simulations, and  ϵ  is a small positive constant used to avoid division by zero.
Spatial feasibility, altitude safety, inter-UAV collision avoidance, and formation-switching constraints are imposed to ensure that the planned trajectories are executable in complex urban environments. For spatial feasibility, each waypoint must remain in the feasible region, and each segment connecting adjacent waypoints must avoid buildings, no-fly zones, and high-risk regions. Let  Ω free  and  Ω obs  denote the feasible and infeasible regions, respectively. The constraint is given by:
( x i k , y i k ) Ω free , p i k p i k + 1 ¯ Ω obs =
Altitude safety is enforced by requiring each UAV to fly above the local building height with a predefined safety margin:
H ( x i k , y i k ) + h safe z i k z max
Inter-UAV safety is ensured by imposing a minimum separation between any two UAVs at each time instant:
p i ( t ) p j ( t ) 2 d safe
where  p i t  and  p j t  denote the positions of the  i -th and  j -th UAVs at time  t .
To balance cooperative efficiency and flight safety, three formation modes are adopted: triangular formation, column formation, and free-flight mode. The formation mode is switched according to the environmental risk level as follows:
F ( t ) = F tri , W ( t ) W tri , R ( t ) < R low F line , W line W ( t ) < W tri , R ( t ) < R high F free , R ( t ) R high   or   c ( t ) G R g
where  F tri F line , and  F free  denote the triangular formation, column formation, and free-release mode, respectively.  W tri  and  W line  are the channel-width thresholds required to maintain the triangular and column formations, respectively, whereas  R low  and  R high  are the thresholds for low- and high-risk assessment, respectively.

2.3. Dynamic-Obstacle Prediction and Goal-Region Risk Modeling

Dynamic-obstacle and goal-neighborhood risks are further considered to improve urban flight safety. The future position of each dynamic obstacle is predicted from its current state by incorporating velocity variation, acceleration effects, and stochastic maneuvering behavior. The predicted position of the  i -th obstacle at time  t + Δ t  is formulated as follows:
P ^ q ( t + τ ) = P q ( t ) + v q ( t ) τ + 1 2 a q t τ 2 + ξ q t , 0 τ Δ t
where  P q t = ( x q ( t ) ,   y q ( t ) ,   z q ( t ) )  denotes the position of the  q -th dynamic obstacle, and  v q ( t )  represents its velocity vector. The additional terms account for acceleration-induced motion variation and stochastic maneuvering uncertainty.
To evaluate dynamic-obstacle risk, a spatiotemporal risk region is constructed around each predicted obstacle trajectory. Let  r o  and  r s  denote the obstacle radius and safety buffer radius, respectively. The dynamic-risk region is defined as:
R q ( t + τ ) = ( x , y , z ) ( x , y , z ) P ^ q ( t ) r q + r s
A potential collision risk is detected if any UAV waypoint or trajectory segment enters a predicted risk region within the prediction horizon. Accordingly, the dynamic-risk cost is formulated to quantify the corresponding risk exposure:
J dyn = i = 1 M k = 1 N i q = 1 Q Φ p i k , R q ( t k + τ )
Furthermore, a goal-neighborhood risk region is defined to account for the limited maneuvering space, strict task-completion requirement, and low collision tolerance near the target:
Ω g = ( x , y , z ) ( x , y , z ) p g R g
where  p g  is the target position, and  R g  is the radius of the terminal neighborhood.

3. Methods

To improve global reachability, dynamic-obstacle adaptability, and formation coordination in highly dynamic urban environments, a hierarchical cooperative trajectory planning method is proposed. As shown in Figure 1, the method consists of three layers: global path search, local path refinement, and online cooperative replanning.
As shown in Figure 1, the proposed framework consists of three layers. The global layer uses dynamic-aware ACO to generate risk-aware initial paths. The local refinement layer applies sliding-window GWO to improve trajectory smoothness while preserving global topology. The execution layer integrates cooperative MPC-GWO, reconfigurable formation switching, and terminal safety control to handle dynamic obstacles, inter-UAV conflicts, and goal-region congestion.

3.1. Dynamic-Aware ACO Initial Trajectory Generation

To improve global reachability in complex urban environments, a dynamic-aware ACO algorithm is adopted at the global planning layer. Compared with conventional ACO methods that mainly rely on path length and pheromone intensity, the proposed method incorporates building-height constraints and predicted dynamic-obstacle risk into the heuristic function, thereby enabling safer global path search. The dynamic-aware ACO-based search is formulated as follows:
η m n = 1 d v , g + ε 1 C env v 1 1 + λ R dyn v
In this equation,  d v , g  denotes the distance from the candidate node to the target node,  C env v  represents the environmental cost,  R dyn v  is the dynamic-risk evaluation value within the prediction horizon, and  λ  is the risk-weighting coefficient. This design enables the ants to maintain a goal-oriented search tendency while actively avoiding high-risk regions during the search process.
In the pheromone update stage, an elite-path reinforcement strategy is adopted as follows:
τ m n t + 1 = 1 ρ τ m n t + Q L best
where  ρ  is the pheromone evaporation coefficient,  Q  is the pheromone constant,  L best  denotes the length of the best path in the current iteration, and  P best  represents the set of currently superior paths. After iterative search, an initial global path that satisfies obstacle-avoidance, flight-safety, and dynamic-risk constraints is obtained, as shown in Figure 2.
In Figure 2, the gray regions denote building obstacles, the color heat map represents the pheromone concentration at path nodes, with brighter colors indicating higher concentrations, and the green and red points denote the start and goal positions, respectively.
During the initial exploration stage, ants search randomly within the feasible region, resulting in a dispersed pheromone distribution, as shown in Figure 2a. As the iterations proceed, candidate paths with shorter lengths, lower environmental costs, and smaller dynamic risks receive stronger pheromone reinforcement, as shown in Figure 2b. At convergence, the pheromone is concentrated around the globally optimal or near-optimal paths, while inferior paths gradually decay due to pheromone evaporation, thereby generating an initial reference trajectory that satisfies safety and global-reachability requirements, as shown in Figure 2c.

3.2. Trajectory Optimization Based on Sliding-Window GWO

Although the dynamic-aware ACO provides a globally reachable initial path, grid discretization may cause redundant segments, sharp turns, and poor smoothness. To improve trajectory feasibility, a sliding-window GWO algorithm is applied for local path refinement. The initial path is divided into fixed-length local segments, and the internal nodes of each segment are optimized by GWO while the global path topology is preserved. For the m-th sliding window, the local path segment is defined as:
P i m = { p i m , p i m + 1 , , p i m + L 1 }
where  L  denotes the window length. This strategy prevents local refinement from destroying the connectivity and topological consistency of the global path.
In the local path-refinement stage, path length, smoothness, environmental risk, and infeasibility penalties are used as the main optimization objectives. The objective function can be expressed as follows:
J gwo = ω 1 J 1 + ω 2 J s + ω 3 J e + ω 4 J p
where  J 1  denotes the path-length cost,  J s  represents the smoothness cost,  J e  is the environmental risk cost,  J p  denotes the infeasibility penalty term, and  ω  represents the corresponding weighting coefficients.
During local refinement, infeasible candidates are penalized when waypoints fall inside obstacle regions or adjacent path segments intersect infeasible areas, thereby preserving spatial feasibility. The sliding-window strategy avoids the computational burden of full-path optimization while reducing sharp turns and local redundancies. As a result, a smoother and more executable reference trajectory is obtained for subsequent tracking and online adjustment.

3.3. Online Replanning and Reconfigurable Formation Control Based on Cooperative MPC-GWO

After global search and local refinement, a reference trajectory is obtained. However, highly dynamic obstacles, inter-UAV conflicts, and terminal congestion may still cause local path failure during flight. Therefore, an online cooperative replanning mechanism is introduced in the execution stage, where cooperative MPC-GWO, formation reconfiguration, and goal-neighborhood safety control are integrated to adjust UAV trajectories in a receding-horizon manner.
Local replanning is triggered when the next node or path segment is occupied by a predicted dynamic-obstacle region, when safety separation from another UAV is violated, or when the formation error exceeds the allowable threshold. Once triggered, a local reference segment is generated within a finite prediction horizon, and cooperative MPC-GWO is applied to optimize the trajectory. The corresponding local cost function is formulated as follows:
J local = ω c J cost + ω s J smooth + ω l J len + ω d J dev + ω sep J sep + ω sync J sync + ω f J form
where  J cost J smooth J len J dev J sep J sync , and  J form  denote the local environmental cost, smoothness cost, path-length cost, reference-deviation cost, inter-UAV separation cost, cooperative-arrival cost, and formation-keeping cost under the current formation mode, respectively.
To improve online efficiency, a receding short-window strategy is adopted to optimize only the blocked local region rather than regenerate the entire global path. The previous local solution is used to initialize the current optimization, reducing trajectory oscillations and improving stability, as shown in Figure 3a. Near the goal, fixed formation keeping is relaxed to avoid repeated yielding and local congestion under limited maneuvering space. If the goal region is occupied by dynamic obstacles or other UAVs, the system either waits briefly or inserts a temporary transition waypoint for local detouring. Once the goal becomes reachable, the UAVs approach it sequentially, as shown in Figure 3b.
For formation coordination, three modes are considered: triangular formation, column formation, and free-decoupled mode, as shown in Figure 4. The triangular formation is maintained in open and low-risk areas to preserve cooperative consistency, while the column formation is adopted in narrow but controllable corridors to reduce lateral space occupation. When dynamic risk is high, inter-UAV conflicts become significant, or UAVs approach the terminal neighborhood, the system switches to the free-decoupled mode, where only the inter-UAV safety-distance constraint is retained.
Blue circles denote UAVs, the red circle indicates the reference UAV, solid lines represent formation constraints, dashed circles indicate safety boundaries, and arrows show the mission direction. Overall, the execution layer combines receding-window optimization, formation switching, dynamic-obstacle avoidance, and goal-neighborhood safety control to achieve real-time trajectory adjustment and safe cooperative arrival in highly dynamic environments. It inherits the global guidance and trajectory smoothness provided by the previous layers while improving adaptability to dynamic obstacles, inter-UAV conflicts, and terminal congestion.

3.4. Computational Complexity Analysis

The computational complexity of the proposed framework is analyzed according to its three main stages: dynamic-aware ACO-based global planning, sliding-window GWO-based local trajectory refinement, and cooperative MPC-GWO-based online replanning. Before deriving the computational complexity of each algorithmic stage, the main notations used in the following analysis are summarized in Table 1.
In the global planning stage, the main computational burden arises from iterative node expansion and candidate-node evaluation, where pheromone intensity, goal distance, environmental cost, and predicted dynamic-obstacle risk are jointly considered. Thus, the complexity of the dynamic-aware ACO stage can be approximated as:
O ACO = O M I a N a L p b ( 1 + Q H p )
The term  M I a N a L p b  corresponds to the basic ACO search process, whereas  Q H p  reflects the additional cost of evaluating the predicted dynamic-obstacle risk over the local prediction horizon. If the predicted dynamic-risk map is precomputed and cached over the grid map, an additional preprocessing cost of  O Q H p G  is introduced, while candidate-node risk evaluation can be reduced to grid-value lookup during ACO search.
In the local trajectory-refinement stage, if the initial ACO path is divided into multiple sliding windows, and the internal waypoints of each window are optimized using GWO, the computational complexity of the sliding-window GWO stage is:
O GWO = O M K w I g N g L w
This complexity is lower than that of full-path optimization because only local windows are optimized while the global path topology is preserved. Therefore, the sliding-window strategy reduces the search dimension and improves computational efficiency while maintaining trajectory smoothness and spatial feasibility.
In the online replanning stage, cooperative MPC-GWO optimizes local trajectory segments in a receding-horizon manner. For each candidate solution, the algorithm evaluates dynamic-obstacle risk, inter-UAV separation, reference tracking, formation keeping, and cooperative-arrival costs within the prediction horizon. The dynamic-obstacle evaluation introduces a complexity proportional to  M Q H p , while pairwise inter-UAV safety-distance evaluation introduces a complexity proportional to  M 2 Q H p . Therefore, the computational complexity of one online replanning cycle is approximately:
O online = O ( I m N m H p ( M Q + M 2 ) )
If online replanning is triggered  T  times during the mission, the total online replanning complexity becomes:
O online , total = O ( T I m N m H p ( M Q + M 2 ) )
This expression shows that the online computational burden is mainly affected by the number of UAVs, the number of dynamic obstacles, and the prediction horizon.

4. Results

A 3D urban simulation scenario is constructed in MATLAB R2023b to validate the proposed multi-UAV reconfigurable-formation trajectory planning method. Overall comparisons and three ablation studies are conducted to evaluate its comprehensive performance and the contributions of dynamic-risk prediction, reconfigurable formation, and goal-neighborhood safety control.

4.1. Experimental Setup

A three-dimensional gridded urban height-field model with a planning area of (100 × 100 × H) m was constructed in MATLAB. The baseline scenario included 45 buildings with heights of 3–40 m and mixed geometries, three UAVs with start and goal points distributed in different quadrants, and five highly dynamic obstacles with velocity variation, random turning, and sudden maneuvers. Additional scenarios with different UAV numbers, obstacle densities, and map sizes were designed to evaluate robustness and scalability. For fair comparison, all algorithm parameters were tuned through preliminary experiments, and the final settings are listed in Table 2.
To reduce the influence of stochastic dynamic-obstacle motions, each experiment was repeated 30 times with different random seeds. In each run, the initial states and maneuvering disturbances of dynamic obstacles were randomized, while the UAV start and goal positions were kept fixed for fair comparison. The results are reported as mean ± standard deviation. Statistical significance was tested using the Kruskal–Wallis test followed by Mann–Whitney U tests with Bonferroni correction, and the significance level was set to p < 0.05.

4.2. Overall Performance Experiment

Six comparative methods were employed to evaluate the proposed approach. A1 denotes conventional ACO, which performs static global path planning without trajectory smoothing or online replanning. A2 combines ACO with GWO-based trajectory smoothing but does not include online obstacle avoidance or cooperative replanning. A3 employs MPC-based local replanning for dynamic-obstacle avoidance while following the reference trajectories generated by ACO and GWO. A4 integrates ACO and GWO with MADDPG, where the ACO+GWO trajectories serve as global references, and MADDPG generates local action corrections for dynamic-obstacle avoidance and inter-UAV coordination. A5 adopts the same ACO+GWO-guided planning framework but uses MAPPO to generate cooperative local action corrections. A6 represents the proposed method, which integrates dynamic-aware ACO, cooperative MPC-GWO replanning, reconfigurable formation control, and goal-neighborhood safety coordination. The trajectories generated by the six methods are presented in Figure 5.
As shown in Figure 5, A1 and A2 generate feasible static reference trajectories but lack dynamic-obstacle perception and inter-UAV coordination. A3 improves dynamic-obstacle avoidance through MPC-based local replanning; however, inter-UAV conflicts may still occur because explicit cooperative decision-making is not incorporated. A4 follows the ACO+GWO reference trajectories and employs MADDPG-based local action correction, thereby improving dynamic avoidance and cooperative motion, although its performance remains sensitive to the learned local policy and lacks explicit global dynamic-risk modeling. Compared with A4, A5 achieves a higher mission success rate and fewer dynamic-obstacle collisions. However, its minimum inter-UAV distance is smaller, and occasional inter-UAV collisions still occur. In contrast, A6 integrates dynamic-aware ACO, cooperative MPC-GWO replanning, reconfigurable formation control, and goal-neighborhood safety coordination, producing the safest and most coordinated trajectories. The comprehensive performance comparison of the six multi-UAV cooperative planning methods is presented in Table 3.
Table 3 summarizes the results of 30 Monte Carlo runs. A1 and A2 achieve high reached rates but low success rates because they lack online dynamic-obstacle avoidance and inter-UAV coordination. A3 improves the success rate to 86.7%. However, it primarily performs reactive local optimization and does not explicitly incorporate global dynamic-risk assessment, formation reconfiguration, or terminal coordination. A4 and A5 improve coordination through learned action correction but require extensive offline training and hyperparameter tuning. In contrast, A6 requires no policy training and achieves the highest success rate of 93.3% and composite score of 96.20, with zero dynamic and inter-UAV collisions and the largest minimum safety distances, demonstrating superior safety, coordination, reliability, and interpretability.
The overall performance experiment in the three-dimensional scenario verifies the advantages of the proposed method in multi-UAV cooperation and dynamic-obstacle avoidance in complete spatial environments. However, complex factors in three-dimensional environments, such as height variations and multi-directional obstacle distributions, may jointly affect the performance of different core modules, making it difficult to independently evaluate the contributions of dynamic-obstacle prediction, reconfigurable formation, and goal-neighborhood safety control. To more clearly reveal the mechanisms of these modules and reduce the interference of incidental environmental factors, the following ablation experiments are conducted by projecting the experimental scenario onto a two-dimensional horizontal plane. The two-dimensional scenario retains the main lateral obstacle-avoidance and formation-coordination requirements of low-altitude urban flight while simplifying the analysis dimension, making the performance differences among modules more intuitive and distinguishable.

4.3. Ablation Experiment Analysis

For low-altitude urban flights, more challenging situations arise from the continuous disturbances caused by highly dynamic obstacles and from local congestion or goal occupation near the goal region. Such scenarios may not only lead to frequent local trajectory failures but also damage the formation structure and affect safe terminal arrival. Therefore, ablation experiments were further conducted to verify the core mechanisms of the proposed method and evaluate its robustness and terminal-stage safety-control capability in highly dynamic environments.

4.3.1. Ablation of the Dynamic-Obstacle Prediction and Risk Assessment Model

To evaluate the dynamic-obstacle prediction and risk assessment model, three variants were compared, with the remaining modules unchanged. B1 uses only current obstacle positions for reactive avoidance, B2 applies constant-velocity extrapolation, and B3 adopts the proposed dynamic prediction and time-domain risk assessment. Figure 6 compares their local trajectory adjustments in a complex environment with static buildings and high-speed dynamic obstacles.
As shown in Figure 6, the reactive obstacle-avoidance strategy of B1 lacks spatiotemporal foresight and triggers avoidance only after the UAV approaches an obstacle. This results in several sharp turns and poor trajectory smoothness. Although B2 adopts linear extrapolation prediction, its prediction results rapidly become invalid when dynamic-obstacles perform sudden turns or nonlinear maneuvers. Consequently, the UAV is forced to perform an almost right-angle emergency turn in the central danger region, increasing attitude-control cost and potential collision risk. In contrast, B3 relies on the dynamic prediction and time-domain risk assessment mechanism to identify obstacle-crossing trends in advance, thereby achieving smoother trajectory adjustment with a smaller detour range in complex intersection regions. The dynamic-obstacle avoidance performance comparison is presented in Table 4.
As shown in Table 4, the proposed B3 strategy achieved the highest success rate and the lowest collision frequency under the high-density dynamic-obstacle scenario. Compared with B1 and B2, B3 increased the success rate from 66.67% to 94.44%, while reducing the average collision frequency to 0.03 per run. In addition, B3 reduced the average risk cost by 90.00% and 93.75% compared with B1 and B2, respectively. The average path length and smoothness cost were also reduced, indicating that the dynamic-aware prediction strategy not only improved safety but also generated shorter and smoother trajectories. The runtime of B3 remained comparable to the baseline methods, suggesting that the improved prediction strategy did not introduce significant additional computational burden.

4.3.2. Ablation of the Reconfigurable Formation-Coordination Mechanism

To evaluate reconfigurable formation control in narrow urban spaces, two variants were compared. C1 maintains a fixed triangular formation with rigid obstacle avoidance, whereas C2 adaptively switches among triangular, column, and free-flight modes according to local constraints and risk levels. Figure 7 shows their cooperative obstacle-avoidance trajectories in a narrow-passage scenario.
As shown in Figure 7a, the fixed triangular formation with rigid shifting in C1 maintains the triangular formation throughout the motion process and cannot adjust the formation structure according to the channel width. Therefore, continuous static collisions occur in the narrow-channel region. As shown in Figure 7b, the reconfigurable formation method C2 can switch from a triangular formation to a column formation according to local spatial constraints, thereby reducing the lateral width of the formation and enabling all three UAVs to safely pass through the narrow channel. Table 5 presents the comparison of formation-coordination performance indicators between the two methods.
As shown in Table 5, the fixed triangular formation in C1 results in a large lateral footprint, preventing the UAV group from passing through the narrow channel and leading to a passage success rate of 0.00%. In contrast, C2 adaptively switches among triangular, column, and free-flight modes, increasing the success rate to 100.00%. It also reduces the average risk cost and passing time, with only a slight runtime increase from 0.08 s to 0.09 s. Although active reconfiguration introduces a formation error of 4.50 m, the minimum inter-UAV distance remains 5.00 m, satisfying the safety-separation requirement. These results confirm the effectiveness of the proposed reconfigurable formation strategy in constrained urban environments.

4.3.3. Ablation of the Goal-Neighborhood Safety-Control Strategy

Three variants were compared to evaluate goal-neighborhood safety control. D1 uses a direct goal approach with collision detection only, D2 applies local detouring without waiting or formation release, and D3 adopts the proposed coordinated mechanism integrating waiting, detouring, and formation release. Figure 8 shows the corresponding UAV trajectory distributions near the goal region.
As shown in Figure 8, D1 directly guides UAVs to the goal without goal-neighborhood coordination, making conflicts likely near the target. D2 reduces some conflicts through local detouring, but trajectory congestion remains because waiting and formation release are not included. In contrast, D3 combines waiting, detouring, and formation release, allowing UAVs to avoid conflicts temporarily and enter the goal region in an orderly manner. These results show that D3 improves safe-arrival capability in complex dynamic environments. Table 6 further provides the statistical comparison of dynamic-obstacle avoidance and goal-region conflict handling.
As shown in Table 6, D1 and D2 fail to achieve safe terminal arrival in the tested goal-neighborhood scenario. Although D2 introduces detouring, it lacks waiting and formation-release control, making it unable to avoid conflicts when dynamic obstacles occupy the goal region. In contrast, D3 achieves a success rate of 96.67%, reduces the collision frequency to 0.05, and decreases the terminal conflict value to 0.00. Although additional waiting and slight detours are introduced, D3 shortens the team arrival time from the maximum simulation time of 25.00 s to 12.13 s, demonstrating improved terminal safety and mission-completion efficiency.

4.4. Parameter Sensitivity Analysis

To evaluate the influence of key algorithm parameters on the proposed framework, a parameter sensitivity analysis was conducted. Three representative parameters were selected, including the dynamic-risk weight ( ω r ), the prediction horizon ( H p ), and the formation-maintenance weight ( ω f ). These parameters were considered because they directly affect dynamic-obstacle avoidance, global search behavior, online replanning performance, and formation consistency.
A one-factor-at-a-time strategy was adopted, in which one parameter was varied while the others were fixed at their baseline values. For each setting, 30 Monte Carlo runs were conducted, and the mean performance metrics were recorded. The tested ranges are summarized in Table 7.
Figure 9a shows that increasing the risk weight reduces the risk cost but may increase path length because of more conservative detours. Figure 9b indicates that a longer prediction horizon lowers the hazardous-time ratio by improving conflict prediction, although it increases replanning runtime. Figure 9c shows that a larger formation-maintenance weight reduces formation error, but excessive weighting limits maneuvering flexibility and provides limited success-rate improvement. Overall, the final parameters were selected to balance safety, path efficiency, formation consistency, and computational feasibility.

4.5. Scalability and Real-Time Computational Performance

To evaluate the scalability of the proposed method, four mission-scale scenarios were constructed by progressively increasing the number of UAVs, the number of dynamic obstacles, the number of buildings, and the map size from the baseline 100 × 100 × H m scenario. The purpose of this setting is to separately examine the influence of UAV-team size, dynamic-obstacle density, and environmental scale on planning performance and computational efficiency. Each scenario was repeated 30 times with different random seeds, and the results are reported as mean ± standard deviation. The scalability settings are listed in Table 8, the detailed numerical results are summarized in Table 9, and the corresponding scalability and real-time computational trends are illustrated in Figure 10.
The real-time feasibility of online replanning was evaluated by comparing the average online replanning time with the predefined replanning interval. In this study, the replanning interval was set to 0.5 s. A scenario was considered real-time feasible when  T ¯ o n l i n e < T i n t e r v a l , where  T ¯ o n l i n e  denotes the average online replanning time, and  T i n t e r v a l = 0.5   s  denotes the predefined replanning interval.
As shown in Table 8 and Figure 10, the proposed method maintains stable scalability under the tested mission-scale scenarios. Compared with S1, S2 increases the number of UAVs from three to five while keeping the map size, building number, and dynamic-obstacle number unchanged. The increased UAV number introduces more pairwise inter-UAV separation constraints and cooperative coordination requirements, resulting in a moderate increase in online replanning time.
Compared with S2, S3 further increases the number of dynamic obstacles from five to ten. As a result, more predicted risk regions need to be evaluated within the local prediction horizon, and the probability of UAVs entering hazardous regions increases. Therefore, the hazardous-time ratio increases in S3. Nevertheless, the proposed dynamic-risk prediction and cooperative replanning mechanisms still maintain a high mission success rate and suppress dynamic-obstacle collisions.
Compared with S3, S4 increases the map size from 100 × 100 m to 200 × 200 m and the number of buildings from 45 to 80. This mainly enlarges the feasible grid-search space and increases the burden of dynamic-aware ACO global planning. Therefore, the global planning time increases more significantly than the online replanning time.
In all tested scenarios, the average online replanning time remains below the predefined replanning interval of 0.5 s, indicating that the proposed method satisfies the real-time requirement in the tested small- to medium-scale scenarios. Although the computational time increases with mission complexity, the proposed framework still maintains acceptable scalability due to the combination of global dynamic-aware search, sliding-window local refinement, and receding-horizon online replanning.
The measured computational trends are consistent with the theoretical complexity analysis in Section 3.4. Specifically, the global planning time mainly increases with map size and building density, whereas the online replanning time is more strongly affected by the number of UAVs and dynamic obstacles. For larger UAV swarms and denser urban environments, the computation time increases noticeably, and further acceleration through parallel computation, or distributed replanning would be required.

5. Discussion

The experimental results indicate that the proposed dynamic-aware ACO and cooperative MPC-GWO framework can effectively improve trajectory safety, cooperative obstacle avoidance, and execution feasibility in the tested simulated urban scenarios. Compared with conventional ACO and ACO+GWO methods, the proposed method incorporates dynamic-obstacle prediction, inter-UAV safety constraints, and mission-synchronization requirements, thereby improving adaptability under dynamic disturbances. The dynamic-risk-awareness mechanism reduces the probability of UAVs entering high-risk regions, while the sliding-window GWO improves trajectory smoothness and provides a more stable reference for online replanning.
During execution, cooperative MPC-GWO performs receding-horizon optimization according to dynamic obstacles and inter-UAV relative states, enabling each UAV to balance individual safety and group cooperation. The reconfigurable formation mechanism improves passage capability in narrow or congested regions, and the goal-neighborhood safety-control strategy reduces terminal conflicts caused by simultaneous arrival. These results suggest that multi-UAV trajectory planning should consider not only path efficiency but also dynamic risk, formation adaptability, and orderly terminal arrival.
This study still has several limitations. The proposed method was validated mainly in MATLAB simulations rather than real-flight or hardware-in-the-loop experiments; therefore, perception uncertainty, localization error, communication delay, actuator saturation, UAV dynamic constraints, and complex meteorological disturbances were not fully considered. In addition, the experiments focused on small- to medium-scale UAV formations in simulated urban scenarios, and the applicability to larger UAV swarms and denser real urban environments requires further validation. Moreover, the framework adopts two-dimensional grid search with altitude assignment based on the building-height field, which improves efficiency but cannot fully represent all possible three-dimensional maneuvers. Future work will incorporate higher-fidelity UAV dynamics, real 3D urban maps, hardware-in-the-loop simulation, and real-flight experiments.

6. Conclusions

With the rapid development of urban low-altitude operations, autonomous trajectory planning and cooperative obstacle avoidance for multiple UAVs in complex building environments and dynamic-obstacle-dense regions have become important issues for ensuring safe low-altitude mission execution. This study proposes a three-dimensional cooperative trajectory-planning method for multiple UAVs by integrating dynamic-aware ACO and cooperative MPC-GWO. The proposed framework first uses dynamic-risk-aware ACO to generate safe initial trajectories, then applies sliding-window GWO for local trajectory smoothing, and finally combines cooperative MPC-GWO, reconfigurable formation control, and goal-neighborhood safety control to achieve online obstacle avoidance, inter-UAV separation, and orderly terminal arrival in simulated low-altitude urban scenarios.
Results show that the proposed method achieves good comprehensive performance in dynamic-risk suppression, inter-UAV safety maintenance, and mission-completion reliability. Compared with conventional ACO and ACO+GWO methods, the proposed method effectively reduces the dynamic risk cost and maintains a relatively stable safety separation during multi-UAV cooperative flights. The ablation experiments further show that the dynamic-obstacle prediction and time-domain risk assessment mechanism helps reduce dynamic-obstacle avoidance conflicts; the reconfigurable formation mechanism improves the passage capability of multiple UAVs in constrained spaces such as narrow channels; and the goal-neighborhood safety-control strategy alleviates congestion and local conflicts caused by the simultaneous arrival of multiple UAVs near the goal region. These results demonstrate that the proposed method can improve the safety, cooperation, and robustness of multi-UAV systems in highly dynamic environments while maintaining trajectory feasibility.
Overall, the proposed method integrates global dynamic-risk perception, local trajectory smoothing, online cooperative replanning, and formation reconfiguration into a unified framework, providing a feasible solution for three-dimensional multi-UAV trajectory planning in the tested simulated low-altitude urban scenarios. Compared with three-dimensional trajectory-planning methods that focus only on static obstacles or single-UAV path optimization, this study places greater emphasis on multi-UAV cooperative safety and mission-execution reliability under dynamic-obstacle disturbances. Therefore, the proposed method is more suitable for application scenarios such as urban low-altitude inspection, logistics delivery, disaster monitoring, and swarm-cooperative missions. Future research will further incorporate high-fidelity UAV dynamic models, real three-dimensional urban maps, hardware-in-the-loop simulation, and real-flight experiments to validate the real-time performance, robustness, and engineering applicability of the proposed method in large-scale low-altitude urban missions.

Author Contributions

Writing—original draft, Y.W.; Investigation, J.L. and Z.L.; Conceptualization, Y.W. and P.Z.; Supervision, P.Z.; Methodology, Y.W. and P.Z.; Formal analysis, Y.L. and D.R.; Validation, Y.L., D.R. and H.H. All authors have read and agreed to the published version of the manuscript.

Funding

This research was funded in part by the Supported by Foundation of Shanxi Key Laboratory of Machine Vision and Virtual Reality (No. 447-110103); the Shanxi Provincial Basic Research Program (No. 202403021221121); the Shanxi Provincial Postgraduate Practical Innovation Project (No. 2025SJ026); and the Shanxi Science and Technology Innovation Leading Talent Team for Special Unmanned Systems and Intelligent Equipment (No. 202204051002001).

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The data and code supporting the findings of this study are not publicly available but are available from the corresponding author upon reasonable request.

Acknowledgments

The authors sincerely thank the editors and reviewers for their valuable comments and suggestions on this manuscript and gratefully acknowledge the support of the School of Mechanical and Electrical Engineering and the Research Institute of Intelligent Weapons, North University of China, during this research.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
3DThree-Dimensional
ACOAnt Colony Optimization
GWOGray Wolf Optimizer
ACO+GWOAnt Colony Optimization plus Gray Wolf Optimizer
MPCModel Predictive Control
MPC-GWOModel Predictive Control–Gray Wolf Optimizer
UAVUnmanned Aerial Vehicle
Avg.Average
Env.Environmental

References

  1. Debnath, D.; Vanegas, F.; Sandino, J.; Hawary, A.F.; Gonzalez, F. A review of UAV path-planning algorithms and obstacle avoidance methods for remote sensing applications. Remote Sens. 2024, 16, 4019. [Google Scholar] [CrossRef]
  2. Zhou, Y.; Yan, L.; Han, Y.; Xie, H.; Zhao, Y. A Survey on the Key Technologies of UAV Motion Planning. Drones 2025, 9, 194. [Google Scholar] [CrossRef]
  3. Rahman, M.; Sarkar, N.I.; Lutui, R. A survey on multi-UAV path planning: Classification, algorithms, open research problems, and future directions. Drones 2025, 9, 263. [Google Scholar] [CrossRef]
  4. Li, Y.; Zhang, Z.; Sun, Q.; Huang, Y. An improved ant colony algorithm for multiple unmanned aerial vehicles route planning. J. Frankl. Inst. 2024, 361, 107060. [Google Scholar] [CrossRef]
  5. Pehlivanoglu, V.Y.; Pehlivanolu, P. An efficient path planning approach for autonomous multi-UAV system in target coverage problems. Aircr. Eng. Aerosp. Technol. 2024, 96, 17. [Google Scholar] [CrossRef]
  6. Yi, W. Path Planning for Multi-UAV in a Complex Environment Based on Reinforcement-Learning-Driven Continuous Ant Colony Optimization. Drones 2025, 9, 638. [Google Scholar] [CrossRef]
  7. Lu, Y.; Yan, D.; Wan, Z.; Feng, C. Conflict-Free 3D Path Planning for Multi-UAV Based on Jump Point Search and Incremental Update. Drones 2025, 9, 688. [Google Scholar] [CrossRef]
  8. Cabreira, T.M.; Brisolara, L.B.; Paulo, R.F.J. Survey on coverage path planning with unmanned aerial vehicles. Drones 2019, 3, 4. [Google Scholar] [CrossRef]
  9. Dewangan, R.K.; Shukla, A.; Godfrey, W.W. Three dimensional path planning using Grey wolf optimizer for UAVs. Appl. Intell. 2019, 49, 2201–2217. [Google Scholar] [CrossRef]
  10. Bayerlein, H.; Theile, M.; Caccamo, M.; Gesbert, D. Multi-UAV path planning for wireless data harvesting with deep reinforcement learning. IEEE Open J. Commun. Soc. 2021, 2, 1171–1187. [Google Scholar] [CrossRef]
  11. Chang, X.; Wang, J.; Li, K.; Zhang, X.; Tang, Q. Research on multi-UAV autonomous obstacle avoidance algorithm integrating improved dynamic window approach and ORCA. Sci. Rep. 2025, 15, 14646. [Google Scholar] [CrossRef] [PubMed]
  12. Yan, C.; Xiang, X.; Wang, C. Towards real-time path planning through deep reinforcement learning for a UAV in dynamic environments. J. Intell. Robot. Syst. 2020, 98, 297–309. [Google Scholar]
  13. Wang, D.; Fan, T.; Han, T.; Pan, J. A two-stage reinforcement learning approach for multi-UAV collision avoidance under imperfect sensing. IEEE Robot. Autom. Lett. 2020, 5, 3098–3105. [Google Scholar] [CrossRef]
  14. AlMahamid, F.; Grolinger, K. Autonomous unmanned aerial vehicle navigation using reinforcement learning: A systematic review. Eng. Appl. Artif. Intell. 2022, 115, 105321. [Google Scholar] [CrossRef]
  15. Xue, Y.; Chen, W. Multi-agent deep reinforcement learning for UAVs navigation in unknown complex environment. IEEE Trans. Intell. Veh. 2023, 9, 2290–2303. [Google Scholar] [CrossRef]
  16. Yan, C.; Xiang, X.; Wang, C. Fixed-Wing UAVs flocking in continuous spaces: A deep reinforcement learning approach. Robot. Auton. Syst. 2020, 131, 103594. [Google Scholar] [CrossRef]
  17. Lin, Y.; Saripalli, S. Sampling-Based Path Planning for UAV Collision Avoidance. IEEE Trans. Intell. Transp. Syst. 2017, 18, 3179–3192. [Google Scholar] [CrossRef]
  18. Ma, B.; Ji, Y.; Fang, L. A multi-uav formation obstacle avoidance method combined with improved simulated annealing and an adaptive artificial potential field. Drones 2025, 9, 390. [Google Scholar] [CrossRef]
  19. Liu, Y.; Liu, Z.; Wang, G.; Yan, C.; Wang, X.; Huang, Z. Flexible multi-UAV formation control via integrating deep reinforcement learning and affine transformations. Aerosp. Sci. Technol. 2025, 157, 109812. [Google Scholar] [CrossRef]
  20. Li, Y.; Zhang, P.; Wang, Z.; Rong, D.; Niu, M.; Liu, C. Multi-UAV obstacle avoidance and formation control in unknown environments. Drones 2024, 8, 714. [Google Scholar] [CrossRef]
  21. Liu, X.; Yan, C.; Zhou, H.; Chang, Y.; Xiang, X.; Tang, D. Towards Flocking Navigation and Obstacle Avoidance for Multi-UAV Systems through Hierarchical Weighting Vicsek Model. Aerospace 2021, 8, 286. [Google Scholar] [CrossRef]
  22. Choi, J.; Choi, Y. Development of a Bearing-Based Distributed Control Method for UAV Formation Tracking and Obstacle Avoidance. Aerospace 2025, 12, 1013. [Google Scholar] [CrossRef]
  23. Mirjalili, S.; Mirjalili, S.M.; Lewis, A. Grey Wolf Optimizer. Adv. Eng. Softw. 2014, 69, 46–61. [Google Scholar] [CrossRef]
  24. Stutzle, M.D.T. Ant Colony Optimization; Bradford Company: Denver, CO, USA, 2004. [Google Scholar]
  25. Mayne, D.Q.; Rawlings, J.B.; Rao, C.V.; Scokaert, P.O.M. Constrained model predictive control: Stability and optimality. Automatica 2000, 36, 789–814. [Google Scholar] [CrossRef]
  26. Ramezani, M.; Habibi, H.; Sanchez-Lopez, J.L.; Voos, H. UAV Path Planning Employing MPC-reinforcement Learning Method Considering Collision Avoidance. In 2023 International Conference on Unmanned Aircraft Systems (ICUAS); IEEE: New York, NY, USA, 2023; pp. 507–514. [Google Scholar]
  27. Xu, Z.; Jin, H.; Han, X.; Shen, H.; Shimada, K. Intent prediction-driven model predictive control for UAV planning and navigation in dynamic environments. IEEE Robot. Autom. Lett. 2025, 10, 4946–4953. [Google Scholar] [CrossRef]
Figure 1. Overall framework for cooperative trajectory planning of multiple UAVs.
Figure 1. Overall framework for cooperative trajectory planning of multiple UAVs.
Technologies 14 00454 g001
Figure 2. Schematic illustration of ACO-based global search. (a) Initial exploration stage. (b) Pheromone accumulation stage. (c) Path convergence stage.
Figure 2. Schematic illustration of ACO-based global search. (a) Initial exploration stage. (b) Pheromone accumulation stage. (c) Path convergence stage.
Technologies 14 00454 g002
Figure 3. Schematic illustration of local trajectory replanning and formation release near the goal region under highly dynamic obstacles. (a) Local replanning under strongly dynamic obstacles. (b) Formation release and orderly approach in the goal neighborhood.
Figure 3. Schematic illustration of local trajectory replanning and formation release near the goal region under highly dynamic obstacles. (a) Local replanning under strongly dynamic obstacles. (b) Formation release and orderly approach in the goal neighborhood.
Technologies 14 00454 g003
Figure 4. Schematic illustration of reconfigurable formation-mode switching. (a) Triangular formation in open areas. (b) Column formation in narrow corridors. (c) Free-formation release in high-risk or terminal regions.
Figure 4. Schematic illustration of reconfigurable formation-mode switching. (a) Triangular formation in open areas. (b) Column formation in narrow corridors. (c) Free-formation release in high-risk or terminal regions.
Technologies 14 00454 g004
Figure 5. Three-dimensional cooperative trajectories generated by the six evaluated methods: (a) A1, conventional ACO; (b) A2, ACO+GWO; (c) A3, MPC-based local replanning; (d) A4, ACO+GWO+MADDPG; (e) A5, ACO+GWO+MAPPO; and (f) A6, the proposed method.
Figure 5. Three-dimensional cooperative trajectories generated by the six evaluated methods: (a) A1, conventional ACO; (b) A2, ACO+GWO; (c) A3, MPC-based local replanning; (d) A4, ACO+GWO+MADDPG; (e) A5, ACO+GWO+MAPPO; and (f) A6, the proposed method.
Technologies 14 00454 g005
Figure 6. Local trajectory adjustment process under typical dynamic disturbance conditions. (a) Local trajectory adjustment by B1. (b) Local trajectory adjustment by B2. (c) Local trajectory adjustment by B3.
Figure 6. Local trajectory adjustment process under typical dynamic disturbance conditions. (a) Local trajectory adjustment by B1. (b) Local trajectory adjustment by B2. (c) Local trajectory adjustment by B3.
Technologies 14 00454 g006
Figure 7. Cooperative obstacle avoidance under different formation strategies in a narrow dynamic corridor. (a) Cooperative obstacle avoidance by C1. (b) Cooperative obstacle avoidance by C2.
Figure 7. Cooperative obstacle avoidance under different formation strategies in a narrow dynamic corridor. (a) Cooperative obstacle avoidance by C1. (b) Cooperative obstacle avoidance by C2.
Technologies 14 00454 g007
Figure 8. Comparison of UAV trajectory distributions under the terminal-neighborhood safety-control strategy. (a) Trajectory distribution of D1. (b) Trajectory distribution of D2. (c) Trajectory distribution of D3.
Figure 8. Comparison of UAV trajectory distributions under the terminal-neighborhood safety-control strategy. (a) Trajectory distribution of D1. (b) Trajectory distribution of D2. (c) Trajectory distribution of D3.
Technologies 14 00454 g008
Figure 9. Sensitivity analysis of key algorithm parameters. (a) Risk weight  ω r . (b) Prediction horizon  H p  (c) Formation-maintenance weight  ω f .
Figure 9. Sensitivity analysis of key algorithm parameters. (a) Risk weight  ω r . (b) Prediction horizon  H p  (c) Formation-maintenance weight  ω f .
Technologies 14 00454 g009
Figure 10. Scalability and real-time computational performance under different mission-scale scenarios. (a) Success rate. (b) Hazardous-time ratio. (c) Global planning time. (d) Online replanning time.
Figure 10. Scalability and real-time computational performance under different mission-scale scenarios. (a) Success rate. (b) Hazardous-time ratio. (c) Global planning time. (d) Online replanning time.
Technologies 14 00454 g010
Table 1. Notations used in the computational complexity analysis.
Table 1. Notations used in the computational complexity analysis.
SymbolDefinition
M Number of UAVs
Q Number of dynamic obstacles
G Number of grid cells
N a Number of ACO ants
I a Maximum ACO iterations
L p Average number of nodes
b Average number of candidate neighboring nodes
N g GWO population size
I g Maximum GWO iterations
K w Number of sliding windows
L w Number of trajectory nodes in each sliding window
s w Sliding-window step size
H p Prediction horizon
N m Population size of MPC-GWO
I m Maximum number of MPC-GWO iterations
T Number of online replanning cycles
Table 2. Algorithm parameter settings and tested ranges.
Table 2. Algorithm parameter settings and tested ranges.
ParameterTested RangeValue
N a 20–5030
I a 50–12080
N g 15–4025
I g 20–8040
L w 5–117
s w 1–53
H p 5–1511
Pheromone weight0.5–2.01.0
Heuristic-function weight3.0–8.06.0
Risk weight0.5–4.02.0
Evaporation coefficient0.03–0.150.08
Terminal-neighborhood radius5–2010
Formation-maintenance weight0.1–0.50.2
Inter-UAV separation weight1.0–4.02.5
Cooperative-arrival weight0.1–0.50.25
The objective-function weights were determined through preliminary tests and one-factor-at-a-time sensitivity analyses. The final values were selected from stable parameter regions by balancing success rate, collision frequency, path length, formation error, and online replanning time.
Table 3. Overall performance comparison of multi-UAV cooperative planning methods.
Table 3. Overall performance comparison of multi-UAV cooperative planning methods.
MethodSuccess Rate/%Reached Rate/%Dynamic
Collisions
Inter-UAV
Collisions
Min.
DynDist
Min.
UAVDist
Composite Score
A116.7 ± 37.9100.0 ± 0.02.33 ± 1.663.10 ± 2.421.79 ± 1.331.23 ± 1.2627.86 ± 7.80
A216.7 ± 37.9100.0 ± 0.03.13 ± 2.270.00 ± 0.002.41 ± 3.073.19 ± 0.0031.00 ± 30.37
A386.7 ± 34.6100.0 ± 0.00.00 ± 0.000.50 ± 1.663.66 ± 1.144.15 ± 1.2186.80 ± 14.90
A470.0 ± 46.696.7 ± 18.30.40 ± 0.690.00 ± 0.004.14 ± 1.794.73 ± 0.3480.23 ± 27.61
A590.0 ± 30.596.7 ± 10.20.03 ± 0.180.03 ± 0.184.36 ± 2.583.57 ± 0.7681.73 ± 15.74
A693.3 ± 25.498.9 ± 6.10.00 ± 0.000.00 ± 0.005.62 ± 1.315.18 ± 0.4996.20 ± 5.30
Table 4. Dynamic-obstacle avoidance performance comparison.
Table 4. Dynamic-obstacle avoidance performance comparison.
MethodSuccess (%)Collisions per RunAvg. RiskAvg. Path (m)Avg. SmoothnessRun Time/s
B166.67 ± 14.990.57 ± 1.040.30 ± 0.64217.35 ± 40.10132.41 ± 37.170.61 ± 0.13
B266.67 ±12.631.03 ± 1.160.48 ± 0.53224.03 ± 42.05132.31 ± 34.580.62 ± 0.13
B394.44 ± 15.540.03 ± 0.180.03 ± 0.11179.62 ± 39.6479.87 ± 34.790.55 ± 0.14
Values are reported as mean ± standard deviation over 30 independent Monte Carlo runs.
Table 5. Formation-coordination performance metrics.
Table 5. Formation-coordination performance metrics.
MethodPassage Success/%Avg Risk CostFormation Error/mMin. Distance/mPassing Time/sRuntime/s
C10.00 ± 0.000.43 ± 0.00120.00 ± 0.009.22 ± 0.0014.00 ± 0.000.08 ± 0.0098
C2100.00 ± 0.000.07 ± 0.00144.50 ± 0.025 ± 0.0011.40 ± 0.000.09 ± 0.0098
Values are reported as mean ± standard deviation over 30 Monte Carlo runs.
Table 6. Statistical results of dynamic-obstacle avoidance and terminal-conflict handling.
Table 6. Statistical results of dynamic-obstacle avoidance and terminal-conflict handling.
MethodTerminal ConflictsWaiting StepsDetour Distance/mCollisionsSuccess/%Arrival Time/s
D113.21 ± 1.260.00 ± 0.000.00 ± 0.0013.27 ± 1.270.00 ± 0.0025.00 ± 0.00
D213.06 ± 1.250.00 ± 0.000.30 ± 0.0713.01 ± 1.300.00 ± 0.0024.17 ± 2.24
D30.00 ± 0.0012.59 ± 13.743.58 ± 5.460.05 ± 0.1896.67 ± 10.1712.13 ± 5.55
Values are reported as mean ± standard deviation over 30 Monte Carlo runs. Arrival time denotes the team’s terminal-arrival time. The detour distance of failed methods should not be interpreted as better efficiency, because D1 and D2 usually failed before completing safe terminal arrival.
Table 7. Parameter sensitivity settings.
Table 7. Parameter sensitivity settings.
ParameterTested ValuesSelected ValueMain Influence
ω r 0; 0.5; 1; 1.5; 2; 3; 42Balances dynamic-risk avoidance and path efficiency
H p 5; 7; 9; 11; 13; 15; 1711Balances predictive avoidance and online runtime
ω f 0; 0.1; 0.2; 0.4; 0.6; 0.8; 1.00.2Balances formation keeping and avoidance flexibility
Table 8. Scalability test scenarios.
Table 8. Scalability test scenarios.
ScenarioMap Size/mBuildingsUAVsDynamic ObstaclesMain Changed Factor
S1100 × 1004535Baseline
S2100 × 1004555UAV number
S3100 × 10045510Dynamic-obstacle density
S4200 × 20080510Map size and building density
Table 9. Detailed scalability performance under different mission-scale scenarios.
Table 9. Detailed scalability performance under different mission-scale scenarios.
ScenarioSuccess/%Hazardous-Time/%Dyn.
Collisions
Inter-UAV
Collisions
Global Time/sOnline Time/s
S193.3 ± 21.081.15 ± 0.720.00 ± 0.000.10 ± 0.3242.6 ± 12.80.035 ± 0.01
S296.0 ± 12.641.30 ± 1.160.03 ± 0.180.10 ± 0.3276.4 ± 20.60.050 ± 0.02
S3100.0 ± 0.001.96 ± 2.580.03 ± 0.180.13 ± 0.35117.8 ± 30.40.067 ± 0.02
S494.0 ± 13.490.084 ± 0.1760.13 ± 0.350.30 ± 0.48168.5 ± 42.70.112 ± 0.04
Since the average online replanning time in all scenarios was lower than the predefined replanning interval of 0.5 s, the proposed method satisfied the real-time requirement in the tested scenarios.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Wang, Y.; Zhang, P.; Li, Y.; Liu, J.; Li, Z.; Rong, D.; Han, H. Low-Altitude Multi-UAV Trajectory Planning in Dynamic Urban Environments Using Dynamic-Aware ACO and MPC-GWO. Technologies 2026, 14, 454. https://doi.org/10.3390/technologies14080454

AMA Style

Wang Y, Zhang P, Li Y, Liu J, Li Z, Rong D, Han H. Low-Altitude Multi-UAV Trajectory Planning in Dynamic Urban Environments Using Dynamic-Aware ACO and MPC-GWO. Technologies. 2026; 14(8):454. https://doi.org/10.3390/technologies14080454

Chicago/Turabian Style

Wang, Yuhan, Pengfei Zhang, Yawen Li, Jinshuai Liu, Zhengxuan Li, Dian Rong, and Huiyan Han. 2026. "Low-Altitude Multi-UAV Trajectory Planning in Dynamic Urban Environments Using Dynamic-Aware ACO and MPC-GWO" Technologies 14, no. 8: 454. https://doi.org/10.3390/technologies14080454

APA Style

Wang, Y., Zhang, P., Li, Y., Liu, J., Li, Z., Rong, D., & Han, H. (2026). Low-Altitude Multi-UAV Trajectory Planning in Dynamic Urban Environments Using Dynamic-Aware ACO and MPC-GWO. Technologies, 14(8), 454. https://doi.org/10.3390/technologies14080454

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop