Active Interception for Multi-Target Encirclement by Heterogeneous UAVs: An LSTM-Enhanced Independent PPO Algorithm
Abstract
1. Introduction
- Unlike most existing studies that focus on single-target encirclement and assume UAV homogeneity, this paper studies the many-to-many confrontation problem of heterogeneous UAVs working together to capture multiple actively escaping targets. The targets are modeled as reinforcement learning agents that can continuously learn and evolve escape strategies. This creates a dynamic environment that closely approximates real-world confrontation scenarios.
- Unlike traditional methods that rely on passive pursuit strategies, this paper proposes an active interception decision framework based on trajectory prediction. An LSTM network is introduced to learn and predict the future trajectory of the target in real time. The hunter uses the point of the predicted trajectory as the interception target, thereby achieving a strategic upgrade from passive following to active anticipation.
2. Literature Review
| Authors | Algorithm Category | Specific Algorithm | Targets | Obstacle Type | Energy | Agent Type | Target Behavior |
|---|---|---|---|---|---|---|---|
| [10] | Control | LMPC | 1 | Static | no | Homogeneous | Static |
| [16] | Control | LMPC + Feedback Linearization | 1 | None | no | Homogeneous | Mobile |
| [11] | Control | LMPC + Feedback Linearization | 1 | None | no | Homogeneous | Static |
| [17] | MARL | Co-DQL | 1 | Static and Dynamic | no | Homogeneous | Active Escape |
| [12] | DCC | DCC | 1 | None | no | Homogeneous | Mobile |
| [18] | Control | Cooperative Motion Path Following | 1 | None | no | Homogeneous | Mobile |
| [2] | Control | Distributed Vector Field Control | 1 | None | no | Homogeneous | Static |
| [19] | Hybrid | Hybrid Guidance + MDP + Collaborative Switching | 1 | Static | no | Homogeneous | Mobile |
| [13] | DCC | DCC + Dynamic Observer | 1 | None | no | Homogeneous | Static/Mobile |
| [20] | MARL | MADDPG | 1 | None | no | Homogeneous | Active Escape |
| [21] | Hybrid | Target Clustering + Guidance Law | 2 | None | no | Homogeneous | Mobile |
| [4] | MARL | CEL-MADDPG | 1 | Static | no | Homogeneous | Active Escape |
| [22] | MARL | GCMSA | 1 | Static | no | Homogeneous | Mobile |
| [23] | Artificial Potential Field | Artificial Potential Field | 3 | Static and Dynamic | no | Homogeneous | Mobile |
| [24] | MARL | GCMSA | 1 | Static | no | Homogeneous | Mobile |
| [6] | Hybrid | Greedy Assignment + Dynamic Sectorization | 3 | None | no | Homogeneous | Static |
| [25] | Control | Event-Triggered Distributed Control | 1 | None | no | Homogeneous | Mobile |
| [5] | Neural Network Control | FWNN + Distributed Anti-Synchronization Controller | 3 | None | no | Homogeneous | Mobile |
| [14] | MARL | MADDPG | 1 | Static | no | Homogeneous | Mobile |
| [3] | MARL | EIR-MARL | 1 | Static | Yes | Heterogeneous | Mobile |
| [26] | Game Theory + RL | SPG + TMSAC | 1 | Static | no | Homogeneous | Mobile |
| [27] | MARL | TP-MADDPG | 1 | Static | no | Homogeneous | Mobile |
| [9] | MARL | MAPPO + LSTM | 1 | Static | no | Homogeneous | Mobile |
| [15] | MARL | MAPPO | 1 | None | no | Homogeneous | Active Escape |
| [28] | Probabilistic Graphical Model | PD-PGM | 1 | Static and Dynamic | no | Homogeneous | Mobile |
| [29] | RL | SAC + LSTM | 1 | None | Yes | Homogeneous | Mobile |
| This study | RL | IPPO + LSTM | 5 | Static | Yes | Heterogeneous | Active Escape |
3. Theoretical Foundation
3.1. Problem Definition
3.2. UAV Kinematics Model
4. The Proposed Algorithm: LIPPO
4.1. Algorithm Framework
| Algorithm 1: Framework of LIPPO | |||||
| Input: Emax (total episodes), Tmax (max steps), λ, γ, τ( RL parameters ). | |||||
| Output: Trained actor μθ and critic Qϕ networks for all agents. | |||||
| 1: | Initialize actor networks μθ and critic Qϕ networks for hunters. | ||||
| 2: | Initialize target networks μ′θ and critic Q′ϕ. | ||||
| 3: | Initialize replay buffers RH (hunters), RT (targets). | ||||
| 4: | Initialize LSTM predictor μLSTM and its trajectory buffer Rtraj. | ||||
| 5: | For episode = 1 to Emax do | ||||
| 6: | Reset environment and get initial observations s. | ||||
| 7: | Assign targets to hunters using greedy distance-based rule (once per episode) | ||||
| 8: | For t = 1 to Tmax do | ||||
| 9: | For each agent i, select action ai | ||||
| 10: | Execute actions, observe next states s′, rewards r. | ||||
| 11: | Store transition (s, a, r, s′) in corresponding replay buffer. | ||||
| 12: | s ← s′ | ||||
| 13: | If replay buffers (Size > 1024) are ready and t mod 50 == 0 | ||||
| 14: | For each agent i do | ||||
| 15: | Sample a random minibatch B (BatchSize = 256). | ||||
| 16: | Calculate TD targets and Critic loss. | ||||
| 17: | Update critic parameters φ by minimizing L(φ). | ||||
| 18: | Calculate advantages and the importance sampling ratio. | ||||
| 19: | Calculate PPO’s clipped surrogate objective. | ||||
| 20: | End | ||||
| 21: | Soft-update target networks (τ = 0.01). | ||||
| 22: | End | ||||
| 23: | Train μLSTM with Equation (7). | ||||
| 24: | End | ||||
| 25: | End | ||||
4.2. LSTM-Based Trajectory Prediction for Active Interception
4.2.1. Target–Hunter Assignment and LSTM Network
4.2.2. Online Trajectory Prediction Using LSTM
4.2.3. Active Interception
4.3. LIPPO
4.3.1. State
4.3.2. Action
4.3.3. Reward
4.3.4. Network Update
5. Experiment and Analysis
5.1. Parameter Setting
5.2. Ablation Experiments
5.3. Comparative Experiments
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Hafez, A.T.; Iskandarani, M.; Givigi, S.N.; Yousefi, S.; Rabbath, C.A.; Beaulieu, A. Using Linear Model Predictive Control via Feedback Linearization for dynamic encirclement. In Proceedings of the 2014 American Control Conference, Portland, OR, USA, 4–6 June 2014, 2014; pp. 3868–3873. [Google Scholar]
- Gao, Y.; Bai, C.; Zhang, L.; Quan, Q. Multi-UAV cooperative target encirclement within an annular virtual tube. Aerosp. Sci. Technol. 2022, 128, 107800. [Google Scholar] [CrossRef] [Scilit]
- Chen, J.C.; Wang, Y.; Zhang, Y.; Lu, Y.T.; Shu, Q.H.; Hu, Y.J. Extrinsic-and-Intrinsic Reward-Based Multi-Agent Reinforcement Learning for Multi-UAV Cooperative Target Encirclement. IEEE Trans. Intell. Transp. Syst. 2025, 26, 17653–17665. [Google Scholar] [CrossRef] [Scilit]
- Li, B.; Wang, J.; Song, C.; Yang, Z.; Wan, K.; Zhang, Q. Multi-UAV roundup strategy method based on deep reinforcement learning CEL-MADDPG algorithm. Expert Syst. Appl. 2024, 245, 123018. [Google Scholar] [CrossRef] [Scilit]
- Liu, F.; Yuan, S.H.; Meng, W.; Su, R.; Xie, L.H. Multiple Noncooperative Targets Encirclement by Relative Distance-Based Positioning and Neural Antisynchronization Control. IEEE Trans. Ind. Electron. 2024, 71, 1675–1685. [Google Scholar] [CrossRef] [Scilit]
- Kumar, G.; Ratnoo, A. Cooperative Multiple Target Encirclement via Platooning. In AIAA SCITECH 2024 Forum; AIAA SciTech Forum; American Institute of Aeronautics and Astronautics: Reston, VA, USA, 2024. [Google Scholar]
- Zhang, Y.; Jia, Z.; Dong, C.; Liu, Y.; Zhang, L.; Wu, Q. Recurrent LSTM-based UAV Trajectory Prediction with ADS-B Information. In Proceedings of the IEEE Global Communications Conference (GLOBECOM), Rio de Janeiro, Brazil, 4–8 December 2022; pp. 6475–6480. [Google Scholar]
- Xie, L.; Liu, M.; Xiao, L.; Guo, S. Integrating LSTM-Based Target Prediction with TD3 for Enhanced UAV Pursuit-Evasion. In Proceedings of the 2024 China Automation Congress (CAC), Qingdao, China, 1–3 November 2024; pp. 3541–3546. [Google Scholar]
- Chen, J.; Yu, C.; Li, G.; Tang, W.; Ji, S.; Yang, X.; Xu, B.; Yang, H.; Wang, Y. Online Planning for Multi-UAV Pursuit-Evasion in Unknown Environments Using Deep Reinforcement Learning. IEEE Robot. Autom. Lett. 2025, 10, 8196–8203. [Google Scholar] [CrossRef] [Scilit]
- Iskandarani, M.; Givigi, S.N.; Rabbath, C.A.; Beaulieu, A. Linear Model Predictive Control for the Encirclement of a Target Using a Quadrotor Aircraft. In Proceedings of the 21st Mediterranean Conference on Control and Automation (MED), Platanias, Greece, 25–28 June 2013; pp. 1550–1556. [Google Scholar]
- Hafez, A.T.; Marasco, A.J.; Givigi, S.N.; Iskandarani, M.; Yousefi, S.; Rabbath, C.A. Solving Multi-UAV Dynamic Encirclement via Model Predictive Control. IEEE Trans. Control Syst. Technol. 2015, 23, 2251–2265. [Google Scholar] [CrossRef] [Scilit]
- Wei, X.; Yang, J.; Fan, X. Distributed guidance law design for multi-UAV multi-direction attack based on reducing surrounding area. Aerosp. Sci. Technol. 2020, 99, 105571. [Google Scholar] [CrossRef] [Scilit]
- Jia, J.; Chen, X.; Wang, W.; Zhang, M. Distributed control of target cooperative encirclement and tracking using range-based measurements. Asian J. Control 2023, 25, 4595–4608. [Google Scholar] [CrossRef] [Scilit]
- Niu, Y.; Tian, Y.; Wang, Q. Counter-Encirclement of UAV in Pursuit-Evasion Environment via Improved RL. In Proceedings of the 2024 IEEE International Conference on Unmanned Systems (ICUS), Nanjing, China, 18–20 October 2024; pp. 266–271. [Google Scholar]
- Lin, Y.; Gao, H.; Xia, Y. Distributed Pursuit-Evasion Game Decision-Making Based on Multi-Agent Deep Reinforcement Learning. Electronics 2025, 14, 2141. [Google Scholar] [CrossRef] [Scilit]
- Hafez, A.T.; Iskandaram, M.; Givigi, S.N.; Yousefi, S.; Noureldin, A.; Beaulieu, A. Encirclement of Moving Target Using Linear Model Predictive Control Via Feedback Linearization. In Proceedings of the IEEE International Conference on Systems, Man, and Cybernetics (SMC), San Diego, CA, USA, 5–8 October 2014; pp. 3078–3083. [Google Scholar]
- Wang, X.; Xuan, S.; Ke, L. Cooperatively pursuing a target unmanned aerial vehicle by multiple unmanned aerial vehicles based on multiagent reinforcement learning. Adv. Control Appl. 2020, 2, e27. [Google Scholar] [CrossRef] [Scilit]
- Wang, W.Z.; Chen, X.; Jia, J.B.; Fu, Z.F. Target localization and encirclement control for multi-UAVs with limited information. IET Control Theory Appl. 2022, 16, 1396–1404. [Google Scholar] [CrossRef] [Scilit]
- Jiang, L.; Wei, R.; Wang, D. Multi-UAV Roundup Inspired by Hierarchical Cognition Consistency Learning Based on an Interaction Mechanism. Drones 2023, 7, 462. [Google Scholar] [CrossRef] [Scilit]
- Xia, Q.; Li, P.; Shi, X.; Li, Q.; Cai, W. Research on Target Capturing of UAV Circumnavigation Formation Based on Deep Reinforcement Learning. In International Conference on Autonomous Unmanned Systems; Springer Nature: Singapore, 2023; pp. 3751–3762. [Google Scholar]
- Jia, J.; Chen, X.; Wang, W.; Liao, H.; Zhu, G. Cooperative Control of Multi-UAV for Multi-Targets Encirclement and Tracking Based on Potential Game. In Proceedings of the 2023 42nd Chinese Control Conference (CCC), Tianjin, China, 24–26 July 2023; pp. 3778–3785. [Google Scholar] [CrossRef] [Scilit]
- Wei, Z.; Wei, R. UAV Swarm Rounding Strategy Based on Deep Reinforcement Learning Goal Consistency with Multi-Head Soft Attention Algorithm. Drones 2024, 8, 731. [Google Scholar] [CrossRef] [Scilit]
- Li, J.; Wei, R.; Zhang, Q.; Shi, R.; Jiang, B. Research on Real-Time Roundup and Dynamic Allocation Methods for Multi-Dynamic Target Unmanned Aerial Vehicles. Sensors 2024, 24, 6565. [Google Scholar] [CrossRef] [Scilit]
- Wei, Z.; Wei, R. UAVs Cluster Target Round up Strategy Based on Neighborhood Cognitive Consistency. In Proceedings of the 2024 International Conference on Guidance, Navigation and Control, Changsha, China, 9–11 August 2024; pp. 56–67. [Google Scholar]
- Jia, J.; Chen, X.; Wang, W.; Zhang, M. Event-triggered cooperative control for moving target encirclement and tracking with time-varying pattern by UAV formation. IET Control Theory Appl. 2024, 18, 55–70. [Google Scholar] [CrossRef] [Scilit]
- Yang, K.; Zhu, M.; Guo, X.; Zhang, Y.; Zhou, Y. Stochastic Potential Game-Based Target Tracking and Encirclement Approach for Multiple Unmanned Aerial Vehicles System. Drones 2025, 9, 103. [Google Scholar] [CrossRef] [Scilit]
- Chen, Y.; Shi, Y.; Dai, X.H.; Meng, Q.; Yu, T. Pursuit-evasion game with online planning using deep reinforcement learning. Appl. Intell. 2025, 55, 512. [Google Scholar] [CrossRef] [Scilit]
- Huang, Y.X.; Xiang, X.J.; Yan, C.; Zhou, H.; Tang, D.Q. Hierarchical probabilistic graphical models for multi-UAV cooperative pursuit in dynamic environments. Robot. Auton. Syst. 2025, 185, 104890. [Google Scholar] [CrossRef] [Scilit]
- Zhang, X.; Guo, H.; Yan, T.; Wang, X.; Sun, W.; Fu, W.; Yan, J. Penetration Strategy for High-Speed Unmanned Aerial Vehicles: A Memory-Based Deep Reinforcement Learning Approach. Drones 2024, 8, 275. [Google Scholar] [CrossRef] [Scilit]
- Wang, H.; Zhang, Y.; Bi, S. Game Strategy Prediction for Spacecraft Orbital Pursuit-Evasion Game Based on Long Short-Term Memory. Space-Sci. Technol. 2025, 5, 0279. [Google Scholar] [CrossRef] [Scilit]
- Wang, C.; Chen, D.; Liao, W. Interactive Multiple-Model Learning Filter for Spacecraft Pursuit-Evasion Game Strategy Switch Based on Long Short-Term Memory Network. Aerospace 2024, 11, 894. [Google Scholar] [CrossRef] [Scilit]
- Zhang, J.; Guo, Y.; Zheng, L.; Yang, Q.; Shi, G.; Wu, Y. Real-Time UAV Path Planning Based on LSTM Network. J. Syst. Eng. Electron. 2024, 35, 374–385. [Google Scholar] [CrossRef] [Scilit]








| Variable Name | Parameters |
|---|---|
| Boundary Length | 2.0 |
| Simulation Time Step | 0.5 |
| Number of Obstacles | 3, 6, 9 |
| Obstacle Radius | 0.05~0.1 |
| Number of Hunters | 15 |
| Initial Velocity of Hunters | 0.0 |
| Maximum Speed of Hunters | 0.06~0.08 |
| Maximum Acceleration of Hunters | 0.02 |
| Sensor Detection Range | 0.2 |
| Number of Lidar Rays | 16 |
| Energy Consumption Rate | 0.1~0.3 |
| Number of Targets | 5 |
| Initial Velocity | 0.0 |
| Maximum Speed of Targets | 0.09 |
| Maximum Acceleration of Targets | 0.03 |
| Target Escape Distance | 0.2 |
| Safety Distance Threshold | 0.02 |
| Algorithms | Hyperparameter | Value |
|---|---|---|
| PPO | Learning Rate | 3 × 10−4 |
| Gamma (γ) | 0.95 | |
| Clip Epsilon (ε) | 0.2 | |
| Update Frequency | 50 | |
| Buffer Warm-up Threshold | 1024 | |
| Batch Size | 256 | |
| Num_episodes | 2000 | |
| Max_steps | 150 | |
| Replay Buffer Size | 1 × 105 | |
| Target network soft update rate | 0.01 | |
| Actor and Critic network structure | 256 × 256 | |
| Active Strategy Trigger | 100 Episodes | |
| LSTM | Input/Hidden Size | 4/64 |
| Output Size | 2 | |
| Sequence Length | 10 | |
| Prediction Horizon | 3 | |
| Learning Rate/Batch Size | 1 × 10−4/64 | |
| Buffer Warm-up Threshold | 5000 | |
| Update Frequency | Every 1 Episode |
| Number of Obstacles | Algorithm | Success Rate (%) | Total Energy Consumption | Average Capture Steps |
|---|---|---|---|---|
| 3 | LIPPO | 82.2 ± 1.92 | 300.63 ± 32.07 | 318.32 ± 19.13 |
| NLIPPO | 76.0 ± 3.16 | 308.38 ± 42.18 | 314.44 ± 7.75 | |
| XIPPO | 70.6 ± 4.61 | 284.52 ± 25.78 | 342.30 ± 11.17 | |
| 6 | LIPPO | 73.60 ± 5.86 | 242.59 ± 56.02 | 306.56 ± 13.02 |
| NLIPPO | 70.8 ± 4.08 | 329.54 ± 44.40 | 334.41 ± 20.47 | |
| XIPPO | 64.4 ± 3.85 | 304.12 ± 61.50 | 339.16 ± 20.79 | |
| 9 | LIPPO | 65.2 ± 6.06 | 276.06 ± 35.75 | 317.67 ± 19.28 |
| NLIPPO | 56.8 ± 4.92 | 331.21 ± 99.80 | 286.43 ± 35.07 | |
| XIPPO | 55.0 ± 6.21 | 320.70 ± 30.11 | 347.81 ± 16.74 |
| Number of Obstacles | Algorithm | Success Rate (%) | Total Energy Consumption | Average Capture Steps |
|---|---|---|---|---|
| 3 | LIPPO | 82.2 ± 1.92 | 300.63 ± 32.07 | 318.32 ± 19.13 |
| LMAPPO | 81.0 ± 3.16 | 309.27 ± 60.93 | 295.54 ± 13.39 | |
| MAPPO | 80.6 ± 3.58 | 309.27 ± 60.93 | 295.54 ± 13.39 | |
| ITD3 | 76.8 ± 3.77 | 335.68 ± 116.73 | 244.03 ± 29.49 | |
| IDDPG | 71.4 ± 5.81 | 288.81 ± 61.72 | 252.25 ± 28.19 | |
| 6 | LIPPO | 73.6 ± 5.86 | 242.59 ± 56.02 | 306.56 ± 13.02 |
| LMAPPO | 75.0 ± 2.92 | 337.98 ± 71.51 | 303.14 ± 9.22 | |
| MAPPO | 73.2 ± 2.68 | 305.65 ± 32.43 | 301.24 ± 10.07 | |
| ITD3 | 72.0 ± 6.28 | 350.76 ± 103.20 | 244.70 ± 24.73 | |
| IDDPG | 73.0 ± 3.94 | 283.98 ± 64.90 | 258.21 ± 17.71 | |
| 9 | LIPPO | 65.2 ± 6.06 | 276.06 ± 35.75 | 317.67 ± 19.28 |
| LMAPPO | 63.80 ± 4.15 | 309.30 ± 79.94 | 309.05 ± 12.53 | |
| MAPPO | 61.6 ± 3.58 | 281.58 ± 23.09 | 328.88 ± 7.45 | |
| ITD3 | 56.6 ± 6.39 | 422.81 ± 217.20 | 276.59 ± 34.80 | |
| IDDPG | 57.4 ± 4.80 | 366.03 ± 153.20 | 270.75 ± 42.98 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Song, Y.; Chen, H. Active Interception for Multi-Target Encirclement by Heterogeneous UAVs: An LSTM-Enhanced Independent PPO Algorithm. Designs 2026, 10, 26. https://doi.org/10.3390/designs10020026
Song Y, Chen H. Active Interception for Multi-Target Encirclement by Heterogeneous UAVs: An LSTM-Enhanced Independent PPO Algorithm. Designs. 2026; 10(2):26. https://doi.org/10.3390/designs10020026
Chicago/Turabian StyleSong, Yuxin, and Hanning Chen. 2026. "Active Interception for Multi-Target Encirclement by Heterogeneous UAVs: An LSTM-Enhanced Independent PPO Algorithm" Designs 10, no. 2: 26. https://doi.org/10.3390/designs10020026
APA StyleSong, Y., & Chen, H. (2026). Active Interception for Multi-Target Encirclement by Heterogeneous UAVs: An LSTM-Enhanced Independent PPO Algorithm. Designs, 10(2), 26. https://doi.org/10.3390/designs10020026

