SA-DSM-MADDPG for Multi-UAV Cooperative Encirclement in Obstacle-Rich Pursuit–Evasion Scenarios
Highlights
- We propose SA-DSM-MADDPG for 3v1 multi-UAV cooperative encirclement, integrating a self-attention critic, double-screened experience replay (PER + relevance screening), and curriculum learning.
- Experimental results show that SA-DSM-MADDPG achieves a higher cooperative capture success rate and more stable convergence than the MADDPG baseline in obstacle-rich environments.
- The improved success rate indicates that combining interaction-aware coordination modeling (attention), phase-relevant sample selection (DSM), and staged reward shaping (curriculum) effectively mitigates sparse-feedback learning issues in pursuit–evasion tasks.
- The proposed design provides a practical guideline for developing more reliable multi-drone interception/containment decision policies in cluttered environments under the CTDE paradigm.
Abstract
1. Introduction
Contributions
- We propose SA-DSM-MADDPG, an integrated CTDE framework for multi-UAV cooperative encirclement that addresses three coupled challenges in MADDPG-based training: dynamic inter-agent dependency modeling, uneven replay quality, and sparse stage-dependent supervision.
- We develop a self-attention-enhanced centralized critic and a double-screened experience replay mechanism to improve coordination learning and replay efficiency. The self-attention module enables each agent to emphasize more relevant teammate information under time-varying interactions, while the DSM mechanism combines prioritized sampling and relevance screening to suppress weakly informative transitions and improve training stability.
- We design a curriculum learning strategy with staged reward shaping for pursuit–evasion encirclement, which provides denser guidance in early training and progressively strengthens task objectives, thereby improving convergence behavior and cooperative performance in obstacle-rich environments.
- Extensive experiments in 3v1 cooperative encirclement scenarios demonstrate that the proposed method achieves faster convergence, improved training stability, and higher success rates than strong baselines, including approximate success-rate gains of 22 percentage points over MADDPG and 35 percentage points over MAPPO.
2. Problem Formulation
2.1. Task Description
2.2. UAV Dynamics and Control Constraints
2.3. Geometry, Safety Constraints, and Terminal Conditions
2.3.1. UAV–UAV Safety Constraint
2.3.2. UAV–Obstacle Safety Constraint
2.3.3. Boundary Constraint
2.3.4. Capture Criterion
2.3.5. Encirclement Criterion
2.4. Observation and Action Spaces
Observation
2.5. Environment Parameters
3. Method: SA-DSM-MADDPG
3.1. Overview
3.2. MADDPG Backbone Under CTDE
3.3. Self-Attention Enhanced Critic (SA)
3.4. Double-Screened Experience Replay (DSM)
3.4.1. First Screening: Prioritized Experience Replay
3.4.2. Second Screening: Relevant Experience Learning
3.5. Curriculum Learning with Staged Reward Shaping
- Search stage: The evader is still outside the effective encirclement region, and the pursuers mainly need to approach the target while reducing their overall distance to the evader.
- Transition stage: The pursuers begin to form an enclosing structure around the evader, but the encirclement is not yet stable or sufficiently compact.
- Capture stage: The evader is enclosed by the pursuer formation and the pursuers further tighten the enclosure until the final capture condition is satisfied.
3.5.1. Stage-Dependent Task Reward
3.5.2. Safety Reward
3.5.3. Individual Pursuit Reward
3.5.4. Terminal Success Reward
3.6. Training Procedure
| Algorithm 1: Training procedure of SA-DSM-MADDPG |
![]() |
4. Experiments and Results
4.1. Experimental Setup
Obstacle Arrangement
4.2. Baselines
4.2.1. Representative Baselines
- MAPPO: An on-policy multi-agent proximal policy optimization method used here as a representative MARL baseline for comparison.
- MADDPG: The standard CTDE baseline with centralized critics and decentralized actors, serving as the main off-policy comparison method [6].
4.2.2. Module-Level Variants
- DSM-MADDPG: MADDPG with the proposed DSM replay mechanism, without attention or curriculum learning [15].
4.3. Evaluation Metrics
- Average episode return: Mean cumulative reward per episode, smoothed by a moving average window.
- Success rate: Percentage of episodes that terminate with success (encirclement and capture).
- Time-to-capture: Average number of steps required to achieve capture, computed over successful episodes.
- Safety violation rate: Percentage of episodes terminated by collisions or boundary violations.
4.4. Training Protocol
4.5. Main Results
4.6. Failure-Type Analysis
4.7. Effect of Obstacle Number
4.8. Component Discussion
4.9. Qualitative Analysis
4.10. Discussion of Results
5. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Conflicts of Interest
References
- Zhang, T.; Liu, Z.; Pu, Z.; Yi, J. Multi-Target Encirclement with Collision Avoidance via Deep Reinforcement Learning. In Proceedings of the IEEE International Conference on Robotics and Automation, Philadelphia, PA, USA, 23–27 May 2022. [Google Scholar]
- Zheng, Y.; Zhang, Y.; Zhao, C.; Yang, H.; Li, T.; Ouyang, Q.; Chen, Y. Faster Target Encirclement with Utilization of Obstacles via Multi-Agent Reinforcement Learning. In Proceedings of the Machine Learning Research, Shenzhen, China, 2–5 February 2024. [Google Scholar]
- Qu, X.; Li, C.; Jiang, S.; Liu, G.; Zhang, R. Multi-Agent Reinforcement Learning-Based Cooperative Encirclement Control Method for Autonomous Surface Vehicles. J. Mar. Sci. Eng. 2025, 13, 1558. [Google Scholar] [CrossRef] [Scilit]
- Kouzeghar, M.; Song, Y.; Meghjani, M.; Bouffanais, R. Multi-Target Pursuit by a Decentralized Heterogeneous UAV Swarm Using Deep Multi-Agent Reinforcement Learning. arXiv 2023, arXiv:2303.01799. [Google Scholar]
- Peng, Z.; Wu, G.; Luo, B.; Wang, L. Multi-UAV Cooperative Pursuit Strategy with Limited Visual Field in Urban Airspace: A Multi-Agent Reinforcement Learning Approach. IEEE/CAA J. Autom. Sin. 2025, 12, 1350–1367. [Google Scholar] [CrossRef] [Scilit]
- Lowe, R.; Wu, Y.; Tamar, A.; Harb, J.; Abbeel, P.; Mordatch, I. Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments. In Proceedings of the NIPS’17: Proceedings of the 31st International Conference on Neural Information Processing Systems, Long Beach, CA, USA, 4–9 December 2017. [Google Scholar]
- Rashid, T.; Samvelyan, M.; Schroeder, C.; Farquhar, G.; Foerster, J.; Whiteson, S. QMIX: Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement Learning. In Proceedings of the 35th International Conference on Machine Learning, Stockholm, Sweden, 10–15 July 2018. [Google Scholar]
- Zhang, Y.; Ding, M.; Zhang, J.; Yang, Q.; Shi, G.; Lu, M.; Jiang, F. Multi-UAV Pursuit–Evasion Gaming Based on PSO-M3DDPG. Complex Intell. Syst. 2024, 10, 6867–6883. [Google Scholar] [CrossRef] [Scilit]
- Wang, X.; Wang, M.; Bai, X.; Ma, Z.; Sun, K.; Li, J. EO-MADDPG: An Improved Reinforcement Learning Approach for Multi-UAV Pursuit–Evasion Games. Aerospace 2026, 13, 296. [Google Scholar] [CrossRef] [Scilit]
- Iqbal, S.; Sha, F. Actor-Attention-Critic for Multi-Agent Reinforcement Learning. arXiv 2018, arXiv:1810.02912. [Google Scholar]
- Liu, K.; Zhao, Y.; Wang, G.; Peng, B. SA-MATD3: Self-Attention-Based Multi-Agent Continuous Control Method in Cooperative Environments. arXiv 2021, arXiv:2107.00284. [Google Scholar]
- Bengio, Y.; Louradour, J.; Collobert, R.; Weston, J. Curriculum Learning. In Proceedings of the 26th Annual International Conference on Machine Learning, Montreal, QC, Canada, 14–18 June 2009; pp. 41–48. [Google Scholar]
- Chen, J.; Li, G.; Yu, C.; Yang, X.; Xu, B.; Yang, H.; Wang, Y. A Dual Curriculum Learning Framework for Multi-UAV Pursuit–Evasion in Diverse Environments. arXiv 2023, arXiv:2312.12255. [Google Scholar]
- Lin, Y.; Gao, H.; Xia, Y. Distributed Pursuit–Evasion Game Decision-Making Based on Multi-Agent Reinforcement Learning with Automatic Curriculum Learning. Electronics 2025, 14, 2141. [Google Scholar] [CrossRef] [Scilit]
- Schaul, T.; Quan, J.; Antonoglou, I.; Silver, D. Prioritized Experience Replay. In Proceedings of the International Conference on Learning Representations, San Juan, Puerto Rico, 2–4 May 2016. [Google Scholar]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser, L.; Polosukhin, I. Attention Is All You Need. In Proceedings of the NIPS’17: Proceedings of the 31st International Conference on Neural Information Processing Systems, Long Beach, CA, USA, 4–9 December 2017. [Google Scholar]






| Parameter | Value |
|---|---|
| Environment dimension D | |
| Number of pursuers | 3 |
| Number of evaders | 1 |
| Pursuer radius | |
| Evader radius | |
| Capture radius | |
| UAV safety clearance | |
| Pursuer max speed | |
| Evader max speed | |
| Pursuer max acceleration | |
| Evader max acceleration | |
| Obstacle radius | |
| Max episode length | 200 steps |
| Evader observability | Directly observable (perfect detection) |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Liang, Q.; Yang, Y.; Liang, S.; Li, H. SA-DSM-MADDPG for Multi-UAV Cooperative Encirclement in Obstacle-Rich Pursuit–Evasion Scenarios. Drones 2026, 10, 360. https://doi.org/10.3390/drones10050360
Liang Q, Yang Y, Liang S, Li H. SA-DSM-MADDPG for Multi-UAV Cooperative Encirclement in Obstacle-Rich Pursuit–Evasion Scenarios. Drones. 2026; 10(5):360. https://doi.org/10.3390/drones10050360
Chicago/Turabian StyleLiang, Qing, Yujie Yang, Shihao Liang, and Hui Li. 2026. "SA-DSM-MADDPG for Multi-UAV Cooperative Encirclement in Obstacle-Rich Pursuit–Evasion Scenarios" Drones 10, no. 5: 360. https://doi.org/10.3390/drones10050360
APA StyleLiang, Q., Yang, Y., Liang, S., & Li, H. (2026). SA-DSM-MADDPG for Multi-UAV Cooperative Encirclement in Obstacle-Rich Pursuit–Evasion Scenarios. Drones, 10(5), 360. https://doi.org/10.3390/drones10050360

