HH-MAPPO: A Hierarchical Reinforcement Learning Framework for Dynamic-Scale Target–Attacker–Defender Games
Abstract
1. Introduction
- 1.
- We propose an asymmetric hierarchical policy architecture for dynamic-scale TAD games. In this design, the GMT learns to make energy-aware decisions on how many defenders to deploy, while each defender learns which attacker to intercept. This asymmetric assignment of strategic roles to heterogeneous agents enables end-to-end joint optimization of resource scheduling and task assignment, which goes beyond fixed rule-based methods.
- 2.
- We introduce a dual-stream conditioned role-aware embedding (RAE) architecture that mitigates policy homogeneity in shared-parameter networks. The RAE first modulates each agent’s observation via its role embedding, and it then extracts task-oriented and role-specific features through separate processing streams before additive fusion. This structured mechanism produces complementary and mutually predictable behaviors essential for stable multi-agent coordination, rather than the unstructured divergence caused by naive independent layers.
- 3.
- Extensive experiments in a 3D TAD simulation with continuous attacker arrivals and strict energy constraints demonstrate that the proposed framework consistently outperforms existing baselines in both interception rate and energy efficiency. Ablation studies further reveal a key insight: effective teamwork depends on structured differentiation grounded in shared representations, rather than the mere magnitude of action dissimilarity.
2. Problem Formulation and Environment Modeling
2.1. Problem Formulation
2.2. Environment Modeling
3. Hierarchical Heterogeneous MAPPO Algorithm Framework
3.1. Overall Architecture
3.2. Policy and Value Network Design
3.2.1. Input Space and Observation Filtering
3.2.2. Role-Aware Embedding
3.2.3. Action Space and Network Output
3.2.4. Centralized Critic Network
3.3. Training with Hierarchical Rewards
3.3.1. Hierarchical Reward Design
Upper-Level Reward
- Matching rationality (). This component rewards high-quality, non-redundant defender–attacker pairings. It is the product of an average pairing-quality score and a non-redundancy factor:
- Quantity appropriateness (). This component guides the GMT’s decision on how many new defenders to deploy:
- Penalty term (). This term directly penalizes three undesirable strategic outcomes:
Individual Credit Assignment
- Defender weight . For defender intercepting attacker , its contribution weight combines three factors:
- GMT weight . The GMT’s contribution weight reflects the quality of its deployment decision:
Lower-Level Reward
Hyperparameter Summary
3.3.2. Hierarchical MAPPO Training Objectives
| Algorithm 1: HH-MAPPO for dynamic TAD defense |
![]() |
4. Experimental Results and Analysis
4.1. Environment and Parameters
4.2. Baseline and Ablation Methods
- Heterogeneous MAPPO (HMAPPO): This scheme employs a single-layer control architecture, where the HMAPPO algorithm is directly used to generate motion policies for both the defenders and the GMT. At its core, it implicitly learns multi-agent coordination and target selection strategies during training through a carefully designed reward signal. This signal calculates the potential interaction reward between each defender and all attackers and selects the maximum value to guide learning.
- Hungarian+HMAPPO: This scheme adopts a two-layer decision-making architecture to achieve explicit task coordination. The upper layer utilizes the Hungarian algorithm to compute the optimal one-to-one interception matching between defenders and attackers, generating deterministic task assignment instructions. Based on these fixed instructions, each defender in the lower layer then employs an independent HMAPPO network to control its own motion, focusing solely on intercepting its specifically assigned target. Consequently, the coordination logic is explicitly dictated by the upper-layer algorithm, rather than being learned from rewards.
- Auction+HMAPPO: Similar to Hungarian+HMAPPO but employs an auction algorithm for the upper-layer assignment, testing a different rule-based coordination strategy.
- HH-MAPPO (w/o RAE): An ablation variant that removes the Role-Aware Embedding (RAE) module. This directly tests RAE’s contribution to policy diversification.
- HH-MAPPO (w/o RAE, w/ Independent Output Layer): An alternative method for policy differentiation, where each defender has a unique final output layer within a shared backbone network. This contrasts with our RAE approach.
- HH-MAPPO (w/ RAE and Independent Output Layer): A combined-method that integrates both the RAE module and independent final output layers for each defender. This design is used to investigate the potential complementary effects and assess whether the combination of RAE and independent layers yields superior policy diversification and performance compared to using either mechanism alone.
4.3. Evaluation Metrics
4.4. Comprehensive Experimental Analysis
4.5. Robustness Against Learning-Based Attackers
4.6. Generalization and Scalability
5. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Zhang, M.; Chen, H.; Cai, W. Hunting Task Allocation for Heterogeneous Multi-AUV Formation Target Hunting in IoUT: A Game Theoretic Approach. IEEE Internet Things J. 2024, 11, 9142–9152. [Google Scholar] [CrossRef] [Scilit]
- Chen, M.; Zhou, Z.; Tomlin, C.J. Multiplayer Reach-Avoid Games via Pairwise Outcomes. IEEE Trans. Autom. Control 2017, 62, 1451–1457. [Google Scholar]
- Wu, H.; Ghadami, A.; Bayrak, A.E.; Smereka, J.M.; Epureanu, B.I. Impact of Heterogeneity and Risk Aversion on Task Allocation in Multi-Agent Teams. IEEE Robot. Autom. Lett. 2021, 6, 7065–7072. [Google Scholar] [CrossRef] [Scilit]
- Zhao, B.; Huo, M.; Li, Z.; Feng, W.; Yu, Z.; Qi, M.; Wang, S. Graph-based multi-agent reinforcement learning for collaborative search and tracking of multiple UAVs. Chin. J. Aeronaut. 2025, 38, 103214. [Google Scholar] [CrossRef] [Scilit]
- Wu, Y.; Chen, Q. Mission planning of the aerial-sea cooperative search and clean for the surface pollutant. Chin. J. Aeronaut. 2026, 39, 103997. [Google Scholar] [CrossRef] [Scilit]
- Huang, H.; Savkin, A.V.; Ni, W. Online UAV Trajectory Planning for Covert Video Surveillance of Mobile Targets. IEEE Trans. Autom. Sci. Eng. 2022, 19, 735–746. [Google Scholar] [CrossRef] [Scilit]
- Deng, Z.; Kong, Z. Multi-Agent Cooperative Pursuit-Defense Strategy Against One Single Attacker. IEEE Robot. Autom. Lett. 2020, 5, 5772–5778. [Google Scholar] [CrossRef] [Scilit]
- Li, Y.; Hu, B.; Li, T.; Guan, Z.-H. Attack-Defense Game of Heterogeneous Multi-Agent Systems Under Actuator Faults. IEEE Trans. Autom. Sci. Eng. 2025, 22, 14699–14711. [Google Scholar] [CrossRef] [Scilit]
- Ruan, W.; Sun, Y.; Deng, Y.; Duan, H. Hawk-Pigeon Game Tactics for Unmanned Aerial Vehicle Swarm Target Defense. IEEE Trans. Ind. Inform. 2023, 19, 11619–11629. [Google Scholar] [CrossRef] [Scilit]
- Wan, K.; Wu, D.; Zhai, Y.; Li, B.; Gao, X.; Hu, Z. An improved approach towards multi-agent pursuit–evasion game decision-making using deep reinforcement learning. Entropy 2021, 23, 1433. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Garcia, E.; Casbeer, D.W.; Pachter, M. The complete differential game of active target defense. J. Optim. Theory Appl. 2021, 191, 675–699. [Google Scholar] [CrossRef] [Scilit]
- Garcia, E.; Casbeer, D.W.; Von Moll, A.; Pachter, M. Multiple Pursuer Multiple Evader Differential Games. IEEE Trans. Autom. Control 2021, 66, 2345–2350. [Google Scholar] [CrossRef] [Scilit]
- Liao, W.; Liang, T.; Xiong, P.; Wang, C.; Song, A.; Liu, P.X. An Improved Level Set Method for Reachability Problems in Differential Games. IEEE Trans. Syst. Man Cybern. Syst. 2024, 54, 2907–2916. [Google Scholar] [CrossRef] [Scilit]
- Li, S.; Wang, C.; Xie, G. Pursuit-evasion differential games of players with different speeds in spaces of different dimensions. In Proceedings of the 2022 American Control Conference (ACC), Atlanta, GA, USA, 8–10 June 2022; pp. 1299–1304. [Google Scholar]
- Makkapati, V.R.; Tsiotras, P.; Zaccour, G. Optimal Evading Strategies and Task Allocation in Multi-player Pursuit–Evasion Problems. Dyn. Games Appl. 2019, 9, 1168–1187. [Google Scholar] [CrossRef] [Scilit]
- Lin, B.; Qiao, L.; Jia, Z.; Sun, Z.; Sun, M.; Zhang, W. Control Strategies for Target-Attacker-Defender Games of USVs. In Proceedings of the 2021 6th International Conference on Automation, Control and Robotics Engineering (CACRE), Dalian, China, 15–17 July 2021; pp. 191–198. [Google Scholar]
- Shi, Y.; Wang, C.; Liang, D. Determination of barrier surface in Target-Attacker-Defender game with capture radius for superior pursuer. Neurocomputing 2025, 638, 130156. [Google Scholar] [CrossRef] [Scilit]
- Li, Y.; Liang, X.; Dang, Z. Nash-equilibrium strategies of orbital Target-Attacker-Defender game with a non-maneuvering target. Chin. J. Aeronaut. 2024, 37, 365–379. [Google Scholar] [CrossRef] [Scilit]
- Gong, X.; Chen, W.; Chen, Z. Active Target Defense Differential Game: An Integrated Guidance and Control Approach. In Proceedings of the 2024 29th International Conference on Automation and Computing (ICAC), Sunderland, UK, 28–30 August 2024; pp. 1–6. [Google Scholar]
- Liang, L.; Deng, F.; Lu, M.; Chen, J. Analysis of Role Switch for Cooperative Target Defense Differential Game. IEEE Trans. Autom. Control 2021, 66, 902–909. [Google Scholar] [CrossRef] [Scilit]
- Gao, P.; Li, X.; Hu, J. Optimal Strategies and Cooperative Teaming for 3-D Multiplayer Reach-Avoid Games. IEEE Trans. Cogn. Dev. Syst. 2024, 16, 2085–2099. [Google Scholar] [CrossRef] [Scilit]
- Yan, R.; Duan, X.; Shi, Z.; Zhong, Y.; Bullo, F. Matching-based capture strategies for 3D heterogeneous multiplayer reach-avoid differential games. Automatica 2022, 140, 110207. [Google Scholar] [CrossRef] [Scilit]
- Garcia, E.; Casbeer, D.W.; Pachter, M. Optimal Strategies for a Class of Multi-Player Reach-Avoid Differential Games in 3D Space. IEEE Robot. Autom. Lett. 2020, 5, 4257–4264. [Google Scholar] [CrossRef] [Scilit]
- Lin, Z.; He, Z.; Wang, X.; Su, W.; Tan, J.; Deng, Y.; Xie, S. Cross-Scale Fuzzy Holistic Attention Network for Diabetic Retinopathy Grading from Fundus Images. IEEE Trans. Emerg. Top. Comput. Intell. 2025, 9, 2164–2178. [Google Scholar] [CrossRef] [Scilit]
- Qin, P.; Fu, Y.; Zhang, J.; Geng, S.; Liu, J.; Zhao, X. DRL-Based Resource Allocation and Trajectory Planning for NOMA-Enabled Multi-UAV Collaborative Caching 6G Network. IEEE Trans. Veh. Technol. 2024, 73, 8750–8764. [Google Scholar] [CrossRef] [Scilit]
- Lin, Z.; He, Z.; Yao, R.; Wang, X.; Liu, T.; Deng, Y.; Xie, S. Deep Dual Attention Network for Precise Diagnosis of COVID-19 from Chest CT Images. IEEE Trans. Artif. Intell. 2024, 5, 104–114. [Google Scholar] [CrossRef] [Scilit]
- Wei, J.; Guo, Y.; Wang, H.; Gu, J.; Zhang, Y.; Yi, J.; Chen, X.; Ding, G. Fluid Antenna Array-Enabled AAV Covert Communications Against Active Warden. IEEE Trans. Wirel. Commun. 2026, 25, 9030–9045. [Google Scholar] [CrossRef] [Scilit]
- Gan, W.; Qiao, L. Many-Versus-Many AUV Attack-Defense Game in 3-D Scenarios Using Hierarchical Multiagent Reinforcement Learning. IEEE Internet Things J. 2025, 12, 23479–23494. [Google Scholar] [CrossRef] [Scilit]
- Qian, H.; Chen, Z.; Wang, X.; Xiao, B.; Meng, L.; Ma, Y. A swarm-independent behaviors-based orbit maneuvering approach for target-attacker-defender games of satellites. Inf. Sci. 2025, 699, 121790. [Google Scholar] [CrossRef] [Scilit]
- Gan, W.; Qiao, L. An Adaptive Deep Reinforcement Learning Framework for AUV Attack-Defense Games. IEEE Trans. Intell. Transp. Syst. 2025, 26, 16320–16335. [Google Scholar] [CrossRef] [Scilit]
- Wang, C.-C.; Wang, Y.-L.; Shi, P.; Wang, F. Scalable-MADDPG-Based Cooperative Target Invasion for a Multi-USV System. IEEE Trans. Neural Netw. Learn. Syst. 2024, 35, 17867–17877. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Zhang, Z.; Liang, X.; Chen, C.; Liu, D.; Yu, C.; Song, Y.; Li, W. Distributed penetration strategy for multi-USV based on scalable deep reinforcement learning. Ocean Eng. 2026, 345, 123793. [Google Scholar] [CrossRef] [Scilit]
- Campbell, R.G.; Eirinaki, M.; Park, Y. Scalable and Autonomous Network Defense Using Reinforcement Learning. IEEE Access 2024, 12, 92919–92930. [Google Scholar] [CrossRef] [Scilit]
- Li, W.; Qiu, Z.; Shao, S.; Song, A. MDDP: Making Decisions from Different Perspectives in Multiagent Reinforcement Learning. IEEE Trans. Games 2024, 16, 621–634. [Google Scholar] [CrossRef] [Scilit]
- Sun, S.; Liu, H.; Xu, K.; Ding, B. Leaders and Collaborators: Addressing Sparse Reward Challenges in Multi-Agent Reinforcement Learning. IEEE Trans. Emerg. Top. Comput. Intell. 2025, 9, 1976–1989. [Google Scholar] [CrossRef] [Scilit]
- Azaki, Z.; Dumon, J.; Offermann, A.; Meslem, N.; Susbielle, P.; Negre, A.; Hably, A. Magnus-Effect Winged Hybrid UAV System: Improved Energy Efficient and Autonomy Through Control Allocation Strategy. IEEE Trans. Aerosp. Electron. Syst. 2025, 61, 1610–1629. [Google Scholar] [CrossRef] [Scilit]
- Xing, X.; Xia, H. Multi-agent reinforcement learning with layered autonomy and collaboration for enhanced collaborative confrontation. Chin. J. Aeronaut. 2026, 39, 103747. [Google Scholar] [CrossRef] [Scilit]
- Pan, Z.; Zhang, C.; Xia, Y.; Xiong, H.; Shao, X. An Improved Artificial Potential Field Method for Path Planning and Formation Control of the Multi-UAV Systems. IEEE Trans. Circuits Syst. II Express Briefs 2022, 69, 1129–1133. [Google Scholar] [CrossRef] [Scilit]
- Huang, J.; Guo, Y.; Yuan, H.; Cui, X.; Chen, X. Dynamic Resource Allocation in Air-Ground Defense via Heterogeneous Multi-Agent Tracking with Cross-Temporal State Rewards. Phys. Commun. 2026, 75, 103019. [Google Scholar] [CrossRef] [Scilit]
- Yu, C.; Velu, A.; Vinitsky, E.; Gao, J.; Wang, Y.; Bayen, A.; Wu, Y. The Surprising Effectiveness of PPO in Cooperative Multi-Agent Games. Adv. Neural Inf. Process. Syst. 2022, 35, 24611–24624. [Google Scholar] [CrossRef] [Scilit]














| Agent Type | Defender | Attacker | GMT |
|---|---|---|---|
| Max Speed () | 1.0 | 0.8 | 0.5 |
| Max Acceleration () | 1.0 | 0.8 | 0.5 |
| Initial Quantity | 2 | 1 | 1 |
| Max Quantity | 5 | 5/7 | 1 |
| Initial Energy | 100 | / | 100 |
| Energy Utilization Efficiency () | 0.8 | / | 0.9 |
| Propulsion Efficiency Coefficient () | 0.75 | / | 0.5 |
| Deceleration Coefficient () | 0.1 | / | 0.08 |
| Minimum Operating Energy () | 10 | / | 5 |
| Capture Radius | 1.0 | 1.5 | / |
| Parameter | Description | Default |
|---|---|---|
| Weights for | 0.4, 0.4, 0.2 | |
| Penalty coefficients (coverage, excess, repeat) | 0.5, 0.3, 0.2 | |
| Weights for | 0.5, 0.3, 0.2 | |
| Weights for defender’s | 0.6, 0.3, 0.1 | |
| Weights for GMT’s | 0.7, 0.3 | |
| Maximum distance for threat evaluation | 10.0 | |
| Maximum distance for a valid match | 8.0 | |
| Clipping bounds for individual rewards | , 10 |
| Method | 5D_vs_5A | 5D_vs_7A | 3D_vs_5A | 3D_vs_7A | ||||
|---|---|---|---|---|---|---|---|---|
| IR (%) | TE | IR (%) | TE | IR (%) | TE | IR (%) | TE | |
| HMAPPO | 41.77 ± 4.10 | 154.77 ± 0.23 | 38.80 ± 4.80 | 152.60 ± 0.28 | 45.57 ± 3.66 | 115.97 ± 0.11 | 39.64 ± 4.98 | 110.44 ± 0.12 |
| Auction+HMAPPO | 31.22 ± 3.22 | 145.52 ± 0.18 | 28.17 ± 2.19 | 214.50 ± 0.13 | 48.69 ± 4.75 | 172.32 ± 0.05 | 46.54 ± 1.23 | 178.71 ± 0.16 |
| Hungarian+HMAPPO | 77.76 ± 5.66 | 222.58 ± 0.46 | 68.03 ± 4.01 | 244.07 ± 1.30 | 65.88 ± 3.54 | 151.67 ± 0.10 | 57.83 ± 3.74 | 146.55 ± 0.07 |
| HH-MAPPO | 93.64 ± 1.43 | 259.97 ± 0.77 | 84.80 ± 2.82 | 232.52 ± 1.60 | 79.49 ± 2.02 | 154.08 ± 0.09 | 80.75 ± 3.33 | 150.67 ± 0.04 |
| Method | 3D_vs_5A | 3D_vs_7A | 5D_vs_5A | 5D_vs_7A |
|---|---|---|---|---|
| Hungarian+HMAPPO | 79.76 ± 3.32 | 77.15 ± 3.05 | 92.64 ± 4.27 | 88.00 ± 3.16 |
| HH-MAPPO (Ours) | 95.97 ± 0.68 | 86.07 ± 3.58 | 93.24 ± 2.28 | 93.57 ± 1.80 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Huang, J.; Guo, Y.; Chen, X.; Wei, J.; Yi, J.; Chen, X.; Chen, L. HH-MAPPO: A Hierarchical Reinforcement Learning Framework for Dynamic-Scale Target–Attacker–Defender Games. Entropy 2026, 28, 793. https://doi.org/10.3390/e28070793
Huang J, Guo Y, Chen X, Wei J, Yi J, Chen X, Chen L. HH-MAPPO: A Hierarchical Reinforcement Learning Framework for Dynamic-Scale Target–Attacker–Defender Games. Entropy. 2026; 28(7):793. https://doi.org/10.3390/e28070793
Chicago/Turabian StyleHuang, Junhui, Yan Guo, Xiliang Chen, Jianyu Wei, Jiawei Yi, Xinliang Chen, and Lifeng Chen. 2026. "HH-MAPPO: A Hierarchical Reinforcement Learning Framework for Dynamic-Scale Target–Attacker–Defender Games" Entropy 28, no. 7: 793. https://doi.org/10.3390/e28070793
APA StyleHuang, J., Guo, Y., Chen, X., Wei, J., Yi, J., Chen, X., & Chen, L. (2026). HH-MAPPO: A Hierarchical Reinforcement Learning Framework for Dynamic-Scale Target–Attacker–Defender Games. Entropy, 28(7), 793. https://doi.org/10.3390/e28070793


