Energy-Harvesting-Assisted UAV Swarm Anti-Jamming Communication Based on Multi-Agent Reinforcement Learning
Highlights
- This paper proposes an energy-harvesting-assisted anti-jamming communication framework for unmanned aerial vehicle (UAV) swarm networks, aiming to optimize the long-term trade-off between transmission success rate and energy consumption.
- A multi-agent reinforcement learning-based solution is developed, enabling each cluster head to independently learn its channel selection, power control, and energy-harvesting time allocation policies through sharing spatiotemporal correlations among neighboring nodes.
- The proposed solution shows superior performance in simulations, achieving a higher transmission success rate with lower energy consumption across varying numbers of channels.
- The research provides a novel anti-jamming framework for intelligent spectrum management in dynamic jamming environments, offering guidance for the design of practical anti-jamming and energy-efficient UAV networks.
Abstract
1. Introduction
1.1. Motivation and Contribution
- We formulate a multi-domain energy-efficient anti-jamming communication optimization problem to achieve a long-term trade-off between data transmission success rate and energy consumption. Furthermore, this problem is modeled as a finite-horizon decentralized partially observable Markov decision process (Dec-POMDP).
- We propose an energy-harvesting-assisted anti-jamming communication scheme based on the multi-agent independent advantage actor–critic (IA2C) approach. In this scheme, each cluster head (CH) obtains discounted observations from neighboring CHs by applying a spatial discount factor based on the network topology. As a result, each CH agent can sequentially reduce the influence of decisions made by other CH agents on its own anti-jamming decision-making. Moreover, each CH maintains its own actor–critic network and performs distributed training using advantage-based updates.
- Simulation results show that the proposed scheme effectively improves the data transmission success rate while reducing energy consumption. In addition, the proposed scheme consistently achieves superior performance under varying numbers of channels.
1.2. Related Work
2. System Model and Problem Formulation
2.1. Network Model
2.2. Wireless Transmission Model
2.3. Mobility Model
2.4. Energy Harvesting Model
2.5. Problem Formulation
3. Proposed Energy-Harvesting-Aided Anti-Jamming Communication Framework
3.1. Framework
- Wideband Spectrum Sensing Phase : Each CH first performs wideband spectrum sensing to perceive the current spectral environment and obtain relevant channel information, while also collecting its local state information for decision-making.
- Decision-making phase : Each CH allocates the time between energy harvesting and data transmission based on the observed intra-cluster channel information and its residual energy, and determines the transmission channel and transmit power for communicating with its CMs.
- Execution phase : Each CH first harvests energy during the energy-harvesting phase according to the selected time-splitting ratio. After the harvesting interval, it switches to the transmission phase and initiates synchronized data transmission to all its CMs on the selected channel with the selected transmit power.
- Learning Phase : The CHs collaboratively learn the jammer’s channel-interference patterns and the energy-loss levels from environmental feedback, thereby providing training information for energy-efficient anti-jamming decisions in the next TS.
3.2. Dec-POMDP Design
3.2.1. Observation Space
3.2.2. Action Space
3.2.3. System Reward
4. Proposed Energy-Harvesting-Assisted Anti-Jamming Communication Scheme Based on Multi-Agent IA2C
4.1. Time–Space-Based Extended Dec-POMDP Model
4.2. Energy-Harvesting-Assisted Multi-Agent IA2C Anti-Jamming Communication Algorithm
| Algorithm 1 Proposed Energy-Harvesting-Assisted Multi-Agent IA2C Anti-Jamming Algorithm |
|
5. Simulation Results
- Consensus-update-based anti-jamming scheme (CU) [34]: Each CH agent adopts a consensus-update mechanism to perform distributed training, allocating the time for energy harvesting and data transmission, while simultaneously deciding the channel selection and transmit power.
- Multi-agent DDRQN-based anti-jamming scheme (DDRQN) [35]: Each CH agent employs the multi-agent DDRQN method. It makes independent decisions using only local observations, including the time allocation ratio between energy harvesting and data transmission, the transmission channel, and the transmit power.
- Received-signal-strength-indication-based anti-jamming scheme (RSSI): Each CH agent randomly allocates the time between energy harvesting and data transmission, and then performs data transmission on the channel with the smallest path loss using the maximum transmit power.
- Random scheme: Each CH agent randomly allocates the time between energy harvesting and data transmission according to its local observations, and randomly selects the transmission channel and transmit power.
5.1. Comparison of Convergence Performance of Different Schemes
5.2. Performance Comparison Versus the Number of Available Channels
6. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Wei, Z.; Zhu, M.; Zhang, N.; Wang, L.; Zou, Y.; Meng, Z.; Wu, H.; Feng, Z. UAV-assisted data collection for Internet of Things: A survey. IEEE Internet Things J. 2022, 9, 15460–15483. [Google Scholar] [CrossRef] [Scilit]
- Li, B.; Fei, Z.; Zhang, Y.; Guizani, M. Secure UAV communication networks over 5G. IEEE Wirel. Commun. 2019, 26, 114–120. [Google Scholar] [CrossRef] [Scilit]
- Alqudsi, Y.; Makaraci, M. UAV swarms: Research, challenges, and future directions. J. Eng. Appl. Sci. 2025, 72, 12. [Google Scholar] [CrossRef] [Scilit]
- Fotouhi, A.; Qiang, H.; Ding, M.; Hassan, M.; Giordano, L.G.; Garcia-Rodriguez, A.; Yuan, J. Survey on UAV cellular communications: Practical aspects, standardization advancements, regulation, and security challenges. IEEE Commun. Surv. Tutor. 2019, 21, 3417–3442. [Google Scholar] [CrossRef] [Scilit]
- Al-Hourani, A. On the probability of line-of-sight in urban environments. IEEE Wirel. Commun. Lett. 2020, 9, 1178–1181. [Google Scholar] [CrossRef] [Scilit]
- Hua, M.; Wang, Y.; Wu, Q.; Dai, H.; Huang, Y.; Yang, L. Energy-efficient cooperative secure transmission in multi-UAV-enabled wireless networks. IEEE Trans. Veh. Technol. 2019, 68, 7761–7775. [Google Scholar] [CrossRef] [Scilit]
- Azari, M.M.; Geraci, G.; Garcia-Rodriguez, A.; Pollin, S. UAV-to-UAV communications in cellular networks. IEEE Trans. Wirel. Commun. 2020, 19, 6130–6144. [Google Scholar] [CrossRef] [Scilit]
- Yin, Z.; Li, J.; Wang, Z.; Qian, Y.; Lin, Y.; Shu, F.; Chen, W. UAV communication against intelligent jamming: A stackelberg game approach with federated reinforcement learning. IEEE Trans. Green Commun. Netw. 2024, 8, 1796–1808. [Google Scholar] [CrossRef] [Scilit]
- Elleuch, I.; Pourranjbar, A.; Kaddoum, G. A novel distributed multi-agent reinforcement learning algorithm against jamming attacks. IEEE Commun. Lett. 2021, 25, 3204–3208. [Google Scholar] [CrossRef] [Scilit]
- Lv, Z.; Xiao, L.; Du, Y.; Niu, G.; Xing, C.; Xu, W. Multi-agent reinforcement learning based UAV swarm communications against jamming. IEEE Trans. Wirel. Commun. 2023, 22, 9063–9075. [Google Scholar] [CrossRef] [Scilit]
- Yao, F.; Jia, L. A collaborative multi-agent reinforcement learning anti-jamming algorithm in wireless networks. IEEE Wirel. Commun. Lett. 2019, 8, 1024–1027. [Google Scholar] [CrossRef] [Scilit]
- Zhou, Q.; Li, Y.; Niu, Y. Intelligent anti-jamming communication for wireless sensor networks: A multi-agent reinforcement learning approach. IEEE Open J. Commun. Soc. 2021, 2, 775–784. [Google Scholar] [CrossRef] [Scilit]
- Yin, Z.; Lin, Y.; Zhang, Y.; Qian, Y.; Shu, F.; Li, J. Collaborative multiagent reinforcement learning aided resource allocation for UAV anti-jamming communication. IEEE Internet Things J. 2022, 9, 23995–24008. [Google Scholar] [CrossRef] [Scilit]
- Zhang, Y.; Xu, Y.; Xu, Y.; Yang, Y.; Luo, Y.; Wu, Q.; Liu, X. A multi-leader one-follower Stackelberg game approach for cooperative anti-jamming: No pains, no gains. IEEE Commun. Lett. 2018, 22, 1680–1683. [Google Scholar] [CrossRef] [Scilit]
- Saif, A.; Dimyati, K.; Noordin, K.A.; Deepak, G.; Shah, N.S.M.; Abdullah, Q.; Mohamad, M. An efficient energy harvesting and optimal clustering technique for sustainable postdisaster emergency communication systems. IEEE Access 2021, 9, 78188–78202. [Google Scholar] [CrossRef] [Scilit]
- Wang, C.; Yu, W.; Zhu, F.; Ou, J.; Fan, C.; Ou, J.; Fan, D. UAV-Aided Multiuser Mobile Edge Computing Networks with Energy Harvesting. Wirel. Commun. Mob. Comput. 2022, 2022, 6723403. [Google Scholar] [CrossRef] [Scilit]
- Liao, C.; Xu, K.; Hu, G.; Xia, X.; Wei, C.; Xie, W.; Li, C.; Wang, Y. Game theory and multi-agent DRL based anti-jamming transmission for integrated air-ground network. IEEE Trans. Veh. Technol. 2024, 73, 19565–19581. [Google Scholar] [CrossRef] [Scilit]
- Yan, L.; Zhijuan, W.; Nuoheng, P.; Tianyu, Z.; Yijin, Z.; Feng, S.; Jun, L. Defending against jamming and interference for Internet of UAVs using cooperative multi-agent reinforcement learning with mutual information. China Commun. 2025, 22, 220–237. [Google Scholar] [CrossRef] [Scilit]
- Wang, W.; Lv, Z.; Lu, X.; Zhang, Y.; Xiao, L. Distributed reinforcement learning based framework for energy-efficient UAV relay against jamming. Intell. Converg. Netw. 2021, 2, 150–162. [Google Scholar] [CrossRef] [Scilit]
- Ma, N.; Xu, K.; Xia, X.; Wei, C.; Su, Q.; Shen, M.; Xie, W. Reinforcement learning-based dynamic anti-jamming power control in UAV networks: An effective jamming signal strength based approach. IEEE Commun. Lett. 2022, 26, 2355–2359. [Google Scholar] [CrossRef] [Scilit]
- Song, F.; Deng, M.; Xing, H.; Liu, Y.; Ye, F.; Xiao, Z. Energy-efficient trajectory optimization with wireless charging in UAV-assisted MEC based on multi-objective reinforcement learning. IEEE Trans. Mob. Comput. 2024, 23, 10867–10884. [Google Scholar] [CrossRef] [Scilit]
- Sekander, S.; Tabassum, H.; Hossain, E. Statistical performance modeling of solar and wind-powered UAV communications. IEEE Trans. Mob. Comput. 2020, 20, 2686–2700. [Google Scholar] [CrossRef] [Scilit]
- Yang, L.; Chen, J.; Hasna, M.O.; Yang, H.C. Outage performance of UAV-assisted relaying systems with RF energy harvesting. IEEE Commun. Lett. 2018, 22, 2471–2474. [Google Scholar] [CrossRef] [Scilit]
- Peng, H.; Wang, L.C. Energy harvesting reconfigurable intelligent surface for UAV based on robust deep reinforcement learning. IEEE Trans. Wirel. Commun. 2023, 22, 6826–6838. [Google Scholar] [CrossRef] [Scilit]
- Yang, H.; Lin, K.; Xiao, L.; Zhao, Y.; Xiong, Z.; Han, Z. Energy harvesting UAV-RIS-assisted maritime communications based on deep reinforcement learning against jamming. IEEE Trans. Wirel. Commun. 2024, 23, 9854–9868. [Google Scholar] [CrossRef] [Scilit]
- Dang-Ngoc, H.; Nguyen, D.N.; Ho-Van, K.; Hoang, D.T.; Dutkiewicz, E.; Pham, Q.V.; Hwang, W.J. Secure swarm UAV-assisted communications with cooperative friendly jamming. IEEE Internet Things J. 2022, 9, 25596–25611. [Google Scholar] [CrossRef] [Scilit]
- Ji, B.; Li, Y.; Cao, D.; Li, C.; Mumtaz, S.; Wang, D. Secrecy performance analysis of UAV assisted relay transmission for cognitive network with energy harvesting. IEEE Trans. Veh. Technol. 2020, 69, 7404–7415. [Google Scholar] [CrossRef] [Scilit]
- Cao, K.; Wang, B.; Ding, H.; Tian, J. Adaptive cooperative jamming for secure communication in energy harvesting relay networks. IEEE Wirel. Commun. Lett. 2019, 8, 1316–1319. [Google Scholar] [CrossRef] [Scilit]
- Cho, S.; Lee, K.; Kang, B.; Koo, K.; Joe, I. Weighted harvest-then-transmit: UAV-enabled wireless powered communication networks. IEEE Access 2018, 6, 72212–72224. [Google Scholar] [CrossRef] [Scilit]
- Li, Y.; Wang, W.; Gao, H.; Wu, Y.; Su, M.; Wang, J.; Liu, Y. Air-to-ground 3D channel modeling for UAV based on Gauss-Markov mobile model. AEU-Int. J. Electron. Commun. 2020, 114, 152995. [Google Scholar] [CrossRef] [Scilit]
- Tsao, K.Y.; Girdler, T.; Vassilakis, V.G. A survey of cyber security threats and solutions for UAV communications and flying ad-hoc networks. Ad Hoc Netw. 2022, 133, 102894. [Google Scholar] [CrossRef] [Scilit]
- Le, N.P.; Huang, X.; Dutkiewicz, E.; Ritz, C.; Phung, S.L.; Bouzerdoum, A.; Franklin, D.; Hanzo, L. Energy-harvesting aided unmanned aerial vehicles for reliable ground user localization and communications under Lognormal-Nakagami-m fading channels. IEEE Trans. Veh. Technol. 2021, 70, 1632–1647. [Google Scholar] [CrossRef] [Scilit]
- Liu, Q.; Li, M.; Yang, J.; Lv, J.; Hwang, K.; Hossain, M.S.; Muhammad, G. Joint power and time allocation in energy harvesting of UAV operating system. Comput. Commun. 2020, 150, 811–817. [Google Scholar] [CrossRef] [Scilit]
- Zhang, K.; Yang, Z.; Liu, H.; Zhang, T.; Basar, T. Fully decentralized multi-agent reinforcement learning with networked agents. In Proceedings of the International Conference on Machine Learning, Stockholm, Sweden, 10–15 July 2018; PMLR: Cambridge, MA, USA, 2018; Volume 80, pp. 5872–5881. [Google Scholar]
- Foerster, J.N.; Assael, Y.M.; de Freitas, N.; Whiteson, S. Learning to communicate to solve riddles with deep distributed recurrent q-networks. arXiv 2016, arXiv:1602.02672. [Google Scholar] [CrossRef] [Scilit]









| Parameter | Notion | Value |
|---|---|---|
| Number of UAV clusters | M | 4 |
| Number of cluster members | I | 2 |
| Number of jammers | N | 4 |
| Number of available channels | L | 6 |
| Transmit power of UAVs | dBm | |
| Transmit power of jammers | 30 dBm | |
| Antenna gain of UAVs/jammers | 3 dBi | |
| Power of AWGN | dBm | |
| Timeslot | s | |
| Execution stage | s | |
| Time ratio factor | ||
| Jamming period | s | |
| Correlation weight factor | ||
| Gaussian variable | 2 | |
| Maximum communication range | 100 m | |
| The radius of motion of the RP | 99 m | |
| The radius of movement of CMs | 1 m | |
| Bandwidth | B | 180 kHz |
| Data packet size | K | 300 KB |
| Energy Harvesting Efficiency | ||
| Spatial discount factor | 1 |
| Parameter | Notion | Value |
|---|---|---|
| Number of episodes | E | 5000 |
| Maximum number of steps per episode | T | 400 |
| Learning rate parameter | 0.002 | |
| Batch size | 32 | |
| Discount factor | 0.99 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Li, Y.; Zhao, T.; Wu, Z.; Lin, Y.; Zhang, Y. Energy-Harvesting-Assisted UAV Swarm Anti-Jamming Communication Based on Multi-Agent Reinforcement Learning. Drones 2026, 10, 294. https://doi.org/10.3390/drones10040294
Li Y, Zhao T, Wu Z, Lin Y, Zhang Y. Energy-Harvesting-Assisted UAV Swarm Anti-Jamming Communication Based on Multi-Agent Reinforcement Learning. Drones. 2026; 10(4):294. https://doi.org/10.3390/drones10040294
Chicago/Turabian StyleLi, Yongfang, Tianyu Zhao, Zhijuan Wu, Yan Lin, and Yijin Zhang. 2026. "Energy-Harvesting-Assisted UAV Swarm Anti-Jamming Communication Based on Multi-Agent Reinforcement Learning" Drones 10, no. 4: 294. https://doi.org/10.3390/drones10040294
APA StyleLi, Y., Zhao, T., Wu, Z., Lin, Y., & Zhang, Y. (2026). Energy-Harvesting-Assisted UAV Swarm Anti-Jamming Communication Based on Multi-Agent Reinforcement Learning. Drones, 10(4), 294. https://doi.org/10.3390/drones10040294

