Multi-Agent Deep Reinforcement Learning Framework for Efficient Aerial Wildfire Fighting †
Abstract
1. Introduction
Motivation
- How can Reinforcement Learning be used to enable the learning of suppression strategies evaluated by integration into an agent-based wildfire simulation?
- To what extent do decomposed reward functions and observations enhance the interpretability of Reinforcement Learning-based suppression strategies?
2. Materials and Methods
2.1. SoSID Toolkit
- Indirect Attack: The objective is to contain the fire by establishing a continuous suppression barrier around the fire perimeter.
- Direct Attack: Agents apply the suppressant to extinguish burning cells directly. This tactic relies on predefined points of interest determined by a weighted function considering water sources, residential areas, agent position, and spread rates.
2.2. Deep Reinforcement Learning
2.3. Dual Decomposition Framework
3. Implementation
3.1. Action Space
3.2. Observation Space
- Local Grid: Enables the agents to perceive their immediate surroundings by providing information about the states of neighboring cells. Each cell within the observation grid corresponds to discrete environmental states, including suppressed, nonflammable, and various burning stages. The size of the local observation grid can be configured at the start of each experiment to adjust the agent’s field of view and thus control how much of the environment is perceived when making decisions.
- Closest N Fire Direction: Provides direction vectors to the N nearest burning cells, with N specified during experiment initialization.
- Cardinal Directions: Provides four direction vectors from the agent’s position to the cardinal extremities of the wildfire perimeter (northernmost, southernmost, easternmost, and westernmost points of the fire boundary), offering a compact representation of the directional distribution of fires.
- Ignition Center Direction: Provides a direction vector from the agent’s position to the ignition center of the fire.
3.3. Reward Function
- Proximity Reward: The Proximity Reward provides dense feedback based on an agent’s distance to the wildfire. Increasing the distance results in a penalty, while decreasing it yields a reward. This encourages agents to systematically locate and approach the fire. Unlike sparse rewards, dense feedback offers continuous guidance, accelerating convergence by reducing the exploration burden. In wildfire suppression, such proximity-based signals are essential for learning effective navigation toward active fire fronts. Without such intermediate rewards, agents would have difficulty recognizing the connection between their movement decisions and the ultimate goal.
- Suppression Reward: The purpose of the Suppression Reward is to encourage agents to actively suppress the wildfire. This reward is calculated directly based on the number of successfully suppressed cells.
- Connected Suppression Reward: Provides the reward based on whether the current suppression is adjacent to already suppressed cells. The goal is to encourage the agent to establish a continuous line of suppression rather than suppressing individual cells in isolation, thereby promoting more effective containment strategies.
- Inactivity Penalty: Applies a negative reward when an agent is adjacent to a burning cell but fails to execute a suppression action. This discourages inactive behavior near the fire front.
- Timestep Penalty: Applies a small penalty at every simulation step to encourage efficient suppression. As the cumulative reward decreases with longer episode lengths, strategies that extinguish the fire more efficiently obtain higher overall returns, aligning policy optimization with operational efficiency.
- Impact Reward: Provides a positive reward based on the reduction in the wildfire’s overall spread rate. This component serves as a long-term incentive, thereby rewarding suppression actions that do not yield an immediate extinction reward if they successfully slow the global propagation of the fire.
3.4. Policy
4. Results
4.1. Experiments
4.2. Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| RL | Reinforcement Learning |
| MARL | Multi-Agent Reinforcement Learning |
| DRL | Deep Reinforcement Learning |
| cc | correction-coefficient |
| PPO | Proximal Policy Optimization |
| SoS | System-of-Systems |
References
- European Forest Fire Information System (EFFIS). Available online: https://forest-fire.emergency.copernicus.eu/ (accessed on 30 September 2025).
- Tanha, R.S.; Hooshmand, F. UAV-based Firefighting by Multi-agent Reinforcement Learning. In Proceedings of the 13th International Conference on Computer and Knowledge Engineering (ICCKE), Mashhad, Iran, 1–2 November 2023. [Google Scholar]
- Blais, M.-A.; Akhloufi, M.A. Drone Swarm Coordination Using Reinforcement Learning for Efficient Wildfires Fighting. SN Comput. Sci. 2024, 5, 314. [Google Scholar] [CrossRef] [Scilit]
- Juozapaitis, Z.; Koul, A.; Fern, A.; Erwig, M.; Doshi-Velez, F. Explainable Reinforcement Learning via Reward Decomposition. In Proceedings of the IJCAI/ECAI Workshop on Explainable Artificial Intelligence, Macao, China, 11 August 2019; Available online: https://www.semanticscholar.org/paper/Explainable-Reinforcement-Learning-via-Reward-Juozapaitis-Koul/d47677337b1083d6bfa940748da0780b2c9faf7d (accessed on 30 September 2025).
- Kilkis, S.; Prakasha, P.S.; Naeem, N.; Nagel, B. A Python Modelling and Simulation Toolkit for Rapid Development of System of Systems Inverse Design (SoSID) Case Studies. In Proceedings of the AIAA AVIATION 2021 Forum, Virtual, 2–6 August 2021. [Google Scholar]
- Prakasha, P.S.; Naeem, N. Establishing a Collaborative Open Source Agent Based Simulation Environment for System of Systems Aviation Problems. 2026; manuscript in preparation; to be submitted.
- Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal Policy Optimization Algorithms. arXiv 2017, arXiv:1707.06347. [Google Scholar] [CrossRef] [Scilit]
- Berner, C.; Chen, G.; Debiak, B.; Edwards, R.; Farr, J.; Goh, C.; Gray, A.; Ho, Y.; Igl, F.; Laskin, M.; et al. Learning to play Dota 2 with self-play reinforcement learning. In Proceedings of the Conference on Neural Information Processing Systems (NeurIPS), Vancouver, BC, Canada, 8–14 December 2019. [Google Scholar]
- Yu, C.; Saran, A.; Liu, Z.; Jain, A.; Ranganath, R.; Iqbal, S.T.; Kumar, S. The Surprising Effectiveness of PPO in Multi-Agent RL. In Proceedings of the Conference on Neural Information Processing Systems (NeurIPS), Virtual, 6–14 December 2021. [Google Scholar]
- Wu, Z.; Liang, E.; Luo, M.; Mika, S.; Gonzalez, J.E.; Stoica, I. RLlib Flow: Distributed Reinforcement Learning is a Dataflow Problem. In Proceedings of the Conference on Neural Information Processing Systems (NeurIPS), Virtual, 6–14 December 2021; Available online: https://proceedings.neurips.cc/paper/2021/file/2bce32ed409f5ebcee2a7b417ad9beed-Paper.pdf (accessed on 26 June 2025).
- Liang, E.; Liaw, R.; Nishihara, R.; Moritz, P.; Fox, R.; Goldberg, K.; Gonzalez, J.E.; Jordan, M.I.; Stoica, I. RLlib: Abstractions for Distributed Reinforcement Learning. arXiv 2018, arXiv:1712.09381. [Google Scholar] [CrossRef] [Scilit]
- Moritz, P.; Nishihara, R.; Wang, S.; Tumanov, A.; Liang, E.; Elibol, I.; Yang, Z.; Paul, W.; Jordan, M.I.; Stoica, I. Ray: A Distributed Framework for Emerging AI Applications. In Proceedings of the 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 2018), Carlsbad, CA, USA, 8–10 October 2018. [Google Scholar]



Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Bardtke, L.; Naeem, N.; Kalliatakis, N.; Prakasha, P.S.; Clemen, T. Multi-Agent Deep Reinforcement Learning Framework for Efficient Aerial Wildfire Fighting. Eng. Proc. 2026, 133, 188. https://doi.org/10.3390/engproc2026133188
Bardtke L, Naeem N, Kalliatakis N, Prakasha PS, Clemen T. Multi-Agent Deep Reinforcement Learning Framework for Efficient Aerial Wildfire Fighting. Engineering Proceedings. 2026; 133(1):188. https://doi.org/10.3390/engproc2026133188
Chicago/Turabian StyleBardtke, Leonard, Nabih Naeem, Nikolaos Kalliatakis, Prajwal Shiva Prakasha, and Thomas Clemen. 2026. "Multi-Agent Deep Reinforcement Learning Framework for Efficient Aerial Wildfire Fighting" Engineering Proceedings 133, no. 1: 188. https://doi.org/10.3390/engproc2026133188
APA StyleBardtke, L., Naeem, N., Kalliatakis, N., Prakasha, P. S., & Clemen, T. (2026). Multi-Agent Deep Reinforcement Learning Framework for Efficient Aerial Wildfire Fighting. Engineering Proceedings, 133(1), 188. https://doi.org/10.3390/engproc2026133188

