Balancing Energy and Mission Time in UAV Site Servicing on Graph Maps Through Dynamic Battery-Threshold Double Deep Q-Learning
Abstract
1. Introduction
1.1. Literature Review
1.2. Contributions and Paper Outline
- We propose a Double Deep Q-Network (DDQN) framework with a graph-aware Dynamic Battery-Threshold (DBT) reward based on Dijkstra shortest-path distances to the nearest recharge node. Unlike fixed battery safety margins, DBT defines a topology-aware energy reserve that adapts to heterogeneous graphs without per-map tuning, resulting, as proof of our experiments, in a more efficient trade-off between mission performance and battery safety. We adopt a value-based RL algorithm [28] rather than a policy-gradient one such as PPO, which has previously been shown to require ad hoc convergence tricks, like action masking and discount-factor scheduling, on comparable problems with large maps and sparse rewards [5].
- We analyze the scalability of the proposed DDQN policy by increasing the graph dimension and the corresponding network size. To make policy training tractable as the state–action space grows with map complexity, we couple DDQN with PER, a combination already adopted for battery-constrained UAV path planning [21]. Across all benchmark graphs, the learned policy attains a 100% task completion rate with always enough final battery to at least reach a charging station after the successful fix. We further observe, in a continuous multi-task deployment where the agent’s battery is not reset from one test to another, a 0% discharge rate while minimizing mission length w.r.t. fixed-threshold policies, either too aggressive or conservative.
- We compare the proposed DBT-DDQN policy against a model-based, pseudo-optimal planner. The results show that the RL policy achieves competitive mission efficiency, degrading less with increasing graph complexity and stochastic battery dynamics, all while maintaining significantly lower computational requirements.
2. Energy-Aware Site-Servicing on a Graph
2.1. Graph and Mission Model
- Target nodes : Locations where failures occur and must be serviced.
- Depot nodes : Locations where the UAV collects servicing tools requested by targets. Tools can represent more general features that the UAV must possess in order to solve a specific requested task (e.g., a water tank in case of a wildfire, or package transportation, etc.).
- Charging station nodes : Locations where the UAV can recharge its battery, allowing uninterrupted autonomous operability. Recharge can occur autonomously via a landing policy on a charging platform or be conducted on-site by human operators.
- Regular nodes : Transit nodes with no special mission function.
- These subsets are not necessarily disjoint: a node simultaneously offers tool pick-up and recharging capabilities. The only structural requirement is , i.e., failures occur exclusively at non-service locations.
2.2. Agent’s Energy Model
2.3. Problem Statement
3. Reinforcement Learning Policy Design
3.1. MDP Formulation
3.1.1. State Space
3.1.2. Action Space
3.1.3. Agent’s State Model
3.2. The Dynamic Battery-Threshold Reward
3.3. DDQN Training
4. Simulation Setup
4.1. Benchmark Evaluation Maps
4.2. Policies Implementation Details
5. Numerical Results
5.1. Training and Monte Carlo Evaluation Across Graphs
5.2. Sample DBT-DDQN Trajectories on Graphs
5.3. Ablation Study: Hidden MLP Size
5.4. The Contribution of the DBT Reward
5.5. Pseudo-Optimality Analysis
6. Discussions and Conclusions
Author Contributions
Funding
Institutional Review Board Statement
Informed Consent Statement
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| UAV | Unmanned Aerial Vehicle |
| DDQN | Double Deep Q-Network |
| DBT | Dynamic Battery Threshold |
| RL | Reinforcement Learning |
| MDP | Markov Decision Process |
| MLP | Multi-Layer Perceptron |
| MCTS | Monte Carlo Tree Search |
| PER | Prioritized Experience Replay |
| MILP | Mixed-Integer Linear Programming |
| UGV | Unmanned Ground Vehicle |
| RTH | Return To Home |
| PPO | Proximal Policy Optimization |
| IDDQN | Improved Double Deep Q-network |
References
- Mohsan, S.A.H.; Othman, N.Q.H.; Li, Y.; Alsharif, M.H.; Khan, M.A. Unmanned aerial vehicles (UAVs): Practical aspects, applications, open challenges, security issues, and future trends. Intell. Serv. Robot. 2023, 16, 109–137. [Google Scholar] [CrossRef] [PubMed]
- Bouček, Z.; Flídr, M. Mission Planner for UAV Battery Replacement. In Proceedings of the 2024 IEEE International Conference on Multisensor Fusion and Integration for Intelligent Systems (MFI), Pilsen, Czech Republic, 4–6 September 2024; pp. 1–6. [Google Scholar] [CrossRef]
- Zeng, Y.; Zhang, R. Energy-efficient UAV communication with trajectory optimization. IEEE Trans. Wirel. Commun. 2017, 16, 3747–3760. [Google Scholar] [CrossRef]
- Theile, M.; Bayerlein, H.; Nai, R.; Gesbert, D.; Caccamo, M. UAV Coverage Path Planning under Varying Power Constraints using Deep Reinforcement Learning. In Proceedings of the 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), Las Vegas, NV, USA, 25–29 October 2020; pp. 1444–1449. [Google Scholar] [CrossRef]
- Theile, M.; Bayerlein, H.; Caccamo, M.; Sangiovanni-Vincentelli, A.L. Learning to recharge: UAV coverage path planning through deep reinforcement learning. arXiv 2023, arXiv:2309.03157. [Google Scholar]
- Zhang, Y.; Zheng, H.; Zhai, X. Deep Reinforcement Learning Based UAV Mission Planning with Charging Module. In Proceedings of the 2023 4th International Conference on Computing, Networks and Internet of Things, Xiamen, China, 26–28 May 2023; pp. 658–662. [Google Scholar]
- Grando, L.; Jaramillo, J.F.G.; Leite, J.R.E.; Ursini, E.L. Modeling and Simulation of Battery Recharging for UAVs Applications: Smart Farming, Disaster Recovery, and Dengue Focus Detections. In Proceedings of the 2024 Winter Simulation Conference (WSC); IEEE: Piscataway, NJ, USA, 2024; pp. 2832–2843. [Google Scholar]
- Garey, M.R.; Johnson, D.S. A Guide to the Theory of NP-Completeness. In Computers and Intractability; W.H. Freeman and Company: New York, NY, USA, 1990; pp. 37–79. [Google Scholar]
- Sundar, K.; Rathinam, S. Algorithms for routing an unmanned aerial vehicle in the presence of refueling depots. IEEE Trans. Autom. Sci. Eng. 2013, 11, 287–294. [Google Scholar] [CrossRef]
- Ribeiro, R.G.; Cota, L.P.; Euzébio, T.A.M.; Ramírez, J.A.; Guimarães, F.G. Unmanned-Aerial-Vehicle Routing Problem With Mobile Charging Stations for Assisting Search and Rescue Missions in Postdisaster Scenarios. IEEE Trans. Syst. Man Cybern. Syst. 2022, 52, 6682–6696. [Google Scholar] [CrossRef]
- Yu, K.; Budhiraja, A.K.; Tokekar, P. Algorithms for Routing of Unmanned Aerial Vehicles with Mobile Recharging Stations. In Proceedings of the 2018 IEEE International Conference on Robotics and Automation (ICRA), Brisbane, Australia, 21–25 May 2018; pp. 5720–5725. [Google Scholar] [CrossRef]
- Alyassi, R.; Khonji, M.; Karapetyan, A.; Chau, S.C.K.; Elbassioni, K.; Tseng, C.M. Autonomous Recharging and Flight Mission Planning for Battery-Operated Autonomous Drones. IEEE Trans. Autom. Sci. Eng. 2023, 20, 1034–1046. [Google Scholar] [CrossRef]
- Mohabbati-Kalejahi, N.; Alavi, S.; Toragay, O. A mixed-integer programming framework for drone routing and scheduling with flexible multiple visits in highway traffic monitoring. Mathematics 2025, 13, 2427. [Google Scholar] [CrossRef]
- Ramasamy, S.; Reddinger, J.P.F.; Dotterweich, J.M.; Childers, M.A.; Bhounsule, P.A. Cooperative route planning of multiple fuel-constrained unmanned aerial vehicles with recharging on an unmanned ground vehicle. In Proceedings of the 2021 International Conference on Unmanned Aircraft Systems (ICUAS); IEEE: Piscataway, NJ, USA, 2021; pp. 155–164. [Google Scholar]
- Maini, P.; Sundar, K.; Singh, M.; Rathinam, S.; Sujit, P.B. Cooperative Aerial–Ground Vehicle Route Planning With Fuel Constraints for Coverage Applications. IEEE Trans. Aerosp. Electron. Syst. 2019, 55, 3016–3028. [Google Scholar] [CrossRef]
- Chen, Z.; Hu, Z.; Bao, Z.; Xu, W. UAV Charging Station Planning and Route Optimization Considering Stochastic Delivery Demand. IEEE Trans. Transp. Electrif. 2024, 10, 9328–9341. [Google Scholar] [CrossRef]
- Fagundes-Junior, L.A.; de Carvalho, K.B.; Ferreira, R.S.; Brandão, A.S. Machine learning for unmanned aerial vehicles navigation: An overview. SN Comput. Sci. 2024, 5, 256. [Google Scholar] [CrossRef]
- Zhao, C.; Liu, J.; Yoon, S.U.; Li, X.; Li, H.; Zhang, Z. Energy constrained multi-agent reinforcement learning for coverage path planning. In Proceedings of the 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); IEEE: Piscataway, NJ, USA, 2023; pp. 5590–5597. [Google Scholar]
- Fu, H.; Li, Z.; Zhang, W.; Feng, Y.; Zhu, L.; Long, Y.; Li, J. Path Planning for Agricultural UAVs Based on Deep Reinforcement Learning and Energy Consumption Constraints. Agriculture 2025, 15, 943. [Google Scholar] [CrossRef]
- Fan, M.; Wu, Y.; Liao, T.; Cao, Z.; Guo, H.; Sartoretti, G.; Wu, G. Deep Reinforcement Learning for UAV Routing in the Presence of Multiple Charging Stations. IEEE Trans. Veh. Technol. 2023, 72, 5732–5746. [Google Scholar] [CrossRef]
- Ni, J.; Gu, Y.; Gu, Y.; Zhao, Y.; Shi, P. UAV coverage path planning with limited battery energy based on improved deep double Q-network. Int. J. Control Autom. Syst. 2024, 22, 2591–2601. [Google Scholar] [CrossRef]
- Chu, N.H.; Hoang, D.T.; Nguyen, D.N.; Van Huynh, N.; Dutkiewicz, E. Joint Speed Control and Energy Replenishment Optimization for UAV-Assisted IoT Data Collection with Deep Reinforcement Transfer Learning. IEEE Internet Things J. 2023, 10, 5778–5793. [Google Scholar] [CrossRef]
- Li, X.; Yao, L.; Li, M.; Zhang, B. Reinforcement Learning Based Collaborative Path Planning Research for UAVs and Unmanned Vehicles. In Proceedings of the International Conference on Machine Learning and Intelligent Computing; PMLR: Cambridge, MA, USA, 2025; pp. 595–603. [Google Scholar]
- Mondal, M.S.; Ramasamy, S.; Bhounsule, P. OptiRoute: A Heuristic-assisted Deep Reinforcement Learning Framework for UAV-UGV Collaborative Route Planning. arXiv 2023, arXiv:2309.09942. [Google Scholar]
- Mondal, M.S.; Ramasamy, S.; Humann, J.D.; Dotterweich, J.M.; Reddinger, J.P.F.; Childers, M.A.; Bhounsule, P. An attention-aware deep reinforcement learning framework for uav-ugv collaborative route planning. In Proceedings of the 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS); IEEE: Piscataway, NJ, USA, 2024; pp. 13687–13694. [Google Scholar]
- Mondal, M.S.; Ramasamy, S.; Rownak, R.; Russo, L.; Humann, J.D.; Dotterweich, J.M.; Bhounsule, P. Risk-Aware Energy-Constrained UAV-UGV Cooperative Routing Using Attention-Guided Reinforcement Learning. In Proceedings of the 2025 IEEE International Conference on Robotics and Automation (ICRA); IEEE: Piscataway, NJ, USA, 2025; pp. 13000–13006. [Google Scholar]
- Gemignani, G.; Bongiorni, M.; Pollini, L. An Energy-aware Decision-making scheme for Mobile Robots on a Graph map based on Deep Reinforcement Learning. In Proceedings of the 2024 18th International Conference on Control, Automation, Robotics and Vision (ICARCV); IEEE: Piscataway, NJ, USA, 2024; pp. 460–466. [Google Scholar]
- Van Hasselt, H.; Guez, A.; Silver, D. Deep reinforcement learning with double q-learning. In Proceedings of the AAAI Conference on Artificial Intelligence, Phoenix, AZ, USA, 12–17 February 2016; Volume 30. [Google Scholar]
- Meng, W.; Zhang, X.; Zhou, L.; Guo, H.; Hu, X. Advances in UAV path planning: A comprehensive review of methods, challenges, and future directions. Drones 2025, 9, 376. [Google Scholar] [CrossRef]
- Dai, W.; Rai, U.; Chiun, J.; Cao, Y.; Sartoretti, G. Heterogeneous multi-robot task allocation and scheduling via reinforcement learning. IEEE Robot. Autom. Lett. 2025, 10, 2654–2661. [Google Scholar] [CrossRef]
- Gemignani, G.; Casini, S.; Rosellini, V.; Bucchioni, G.; Pollini, L. Preliminary Design of Human-like Decentralised Task Assignment for Heterogeneous Unmanned Vehicles using Reinforcement Learning. In Proceedings of the AIAA SCITECH 2026 Forum, Orlando, FL, USA, 12–16 January 2026; p. 0326. [Google Scholar]
- Mnih, V.; Kavukcuoglu, K.; Silver, D.; Rusu, A.A.; Veness, J.; Bellemare, M.G.; Graves, A.; Riedmiller, M.; Fidjeland, A.K.; Ostrovski, G.; et al. Human-level control through deep reinforcement learning. Nature 2015, 518, 529–533. [Google Scholar] [CrossRef] [PubMed]
- Dijkstra, E.W. A note on two problems in connexion with graphs. In Edsger Wybe Dijkstra: His Life, Work, and Legacy; Association for Computing Machinery (ACM): New York, NY, USA; Morgan & Claypool: New York, NY, USA, 2022; pp. 287–290. [Google Scholar]
- Schaul, T.; Quan, J.; Antonoglou, I.; Silver, D. Prioritized experience replay. arXiv 2015, arXiv:1511.05952. [Google Scholar]
- Towers, M.; Kwiatkowski, A.; Balis, J.; De Cola, G.; Deleu, T.; Goulão, M.; Andreas, K.; Krimmel, M.; Kg, A.; Perez-Vicente, R.; et al. Gymnasium: A standard interface for reinforcement learning environments. Adv. Neural Inf. Process. Syst. 2026, 38. [Google Scholar] [CrossRef]









| Graph | N | M | ||||
|---|---|---|---|---|---|---|
| G1 | 35 | 4 | 2 | 2 | 3 | / |
| G2 | 64 | 6 | 3 | 3 | 5 | 0.80 |
| G3 | 100 | 10 | 4 | 4 | 6 | 0.75 |
| G4 | 144 | 18 | 6 | 8 | 8 | 0.70 |
| Graph | State Dim. | Action Dim. | MLP |
|---|---|---|---|
| G1 | 13 | 40 | |
| G2 | 16 | 71 | |
| G3 | 18 | 108 | |
| G4 | 21 | 154 |
| Parameter | |||||
|---|---|---|---|---|---|
| Value | 200 | 50 | 25 | 10 | 3 |
| Hyperparameter | Value | Hyperparameter | Value |
|---|---|---|---|
| Learning rate | Batch size | 128 | |
| Exploration factor | Buffer size | ||
| Target network sync rate | Discount factor | ||
| Environment transitions per step | 10 | PER priority exponent [34] | |
| PER importance sampling [34] | Backpropagation optimizer | AdamW |
| Metric | G1 | G2 | G3 | G4 |
|---|---|---|---|---|
| Episodic reward (mean ± std) | 178.15 ± 20.85 | 162.82 ± 29.32 | 152.18 ± 41.59 | 140.75 ± 51.24 |
| Episode length (mean ± std) | 7.26 ± 2.19 | 11.21 ± 3.81 | 12.92 ± 4.29 | 15.84 ± 5.74 |
| Termination rate (%) | 100.00 | 100.00 | 100.00 | 100.00 |
| Infeasible actions rate (%) | 0.00 | 0.00 | 0.00 | 0.00 |
| Recharges rate (%) | 26.55 | 42.15 | 53.18 | 55.96 |
| Positive termination rate (16) (%) | 99.97 | 100.00 | 99.99 | 99.88 |
| Safety battery margin range * | [0.029, 0.041] | / | [0.051, 0.051] | [0.011, 0.119] |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Gemignani, G.; Pollini, L. Balancing Energy and Mission Time in UAV Site Servicing on Graph Maps Through Dynamic Battery-Threshold Double Deep Q-Learning. Electronics 2026, 15, 2984. https://doi.org/10.3390/electronics15142984
Gemignani G, Pollini L. Balancing Energy and Mission Time in UAV Site Servicing on Graph Maps Through Dynamic Battery-Threshold Double Deep Q-Learning. Electronics. 2026; 15(14):2984. https://doi.org/10.3390/electronics15142984
Chicago/Turabian StyleGemignani, Gabriele, and Lorenzo Pollini. 2026. "Balancing Energy and Mission Time in UAV Site Servicing on Graph Maps Through Dynamic Battery-Threshold Double Deep Q-Learning" Electronics 15, no. 14: 2984. https://doi.org/10.3390/electronics15142984
APA StyleGemignani, G., & Pollini, L. (2026). Balancing Energy and Mission Time in UAV Site Servicing on Graph Maps Through Dynamic Battery-Threshold Double Deep Q-Learning. Electronics, 15(14), 2984. https://doi.org/10.3390/electronics15142984

