Optimal Path Planning for High-Altitude Low-Speed Aerostats Under Complex Constraints
Highlights
- Proposes a Markov decision process (MDP)-based path planning method that effectively handles uncertainties in forecasted wind fields for high-altitude low-speed aerostats.
- Enables optimal 2D and 3D global path planning and quantitative regional reachability analysis under complex constraints like limited propulsion and flight time.
- Provides a practical theoretical basis and strategic guidance for the rapid emergency deployment and operational planning of aerostats.
- Offers quantitative assessment for selecting optimal launch points and formulating flight strategies by evaluating the impact of wind resistance on reachability.
Abstract
1. Introduction
- An integrated MDP formulation that jointly addresses probabilistic wind fields and the complex physical constraints that are inherent to high-altitude low-speed aerostats.
- A 3D optimal path planning algorithm that enables strategic altitude-layer selection to exploit wind shear, coupled with a quantitative regional reachability analysis tool for mission pre-assessment.
- Comprehensive simulation studies validating the method’s efficacy in generating time-optimal 2D/3D paths under uncertainty and constraints, demonstrating its practical utility.
2. Related Work
2.1. Planning Characteristics and Physical Constraints of Aerostats
2.2. Evolution of Path Planning Methodologies
2.3. Markov Decision Processes in Stochastic Planning for Constrained Systems
2.4. Identified Research Gap and Our Contribution
3. Models
3.1. Overview of the MDP
3.2. Modeling Wind Uncertainty
3.2.1. Spatial Distribution of Wind Fields
3.2.2. Establishing the Uncertain Wind Fields
4. Approach
4.1. Research Object
4.2. Problem Setup
4.3. The Key Parameters of the MDP
4.3.1. Spatial Discretization and Transition
4.3.2. The Set of Actions
4.3.3. Transition Probabilities
4.3.4. Immediate Reward
4.3.5. Discount Factor
4.4. The Flow of Optimal Path Planning
5. Results
5.1. 2D Optimal Path Planning

5.1.1. Regional Reachability Analysis
5.1.2. The Optimal Path and Strategy Planning
5.2. 3D Optimal Path Planning
5.2.1. Regional Reachability Analysis
5.2.2. The Optimal Path and Strategy Planning
- Start Points 1 and 2: Both are west of the goal, but Start Point 1 is north and Start Point 2 is south. To better exploit the 3D wind field, their optimal paths exhibit distinct altitude variations: From Start Point 1, the altitude increases, then decreases, and increases again. From Start Point 2, the altitude first decreases, then remains stable, then increases.
- Start Points 3 and 4: Their paths show similar altitude changes because both are east of the goal and require locating optimal eastward wind layers to minimize flight time.
- Start Points 1 and 5: Both are southwest of the goal but at different distances. Their strategies for locating favorable wind layers differ, leading to disparate altitude adjustments.
- Start Point 4 achieves the shortest flight time for the simulated conditions, making it the most favorable among the evaluated deployment options in this scenario. 3D paths from the same start point consistently outperform 2D paths in terms of flight time, demonstrating the aerostat’s capability to leverage vertical wind layers for rapid deployment.
- Start Point 2 is closest to the goal, yet its expected flight time is not the shortest, underscoring the dominant influence of wind fields on path efficiency. The aerostat’s altitude adjustment capability significantly enhances deployment efficiency by enabling adaptive navigation through vertical wind layers.
5.3. Comparative Analysis and Performance Evaluation
5.3.1. Baseline Methods
- Fixed-altitude MDP (2D-MDP): This method utilizes the same MDP core (state transition, reward) as our proposed framework but is constrained to operate at a single, fixed altitude (the initial flight level of 19,400 m). The vertical action set is disabled. This baseline serves to isolate and quantify the performance gain that is solely attributable to the strategic 3D altitude selection capability of our full method.
- Greedy Heuristic algorithm: At each decision step, this reactive planner selects the combination of allowable horizontal and vertical actions that immediately minimizes the Euclidean distance to the goal position, given the current local wind estimate and propulsion limits. It performs no long-term value iteration or planning. This baseline represents a myopic, locally optimal strategy and highlights the value of global, foresighted planning under uncertainty that is offered by the MDP framework.
5.3.2. Evaluation Metrics
5.3.3. Comparative Results and Discussion
- Superiority in deployment efficiency and reliability: The proposed 3D-MDP method consistently outperforms both baselines across all start points and all metrics. It achieves the shortest average flight times (AFT) and the highest success rates (SR). For instance, at SP1, 3D-MDP reduces the AFT by 21.5% compared to the 2D-MDP and by 31.2% compared to the Greedy algorithm, while simultaneously improving SR by 7 and 22 percentage points, respectively. This conclusively demonstrates that integrating strategic altitude control into a global MDP planner is highly effective for minimizing deployment time and ensuring mission success under uncertainty.
- Enhanced robustness and predictability: A key, clear finding is the superior temporal robustness of the 3D-MDP method, evidenced by its consistently lowest standard deviation (σ) in flight time. This indicates that our method not only plans faster paths but also plans more consistent and predictable ones. The lower variance means the actual flight time is less susceptible to the specific instantiations of wind uncertainty, providing mission planners with more reliable time estimates—a critical practical advantage.
- Disentangling the contributions of planning foresight and altitude control:
- (1)
- Global planning vs. myopic reaction: The substantial performance gap between 3D-MDP and the Greedy Heuristic (e.g., ~50% longer flight times and ~20% lower success rates for Greedy at SP3) provides direct evidence that a foresighted, global optimization strategy is fundamentally superior to a reactive, locally optimal one in a stochastic, wind-dominated environment.
- (2)
- Value of 3D altitude optimization: The consistent and significant advantage of 3D-MDP over the 2D-MDP baseline directly quantifies the value of active altitude selection. By dynamically choosing flight levels to exploit favorable wind layers (as visualized in Section 4.2), the 3D-MDP method achieves faster and more reliable paths even than what is possible with an optimal planner restricted to a fixed altitude.
6. Discussion
- Effective handling of uncertain wind fields. The method effectively addresses the optimal path planning problem for high-altitude low-speed aerostats under uncertain wind fields, providing a more practically relevant environmental foundation for path planning.
- Quantitative assessment of wind resistance impact. It quantitatively evaluates the impact of factors such as wind resistance capability on regional reachability, offering guidance for selecting optimal starting points for aerostats.
- Comprehensive path and strategy generation. The method generates optimal paths and strategies for all positions within a given flight area, providing a theoretical basis for the practical deployment and actuation strategies of aerostats.
7. Conclusions
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
References
- Belmont, A.D.; Dartt, D.G.; Nastrom, G.D. Variations of stratospheric zonal winds, 20–65 km, 1961–1971. J. Appl. Meterology 2010, 14, 585–594. [Google Scholar] [CrossRef]
- Hong, Y.J. Aircraft Technology of Near Space; National Defense Industry Press: Beijing, China, 2012. [Google Scholar]
- Zhai, J.Q.; Yang, X.X.; Dend, X.L. Global path planning of stratospheric aerostat in uncertain wind field. Beijing Univ. Aeronaut. Astronaut. 2023, 49, 1116–1126. [Google Scholar] [CrossRef]
- Lin, K.; Ma, Y.P.; Zheng, Z.W.; Wu, Z. Height control of stratospheric aerostat based on secondary airbag. Beijing Univ. Aeronaut. Astronaut. 2022, 48, 762–770. [Google Scholar]
- Yang, Y.C.; Cao, H.S.; Zhao, R.; Zhu, R.C.; Song, L. Modeling and numerical simulation of constant-height flight by air-lifting gas mixing for aerostats. J. Natl. Univ. Def. Technol. 2023, 45, 196–204. [Google Scholar]
- Deng, X.; Yang, X.; Zhu, B.; Ma, Z.; Hou, Z. Simulation research and key technologies analysis of intelligent stratospheric aerostat Loon. ACTA Aeronaut. Astronaut. Sin. 2023, 44, 127412. [Google Scholar]
- Huang, D.J.; Jiang, C.F.; Han, K.L. 3D Path Planning Algorithm Based on Deep Reinforcement Learning. Comput. Eng. Appl. 2020, 56, 30–36. [Google Scholar]
- Khatib, O. Real-time obstacle avoidance for manipulators and mobile robots. Int. J. Robot. Res. 1986, 5, 90–98. [Google Scholar] [CrossRef]
- Wolf, M.T.; Blackmore, L.; Kuwata, Y. Probabilistic motion planning of balloons in strong, uncertain wind fields. In Proceedings of the 2010 IEEE International Conference on Robotics and Automation, Anchorage, AK, USA, 3–7 May 2010; IEEE: Piscataway, NJ, USA, 2010; pp. 1123–1129. [Google Scholar]
- Zheng, B.J.; Guo, X.; Wang, Y.F.; Ou, J.J.; Lou, W.J. Deep-Reinforcement-Learning-Based Path Planning Method for Stratospheric Airships in Spatiotemporally Complex Environments. IEEE Trans. Aerosp. Electron. Syst. 2025, 61, 17843–17857. [Google Scholar] [CrossRef]
- Lyu, Z.; Gao, Y.; Chen, J.; Du, H.; Xu, H.; Huang, K.; Kim, D.I. Empowering Intelligent Low-Altitude Economy With Large AI Model Deployment. IEEE Wirel. Commun. 2026. early access. [Google Scholar] [CrossRef]
- Zhao, W.Y.; He, T.R.; Chen, R.; Wei, T.H.; Liu, C.L. State-wise Safe Reinforcement Learning: A Survey. In Proceedings of the 32nd International Joint Conference on Artificial Intelligence (IJCAI), Macao, China, 19–25 August 2023; pp. 6814–6822. [Google Scholar]
- Yu, X.; Zhou, X.; Zhang, Y. Collision-free trajectory generation and tracking for UAVs using Markov decision process in a cluttered environment. J. Intell. Robot. Syst. 2019, 93, 17–32. [Google Scholar] [CrossRef]
- Sun, Y.; Wang, L.; Wu, J. A general overview of path planning methods for autonomous underwater vehicle. Ship Sci. Technol. 2020, 4, 1–7. [Google Scholar]
- Chen, L.; Duan, D.P.; Sun, D.S. Design of a multi-vectored thrust aerostat with a reconfigurable control system. Aerosp. Sci. Technol. 2016, 53, 95–102. [Google Scholar] [CrossRef]
- Zhao, M.; Xiao, C.; Zhou, P.F.; Duan, D.P. Dynamics modeling and simulation of a saucer-shaped stratospheric aerostat with an under-slung nacelle. J. Cent. South Univ. 2017, 24, 1288–1298. [Google Scholar] [CrossRef]
- Bordalba, R.; Ros, L.; Porta, J.M. Randomized Kinodynamic Planning for Constrained Systems. In Proceedings of the 2018 IEEE International Conference on Robotics and Automation (ICRA), Brisbane, QLD, Australia, 21–25 May 2018; IEEE: Piscataway, NJ, USA, 2018; pp. 7079–7086. [Google Scholar]
- Goldberg, D.E. Genetic Algorithms in Search, Optimization and Machine Learning; Addison-Wesley: Boston, MA, USA, 1989. [Google Scholar]
- Sutton, R.S.; Barto, A.G. Reinforcement Learning: An Introduction; MIT Press: Cambridge, MA, USA, 2018; pp. 10–87. [Google Scholar]
- Ruan, X.; Ren, D.; Zhu, X. Mobile robot navigation based on deep reinforcement learning. In Proceedings of the 2019 Chinese Control and Decision Conference (CCDC); IEEE: Piscataway, NJ, USA, 2019; pp. 6174–6178. [Google Scholar]
- Liu, K. Practical Markov Decision Making Process; University of Tsinghua Press: Beijing, China, 2012; pp. 9–12. [Google Scholar]














| Platform | Core Characteristics | Primary Env. Factor | Key Physical and Operational Constraints |
|---|---|---|---|
| High-Altitude Aerostat | Low dynamics, large size, wind-coupled | Wind field | Limited thrust [15], bounded altitude range [4,5], scarce on-board energy [6], coupled dynamics [16] |
| UAV | High maneuverability | Obstacles, other agents | Dynamics, sensor/comm. range, battery life |
| AUV | High maneuverability | Obstacles, other agents | Dynamics, sensor/comm. range, battery life |
| Max. Horizontal Actuation | Number of Reachable Cells | Regional Reachability |
|---|---|---|
| 56 | 12.7% | |
| 251 | 56.9% | |
| 441 | 100% |
| Start Point | Max. Horizontal Actuation | Numbers of Optimal Path Notes | Flight Time (h) |
|---|---|---|---|
| Start Point 1 | 9 | 24.9 | |
| Start Point 2 | 9 | 21.33 | |
| 9 | 15.79 | ||
| 9 | 16.08 | ||
| Start Point 3 | 9 | 17.84 | |
| 9 | 16.45 | ||
| Start Point 4 | 7 | 11.62 | |
| Start Point 5 | 5 | 20.91 |
| Start Point | Number of Altitude Controls | Numbers of Path Notes | Flight Time (h) |
|---|---|---|---|
| Start Point 1 | 7 | 9 | 19.88 |
| Start Point 2 | 5 | 9 | 9.29 |
| Start Point 3 | 2 | 9 | 9.51 |
| Start Point 4 | 3 | 7 | 7.78 |
| Start Point 5 | 2 | 5 | 12.62 |
| Start Point | Method | Avg. Flight Time (h) ↓ | Std. Dev. (h) ↓ | Success Rate (%) ↑ |
|---|---|---|---|---|
| Start Point 1 | 3D-MDP (Proposed) | 19.88 | 2.1 | 92 |
| 2D-MDP (Fixed-Altitude) | 25.34 | 3.3 | 85 | |
| Greedy Heuristic | 28.91 | 5.8 | 70 | |
| Start Point 2 | 3D-MDP (Proposed) | 9.29 | 1.2 | 98 |
| 2D-MDP (Fixed-Altitude) | 12.15 | 2.8 | 90 | |
| Greedy Heuristic | 15.40 | 4.5 | 75 | |
| Start Point 3 | 3D-MDP (Proposed) | 9.51 | 1.5 | 97 |
| 2D-MDP (Fixed-Altitude) | 11.92 | 2.9 | 88 | |
| Greedy Heuristic | 17.84 | 5.2 | 72 | |
| Start Point 4 | 3D-MDP (Proposed) | 7.78 | 0.9 | 99 |
| 2D-MDP (Fixed-Altitude) | 10.23 | 2.0 | 93 | |
| Greedy Heuristic | 13.57 | 3.8 | 80 | |
| Start Point 5 | 3D-MDP (Proposed) | 12.62 | 1.8 | 95 |
| 2D-MDP (Fixed-Altitude) | 16.45 | 3.2 | 87 | |
| Greedy Heuristic | 20.91 | 4.7 | 78 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Zhai, J.; Wu, X.; Zhang, Y.; Ye, H.; Wang, Z.; Yin, P. Optimal Path Planning for High-Altitude Low-Speed Aerostats Under Complex Constraints. Drones 2026, 10, 128. https://doi.org/10.3390/drones10020128
Zhai J, Wu X, Zhang Y, Ye H, Wang Z, Yin P. Optimal Path Planning for High-Altitude Low-Speed Aerostats Under Complex Constraints. Drones. 2026; 10(2):128. https://doi.org/10.3390/drones10020128
Chicago/Turabian StyleZhai, Jiaqi, Xiaolong Wu, Yongdong Zhang, Hu Ye, Ziwei Wang, and Peng Yin. 2026. "Optimal Path Planning for High-Altitude Low-Speed Aerostats Under Complex Constraints" Drones 10, no. 2: 128. https://doi.org/10.3390/drones10020128
APA StyleZhai, J., Wu, X., Zhang, Y., Ye, H., Wang, Z., & Yin, P. (2026). Optimal Path Planning for High-Altitude Low-Speed Aerostats Under Complex Constraints. Drones, 10(2), 128. https://doi.org/10.3390/drones10020128
