A Safe Maritime Path Planning Fusion Algorithm for USVs Based on Reinforcement Learning A* and LSTM-Enhanced DWA
Abstract
1. Introduction
- The heuristic function of the A* algorithm is adaptively optimized using a reinforcement learning approach to improve path quality.
- A dynamic five-neighborhood search range is adopted to enhance the algorithm’s search speed.
- An LSTM network, integrated with KF, is employed to predict the trajectories of dynamic obstacles. Enabling the LSTM to provide stable, accurate, and early judgments of obstacle movements to mitigate collision risks.
- The evaluation function of the DWA is reformulated to strictly incorporate International Regulations for Preventing Collisions at Sea. LSTM enables the USVs to execute anticipatory and compliant avoidance maneuvers.
2. Environmental Map
3. Reinforcement A*
3.1. Conventional A*
- : the actual cost from the start node to the current node ;
- : the heuristic estimated cost from the current node to the goal node.
3.2. Q-Learning
- : the immediate reward function;
- : the environmental state at the current time step;
- : the action selected under state ;
- : the new state reached after executing action ;
- : the updated Q-value after taking action , in state , at time t + 1;
- : the discount factor, which balances the importance of immediate rewards and future returns;
- : the maximum Q-value among all possible actions in the next state , and signifies that the action belongs to the set of valid actions in the action space .
3.3. Improved A*
3.3.1. Dynamic Five-Neighborhood Search
3.3.2. Reinforcement Learning-Based Adaptive Weighting
- : the value of the enhanced heuristic function, which integrates the conventional heuristic, potential field guidance, and learned experience;
- : the base heuristic function;
- : the potential field value, generated based on the attraction from the target point to guide the path toward convergence with the goal;
- : the action-value function learned through Q-learning, accumulating historical decision-making experience;
- : the weight coefficient for the potential field, controlling the strength of its guidance;
- : the weight coefficient for the Q-value, adjusting the influence of learned experience.
- , , , and : the weight coefficients for each component, namely basic movement cost, progress-toward-goal reward, smoothness penalty, and potential field adjustment, all satisfying the normalization condition .
3.3.3. Path Optimization
3.3.4. Simulation Experiments
4. Improved DWA
4.1. Kinematic Model
- : minimum safe braking distance
- , : obstacle influence radius
- : maximum deceleration
4.2. LSTM
4.3. Adjustment of the Evaluation Function
- : the basic goal-alignment score;
- : the dynamic weight for distance-dependent;
- : the weight for heading.
- : a collection of parameters, where denotes the -th parameter and denotes the scores in EvalDB for heading, obstacle distance, and velocity, respectively; denotes the penalty score for proximity to the dynamic obstacle zone predicted by LSTM. This penalty appears when near dynamic obstacles, increases linearly within the zone, and is set to 0 outside the zone;
- : weight for each parameter;
- : number of motion candidates.
5. A*-DWA Fusion
5.1. Algorithmic Flowchart
5.2. Simulation
6. Discussion
Author Contributions
Funding
Data Availability Statement
Conflicts of Interest
Abbreviations
| USVs | Unmanned Surface Vehicles |
| LSTM | Long Short-Term Memory |
| DWA | Dynamic Window Approach |
| COLREGs | Convention on the Collision Regulations for the Prevention of Collisions at Sea |
| KF | Kalman Filter |
References
- Yao, Y.; Xiao, W.; Miao, P.; Chen, G.; Yang, H.; Chae, C.B.; Wong, K.K. UAV-Relay-Aided Secure Maritime Networks Coexisting with Satellite Networks: Robust Beamforming and Trajectory Optimization. IEEE Trans. Wirel. Commun. 2026, 25, 2342–2358. [Google Scholar] [CrossRef]
- Gu, C.; Deng, D.; Xiang, C. Risk-aware model predictive control framework for collision avoidance in unmanned surface vehicles. Ocean Eng. 2025, 342, 122954. [Google Scholar] [CrossRef]
- Sun, X.; Wang, G.; Fan, Y.; Mu, D. Collision avoidance control for unmanned surface vehicle with COLREGs compliance. Ocean Eng. 2023, 267, 113263. [Google Scholar] [CrossRef]
- Chen, Z.; Zhang, Y.; Zhang, Y.; Nie, Y.; Tang, J.; Zhu, S. A Hybrid Path Planning Algorithm for Unmanned Surface Vehicles in Complex Environment with Dynamic Obstacles. IEEE Access 2019, 7, 126439–126449. [Google Scholar] [CrossRef]
- Liu, L.; Wang, X.; Yang, X.; Liu, H.; Li, J.; Wang, P. Path planning techniques for mobile robots: Review and prospect. Expert Syst. Appl. 2023, 227, 120254. [Google Scholar] [CrossRef]
- Yang, Z.; Shu, J.; Jiang, J.; Han, W.; Wang, Y.; Zhao, L.; Bai, Y. Automated path-planning strategy for robotic inspection of underground utilities based on building information model. Comput. Aided Civ. Infrastruct. Eng. 2025, 40, 5554–5575. [Google Scholar] [CrossRef]
- Liu, Y.; Gao, X.; Wang, B.; Fan, J.; Li, Q.; Dai, W. A passage time–cost optimal A* algorithm for cross-country path planning. Int. J. Appl. Earth Obs. Geoinf. 2024, 130, 103907. [Google Scholar] [CrossRef]
- Lin, Z.; Wu, K.; Shen, R.; Yu, X.; Huang, S. An Efficient and Accurate A-Star Algorithm for Autonomous Vehicle Path Planning. IEEE Trans. Veh. Technol. 2024, 73, 9003–9008. [Google Scholar] [CrossRef]
- Huang, Y.; Li, Y.; Zhang, Z.; Sun, Q. A novel path planning approach for AUV based on improved whale optimization algorithm using segment learning and adaptive operator selection. Ocean Eng. 2023, 280, 114591. [Google Scholar] [CrossRef]
- Long, X.; Cai, W.; Yang, L.; Huang, H. Improved particle swarm optimization with reverse learning and neighbor adjustment for space surveillance network task scheduling. Swarm Evol. Comput. 2024, 85, 101482. [Google Scholar] [CrossRef]
- Yu, X.; Luo, W. Reinforcement learning-based multi-strategy cuckoo search algorithm for 3D UAV path planning. Expert Syst. Appl. 2023, 223, 119910. [Google Scholar] [CrossRef]
- Ren, Y.; Chang, Y.; Cui, Z.; Chang, X.; Yu, H.; Li, X.; Wang, Y. Is cooperative always better? Multi-Agent Reinforcement Learning with explicit neighborhood backtracking for network-wide traffic signal control. Transp. Res. Part C Emerg. Technol. 2025, 179, 105265. [Google Scholar] [CrossRef]
- Gu, J.; Wang, Y. A constrained reinforcement learning based approach for cooperative control of multi-UAV in dense obstacle environments. Sci. China Technol. Sci. 2026, 69, 1120601. [Google Scholar] [CrossRef]
- Chen, Y.; Bai, G.; Zhan, Y.; Hu, X.; Liu, J. Path Planning and Obstacle Avoiding of the USV Based on Improved ACO-APF Hybrid Algorithm with Adaptive Early-Warning. IEEE Access 2021, 9, 40728–40742. [Google Scholar] [CrossRef]
- Fox, D.; Burgard, W.; Thrun, S. The dynamic window approach to collision avoidance. IEEE Robot. Autom. Mag. 1997, 4, 23–33. [Google Scholar] [CrossRef]
- Yasuda, S.; Kumagai, T.; Yoshida, H. Safe and Efficient Dynamic Window Approach for Differential Mobile Robots with Stochastic Dynamics Using Deterministic Sampling. IEEE Robot. Autom. Lett. 2023, 8, 2614–2621. [Google Scholar] [CrossRef]
- Lee, D.H.; Lee, S.S.; Ahn, C.K.; Shi, P.; Lim, C.C. Finite Distribution Estimation-Based Dynamic Window Approach to Reliable Obstacle Avoidance of Mobile Robot. IEEE Trans. Ind. Electron. 2021, 68, 9998–10006. [Google Scholar] [CrossRef]
- Chen, H.; Lin, Z.; Chen, Z.; Jian, J.; Liu, C. Adaptive DWA algorithm with decision tree classifier for dynamic planning in USV navigation. Ocean Eng. 2025, 321, 120328. [Google Scholar] [CrossRef]
- Zhang, J.; Ling, H.; Tang, Z.; Song, W.; Lu, A. Path planning of USV in confined waters based on improved A* and DWA fusion algorithm. Ocean Eng. 2025, 322, 120475. [Google Scholar] [CrossRef]
- Guo, H.; Li, Y.; Wang, H.; Wang, C.; Zhang, J.; Wang, T.; Rong, L.; Wang, H.; Wang, Z.; Huo, Y.; et al. Path planning of greenhouse electric crawler tractor based on the improved A* and DWA algorithms. Comput. Electron. Agric. 2024, 227, 109596. [Google Scholar] [CrossRef]
- Vu, V.T.; Vu, M.C. 3D bathymetric images of the Truong Sa Archipelago (Spratly Islands). Reg. Stud. Mar. Sci. 2024, 73, 103509. [Google Scholar] [CrossRef]
- Hart, P.E.; Nilsson, N.J.; Raphael, B. A Formal Basis for the Heuristic Determination of Minimum Cost Paths. IEEE Trans. Syst. Sci. Cybern. 1968, 4, 100–107. [Google Scholar] [CrossRef]
- Dijkstra, E.W. A note on two problems in connexion with graphs. Numer. Math. 1959, 1, 269–271. [Google Scholar] [CrossRef]
- Zhu, Y.; Zhang, G.; Chu, R.; Xiao, H.; Yang, Y.; Wu, X. Research on escape route planning analysis in forest fire scenes based on the improved A* algorithm. Ecol. Indic. 2024, 166, 112355. [Google Scholar] [CrossRef]
- Fu, X.; Huang, Z.; Zhang, G.; Wang, W.; Wang, J. Research on path planning of mobile robots based on improved A* algorithm. PeerJ Comput. Sci. 2025, 11, e2691. [Google Scholar] [CrossRef]
- Hong, S.; Yang, H.; Zhao, T.; Ma, X. Epidemic spreading model of complex dynamical network with the heterogeneity of nodes. Int. J. Syst. Sci. 2016, 47, 2745–2752. [Google Scholar] [CrossRef]
- Wang, X.; Li, Y.; Gan, L.; Ma, Y. An enhanced ship collision-avoidance method using Q-learning-Optimized VO-PSO-Q-learning algorithm. Ocean Eng. 2026, 343, 123395. [Google Scholar] [CrossRef]
- Tan, X.; Han, L.; Gong, H.; Wu, Q. Biologically Inspired Complete Coverage Path Planning Algorithm Based on Q-Learning. Sensors 2023, 23, 4647. [Google Scholar] [CrossRef]
- Liang, J.; Tan, C.; Yan, L.; Zhou, J.; Yin, G.; Yang, K. Interaction-Aware Trajectory Prediction for Safe Motion Planning in Autonomous Driving: A Transformer-Transfer Learning Approach. IEEE Trans. Intell. Transp. Syst. 2025, 26, 17080–17095. [Google Scholar] [CrossRef]
- Wang, T.; Li, J.; Wu, H.; Li, C.; Snoussi, H.; Wu, Y. ResLNet: Deep residual LSTM network with longer input for action recognition. Front. Comput. Sci. 2022, 16, 166334. [Google Scholar] [CrossRef]
- Liu, R.W.; Liang, M.; Nie, J.; Lim, W.Y.B.; Zhang, Y.; Guizani, M. Deep Learning-Powered Vessel Trajectory Prediction for Improving Smart Traffic Services in Maritime Internet of Things. IEEE Trans. Netw. Sci. Eng. 2022, 9, 3080–3094. [Google Scholar] [CrossRef]
- Ahmad, E.; He, Y.; Luo, Z.; Lv, J. A Hybrid Long Short-Term Memory and Kalman Filter Model for Train Trajectory Prediction. IEEE Trans. Intell. Transp. Syst. 2024, 25, 7125–7139. [Google Scholar] [CrossRef]
- Namgung, H.; Kim, J.S. Collision Risk Inference System for Maritime Autonomous Surface Ships Using COLREGs Rules Compliant Collision Avoidance. IEEE Access 2021, 9, 7823–7835. [Google Scholar] [CrossRef]
- Namgung, H. Local Route Planning for Collision Avoidance of Maritime Autonomous Surface Ships in Compliance with COLREGs Rules. Sustainability 2022, 14, 198. [Google Scholar] [CrossRef]
























| Azimuth | Reserved Node | Discard Node | Unreachable Node |
|---|---|---|---|
| 337.5°~22.5° | {1,2,3,4,5} | {6,7,8} | {6,7,8} |
| 22.5°~67.5° | {1,2,3,5,8} | {4,6,7} | {4,6,7} |
| 67.5°~112.5° | {2,3,5,7,8} | {1,4,6} | {1,4,6} |
| 112.5°~155.5° | {3,5,6,7,8} | {1,2,4} | {1,2,4} |
| 155.5°~202.5° | {4,5,6,7,8} | {1,2,3} | {1,2,3} |
| 202.5°~247.5° | {1,4,6,7,8} | {2,3,5} | {2,3,5} |
| 247.5°~292.5° | {1,2,4,6,7} | {3,5,8} | {3,5,8} |
| 292.5°~337.5° | {1,2,3,4,6} | {5,7,8} | {5,7,8} |
| Reward Component | Mathematical Expression | Parameter Explanation | Physical Meaning |
|---|---|---|---|
| Basic Movement Cost | : Euclidean distance of movement | Penalize long paths to encourage shortest path selection | |
| Progress Reward | : progress coefficient | Reward effective progress toward the goal point | |
| Smoothness Penalty | : diagonal indicator function : smoothness coefficient | Apply a slight penalty to diagonal movements to enhance path smoothness | |
| Potential Field Adjustment | : potential field coefficient | Adjust the path based on potential field changes to strengthen goal-directed behavior |
| Algorithm | Path Length | Number of Turns | Smoothness | Expanded Nodes |
|---|---|---|---|---|
| Conventional A* (25 25) | 39.56 | 11 | 0.22 | 292 |
| Improved A* (25 25) | 38.16 | 8 | 0.11 | 276 |
| Conventional A* (50 50) | 78.43 | 27 | 0.31 | 1034 |
| Improved A* (50 50) | 74.41 | 20 | 0.11 | 871 |
| Conventional A* (75 75) | 114.95 | 29 | 0.16 | 1958 |
| Improved A* (75 75) | 111.89 | 20 | 0.06 | 1536 |
| Conventional A* (100 100) | 157.92 | 51 | 0.25 | 3368 |
| Improved A* (100 100) | 152.28 | 33 | 0.12 | 2852 |
| Algorithm | Path Length | Number of Turns | Smoothness | Expanded Nodes |
|---|---|---|---|---|
| Conventional A* (25 25) | 38.39 | 21 | 0.53 | 252 |
| Improved A* (25 25) | 35.68 | 11 | 0.22 | 211 |
| Conventional A* (50 50) | 78.67 | 24 | 0.24 | 789 |
| Improved A* (50 50) | 76.38 | 15 | 0.10 | 630 |
| Conventional A* (75 75) | 117.88 | 44 | 0.25 | 2347 |
| Improved A* (75 75) | 113.34 | 29 | 0.10 | 1876 |
| Conventional A* (100 100) | 155.58 | 51 | 0.27 | 3391 |
| Improved A* (100 100) | 148.30 | 37 | 0.11 | 2763 |
| Algorithm | Path Length | Number of Turns | Smoothness | Expanded Nodes |
|---|---|---|---|---|
| Conventional A* (25 25) | 37.80 | 16 | 0.31 | 218 |
| Improved A* (25 25) | 35.87 | 11 | 0.17 | 168 |
| Conventional A* (50 50) | 78.43 | 26 | 0.28 | 1021 |
| Improved A* (50 50) | 74.91 | 14 | 0.10 | 789 |
| Conventional A* (75 75) | 116.71 | 40 | 0.29 | 2018 |
| Improved A* (75 75) | 111.12 | 28 | 0.08 | 1329 |
| Conventional A* (100 100) | 161.10 | 53 | 0.31 | 4362 |
| Improved A* (100 100) | 155.62 | 38 | 0.17 | 3846 |
| Parameter | Value |
|---|---|
| Input Features | |
| Hidden Layers | 2 (128 units, 64 units) |
| Optimization Algorithm | Adam |
| Initial Learning Rate | 0.005 |
| Mini-Batch Size | 32 |
| Training/Validation Split | 80%/20% |
| Step | Mean Error | Max Error | Min Error | Standard Deviation |
|---|---|---|---|---|
| 10 | 0.156 | 0.378 | 0.009 | 0.107 |
| 20 | 0.786 | 1.903 | 0.428 | 0.545 |
| 30 | 2.0306 | 5.448 | 0.064 | 1.4357 |
| 40 | 4.027 | 12.083 | 0.274 | 2.795 |
| 50 | 6.828 | 22.112 | 0.24 | 4.827 |
| 60 | 10.511 | 35.097 | 0.46 | 7.658 |
| Path Length | Turning Points | Smoothness | Expand Nodes |
|---|---|---|---|
| 81.255 | 22 | 0.8764 | 1680 |
| 74.295 | 3 | 0.0139 | 1435 |
| 155.6 | 20 | 0.0734 | 3681 |
| 148.64 | 4 | 0.0033 | 3214 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Zhang, Z.; Wang, Q.; Wang, X.; Feng, M. A Safe Maritime Path Planning Fusion Algorithm for USVs Based on Reinforcement Learning A* and LSTM-Enhanced DWA. Sensors 2026, 26, 776. https://doi.org/10.3390/s26030776
Zhang Z, Wang Q, Wang X, Feng M. A Safe Maritime Path Planning Fusion Algorithm for USVs Based on Reinforcement Learning A* and LSTM-Enhanced DWA. Sensors. 2026; 26(3):776. https://doi.org/10.3390/s26030776
Chicago/Turabian StyleZhang, Zhenxing, Qiujie Wang, Xiaohui Wang, and Mingkun Feng. 2026. "A Safe Maritime Path Planning Fusion Algorithm for USVs Based on Reinforcement Learning A* and LSTM-Enhanced DWA" Sensors 26, no. 3: 776. https://doi.org/10.3390/s26030776
APA StyleZhang, Z., Wang, Q., Wang, X., & Feng, M. (2026). A Safe Maritime Path Planning Fusion Algorithm for USVs Based on Reinforcement Learning A* and LSTM-Enhanced DWA. Sensors, 26(3), 776. https://doi.org/10.3390/s26030776
