Compensating Environmental Disturbances in Maritime Path Following Using Deep Reinforcement Learning
Abstract
1. Introduction
1.1. Background
1.2. Related Works
1.3. Contribution
- Systematic evaluation of disturbance perception: We conduct a comparative analysis between path following agents operating with implicit environmental awareness and those utilizing explicit wind vector observations. This investigation quantifies the specific performance gains—in terms of stability and precision—achieved by augmenting the observation space with disturbance data.
- Pareto analysis of reward functions: By systematically varying reward weights, we identify and visualize the trade-off (Pareto frontier) between navigational velocity and CTE. This provides insights into the sensitivity of the policy to reward shaping.
- Statistical robustness assessment: To ensure generalization across diverse operational conditions, we employ a domain randomization approach involving stochastic wind vectors and variable path geometries. Performance is assessed using comprehensive statistical metrics—including box plots and scatter plots—over test sets.
2. Methodology
2.1. Framework Architecture
2.2. Scenario Overview and Definition of the Reinforcement Learning Problems
- PF1—Path following without disturbance observation:The objective in this scenario is for the USV to accurately track a predefined reference trajectory without the observation of disturbances. Here, the agent can observe its own motion states as well as its position and orientation relative to relevant path points. Further details are described in the following subsections.
- PF2—Path following with disturbance observation:In addition to the observation variables introduced in PF1, information about wind speed and direction is provided to the agent during training and testing.
2.2.1. Action Space
2.2.2. Observation Space
2.2.3. Reward Function
- Cross-track error (CTE) penalty to penalize the distance from the path:
- Reward for velocity towards the path goal (VTPG) with the actual boat velocity and the distance vector to the look-ahead waypoint. This reward is intended to encourage the agent to travel target-oriented:
- A constant time penalty per step to penalize slow traveling:
2.2.4. Scenario Generation
3. Results and Discussion
3.1. Scenario PF1—Path Following Without Disturbance Observation
3.2. Scenario PF2—Path Following with Disturbance Observation
4. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
Abbreviations
| A3C | Asynchronous Advantage Actor–Critic |
| AI | Artificial Intelligence |
| AMV | Autonomous Marine Vehicle |
| APFs | Artificial Potential Fields |
| CTE | Cross-Track Error |
| DDPG | Deep Deterministic Policy Gradient |
| DQN | Deep Q-Network |
| DRL | Deep Reinforcement Learning |
| EKF | Extended Kalman Filter |
| GNC | Guidance Navigation and Control |
| LOS | Line of Sight |
| LSTM | Long Short-Term Memory |
| MPC | Model Predictive Controller |
| PF | Path Following |
| PPO | Proximal Policy Optimization |
| RL | Reinforcement Learning |
| ROS | Robot Operating System |
| RRT | Rapidly Exploring Random Tree |
| SAC | Soft Actor–Critic |
| SLAM | Simultaneous Localization and Mapping |
| TD3 | Twin-Delayed Deep Deterministic Policy Gradient |
| USV | Unmanned Surface Vehicle |
| VRX | Virtual RobotX |
| VTPG | Velocity Toward Path Goal |
| WAMV | Wave Adaptive Modular Vehicle |
Appendix A

| PF1 | C1_1_1 | C1_2_1 | C1_3_1 | C1_1_2 | C1_2_2 | C1_3_2 | C1_1_3 | C1_2_3 | C1_3_3 | C1_1_1_D1 | C1_1_1_D2 |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Mean of theor. mean velocity/m/s | 2.50 | 0.49 | 0.38 | 2.30 | 1.62 | 0.87 | 1.88 | 2.09 | 2.22 | 2.30 | 2.13 |
| SD of theor. mean velocity/m/s | 0.15 | 0.05 | 0.05 | 0.13 | 0.06 | 0.07 | 0.29 | 0.13 | 0.11 | 0.43 | 0.50 |
| Mean of mean cross-track error/m | 0.42 | 0.32 | 0.34 | 1.22 | 0.43 | 0.35 | 4.44 | 1.15 | 0.88 | 0.90 | 0.89 |
| SD of mean cross-track error/m | 0.08 | 0.02 | 0.01 | 0.26 | 0.03 | 0.03 | 1.44 | 0.29 | 0.19 | 1.57 | 2.06 |
| PF2 | C2_1_1 | C2_2_1 | C2_3_1 | C2_1_2 | C2_2_2 | C2_3_2 | C2_1_3 | C2_2_3 | C2_3_3 | C2_1_1_CL | |
| Mean of theor. mean velocity/m/s | 2.17 | 0.94 | 0.79 | 2.10 | 2.22 | 1.41 | 1.96 | 1.84 | 2.06 | 2.20 | |
| SD of theor. mean velocity/m/s | 0.37 | 0.25 | 0.31 | 0.44 | 0.40 | 0.24 | 0.66 | 0.51 | 0.42 | 0.49 | |
| Mean of mean cross-track error/m | 0.73 | 1.01 | 0.79 | 1.45 | 1.14 | 0.70 | 3.01 | 3.25 | 1.71 | 0.92 | |
| SD of mean cross-track error/m | 0.77 | 2.60 | 1.26 | 1.52 | 1.66 | 1.34 | 1.57 | 1.42 | 0.77 | 0.82 |
References
- Liu, Z.; Zhang, Y.; Yu, X.; Yuan, C. Unmanned surface vehicles: An overview of developments and challenges. Annu. Rev. Control 2016, 41, 71–93. [Google Scholar] [CrossRef] [Scilit]
- Krautwig, B.; Wans, D.; Li, L.; Temmen, T.; Koch, L.; Eisenbarth, M.; Andert, J. Navigating the Trade-Offs: A Quantitative Analysis of Reinforcement Learning Reward Functions for Autonomous Maritime Collision Avoidance. J. Mar. Sci. Eng. 2025, 13, 2233. [Google Scholar] [CrossRef] [Scilit]
- Siciliano, B.; Khatib, O. (Eds.) Springer Handbook of Robotics; Springer: Berlin/Heidelberg, Germany, 2008. [Google Scholar] [CrossRef] [Scilit]
- Chen, C.S.; Lin, C.J.; Lai, C.C.; Lin, S.Y. Velocity Estimation and Cost Map Generation for Dynamic Obstacle Avoidance of ROS Based AMR. Machines 2022, 10, 501. [Google Scholar] [CrossRef] [Scilit]
- Ferguson, D.; Likhachev, M. Efficiently Using Cost Maps For Planning Complex Maneuvers. In Proceedings of International Conference on Robotics and Automation Workshop on Planning with Cost Maps; IEEE: Piscataway, NJ, USA, 2008. [Google Scholar]
- Racinskis, P.; Arents, J.; Greitans, M. Constructing Maps for Autonomous Robotics: An Introductory Conceptual Overview. Electronics 2023, 12, 2925. [Google Scholar] [CrossRef] [Scilit]
- Thrun, S. Robotic mapping: A survey. In Exploring Artificial Intelligence in the New Millennium; Morgan Kaufmann Publishers Inc.: San Francisco, CA, USA, 2003; pp. 1–35. [Google Scholar]
- Yi, C.; Jeong, S.; Cho, J. Map Representation for Robots. Smart Comput. Rev. 2012, 2, 18–27. [Google Scholar] [CrossRef] [Scilit]
- Macenski, S.; Moore, T.; Lu, D.V.; Merzlyakov, A.; Ferguson, M. From the desks of ROS maintainers: A survey of modern & capable mobile robotics algorithms in the robot operating system 2. Robot. Auton. Syst. 2023, 168, 104493. [Google Scholar] [CrossRef] [Scilit]
- Sanchez, M.; Morales, J.; Martínez, J.; Fernández-Lozano, J.; Garcia, A. Automatically Annotated Dataset of a Ground Mobile Robot in Natural Environments via Gazebo Simulations. Sensors 2022, 22, 5599. [Google Scholar] [CrossRef] [Scilit]
- Souissi, O.; Benatitallah, R.; Duvivier, D.; Artiba, A.; Belanger, N.; Feyzeau, P. Path planning: A 2013 survey. In Proceedings of the 2013 International Conference on Industrial Engineering and Systems Management (IESM), Rabat, Morocco, 28–30 October 2013; Curran Associates, Inc.: Red Hook, NY, USA, 2013; pp. 1–8. [Google Scholar]
- LaValle, S.M. Planning Algorithms, reprinted. ed.; Cambridge University Press: New York, NY, USA, 2014; p. cop. 2006. [Google Scholar]
- Karaman, S.; Frazzoli, E. Sampling-based algorithms for optimal motion planning. Int. J. Robot. Res. 2011, 30, 846–894. [Google Scholar] [CrossRef] [Scilit]
- Loe, Ø. Collision Avoidance Concepts for Marine Surface Craft. Trondheim 2007, 19, 111. [Google Scholar]
- Alessandretti, A.; Aguiar, A.P.; Jones, C.N. Trajectory-tracking and path-following controllers for constrained underactuated vehicles using Model Predictive Control. In Proceedings of the 2013 European Control Conference (ECC), Zurich, Switzerland, 17–19 July 2013; IEEE: Piscataway, NJ, USA, 2013; pp. 1371–1376. [Google Scholar] [CrossRef] [Scilit]
- Thyri, E.H.; Breivik, M. Collision avoidance for ASVs through trajectory planning: MPC with COLREGs-compliant nonlinear constraints. Model. Identif. Control A Nor. Res. Bull. 2022, 43, 55–77. [Google Scholar] [CrossRef] [Scilit]
- Kamel, M.S.; Stastny, T.; Alexis, K.; Siegwart, R. Model Predictive Control for Trajectory Tracking of Unmanned Aerial Vehicles Using Robot Operating System. In Robot Operating System (ROS); Springer: Berlin/Heidelberg, Germany, 2017. [Google Scholar] [CrossRef] [Scilit]
- Wallace, M.T.; Streetman, B.; Lessard, L. Model Predictive Planning: Trajectory Planning in Obstruction-Dense Environments for Low-Agility Aircraft. arXiv 2024, arXiv:2309.16024v2. [Google Scholar] [CrossRef] [Scilit]
- Li, Z.; Sun, J. Disturbance Compensating Model Predictive Control With Application to Ship Heading Control. IEEE Trans. Control. Syst. Technol. 2012, 20, 257–265. [Google Scholar] [CrossRef] [Scilit]
- Moser, M.M.; Huang, M.; Abel, D. Model Predictive Control for Safe Path Following in Narrow Inland Waterways for Rudder Steered Inland Vessels*. In Proceedings of the 2023 European Control Conference (ECC), Bucharest, Romania, 13–16 June 2023; IEEE: Piscataway, NJ, USA, 2023; pp. 1–6. [Google Scholar] [CrossRef] [Scilit]
- Wang, Y.; Tong, H.; Fu, M. Line-of-sight guidance law for path following of amphibious hovercrafts with big and time-varying sideslip compensation. Ocean Eng. 2019, 172, 531–540. [Google Scholar] [CrossRef] [Scilit]
- Fossen, T.I.; Breivik, M.; Skjetne, R. Line-of-sight path following of underactuated marine craft. IFAC Proc. Vol. 2003, 36, 211–216. [Google Scholar] [CrossRef] [Scilit]
- Xia, J.; Zhu, X.; Liu, Z.; Luo, Y.; Wu, Z.; Wu, Q. Research on Collision Avoidance Algorithm of Unmanned Surface Vehicle Based on Deep Reinforcement Learning. IEEE Sens. J. 2023, 23, 11262–11273. [Google Scholar] [CrossRef] [Scilit]
- Johansen, T.A.; Fossen, T.I.; Berge, S.P. Constrained Nonlinear Control Allocation with Singularity Avoidance Using Sequential Quadratic Programming. IEEE Trans. Control Syst. Technol. 2004, 12, 211–216. [Google Scholar] [CrossRef] [Scilit]
- Lorenz, U. Reinforcement Learning: Aktuelle Ansätze Verstehen–Mit Beispielen in Java und Greenfoot, 2nd ed.; Springer: Berlin/Heidelberg, Germany, 2024. [Google Scholar] [CrossRef] [Scilit]
- Chen, Z.; Bao, T.; Zhang, B.; Wu, T.; Chu, X.; Zhou, Z. Deep Reinforcement Learning Methods for USV Control: A Review. In Proceedings of the 2023 China Automation Congress (CAC), Chongqing, China, 17–19 November 2023; IEEE: Piscataway, NJ, USA, 2023; pp. 1526–1531. [Google Scholar] [CrossRef] [Scilit]
- Li, S.; Ji, Y.; Liu, J.; Bai, Z.; Hu, J.; Gao, Q. Global Path Planning of Unmanned Surface Vehicles Based on Deep Q Network. In Proceedings of the 2024 43rd Chinese Control Conference (CCC), Kunming, China, 28–31 July 2024; IEEE: Piscataway, NJ, USA, 2024; pp. 3827–3832. [Google Scholar] [CrossRef] [Scilit]
- Zhao, J.; Wang, P.; Li, B.; Bai, C. A DDPG-Based USV Path-Planning Algorithm. Appl. Sci. 2023, 13, 10567. [Google Scholar] [CrossRef] [Scilit]
- Zhong, W.; Li, H.; Meng, Y.; Yang, X.; Feng, Y.; Ye, H.; Liu, W. USV path following controller based on DDPG with composite state-space and dynamic reward function. Ocean Eng. 2022, 266, 112449. [Google Scholar] [CrossRef] [Scilit]
- Deraj, R.; Kumar, R.S.; Alam, M.S.; Somayajula, A. Deep reinforcement learning based controller for ship navigation. Ocean Eng. 2023, 273, 113937. [Google Scholar] [CrossRef] [Scilit]
- Qu, X.; Jiang, Y.; Zhang, R.; Long, F. A Deep Reinforcement Learning-Based Path-Following Control Scheme for an Uncertain Under-Actuated Autonomous Marine Vehicle. J. Mar. Sci. Eng. 2023, 11, 1762. [Google Scholar] [CrossRef] [Scilit]
- Dong, Z.; Chen, L.; Huang, Y.; Chen, P.; Mou, J. Model-based Reinforcement Learning for Ship Path Following with Disturbances. IFAC-PapersOnLine 2024, 58, 247–252. [Google Scholar] [CrossRef] [Scilit]
- Raffin, A.; Hill, A.; Gleave, A.; Kanervisto, A.; Ernestus, M.; Dormann, N. Stable-Baselines3: Reliable Reinforcement Learning Implementations. J. Mach. Learn. Res. 2021, 22, 1–8. [Google Scholar]
- Brockman, G.; Cheung, V.; Pettersson, L.; Schneider, J.; Schulman, J.; Tang, J.; Zaremba, W. Openai gym. arXiv 2016, arXiv:1606.01540. [Google Scholar] [CrossRef] [Scilit]
- Bingham, B.; Agüero, C.; McCarrin, M.; Klamo, J.; Malia, J.; Allen, K.; Lum, T.; Rawson, M.; Waqar, R. Toward Maritime Robotic Simulation in Gazebo. In Proceedings of the OCEANS 2019 MTS/IEEE SEATTLE, Seattle, WA, USA, 27–31 October 2019; IEEE: Piscataway, NJ, USA, 2019; pp. 1–10. [Google Scholar] [CrossRef] [Scilit]
- Quigley, M.; Gerkey, B.; Conley, K.; Faust, J.; Foote, T.; Leibs, J.; Berger, E.; Wheeler, B.; Ng, A. ROS: An open-source Robot Operating System. In Proceedings of the ICRA Workshop on Open Source Software; IEEE: Piscataway, NJ, USA, 2009; Volume 3, p. 5. [Google Scholar]
- Daniel, K.; Nash, A.; Koenig, S.; Felner, A. Theta*: Any-Angle Path Planning on Grids. J. Artif. Intell. Res. 2010, 39, 533–579. [Google Scholar] [CrossRef] [Scilit]
- Abouelazm, A.; Michel, J.; Zöllner, J. A Review of Reward Functions for Reinforcement Learning in the context of Autonomous Driving. In Proceedings of the 2024 IEEE Intelligent Vehicles Symposium (IV), Jeju Island, Republic of Korea, 2–5 June 2024; IEEE: Piscataway, NJ, USA, 2024. [Google Scholar] [CrossRef] [Scilit]
- Florensa, C.; Held, D.; Wulfmeier, M.; Zhang, M.; Abbeel, P. Reverse Curriculum Generation for Reinforcement Learning. arXiv 2018, arXiv:1707.05300. [Google Scholar] [CrossRef] [Scilit]
- Cheng, M.; Yao, J.; Ren, Q. A Model Predictive Control Approach for USV Autonomous Cruising via Disturbance Learning. In Proceedings of the 2024 IEEE 18th International Conference on Control & Automation (ICCA), Reykjavík, Iceland, 18–21 June 2024; IEEE: Piscataway, NJ, USA, 2024; pp. 988–993. [Google Scholar] [CrossRef] [Scilit]
- DNV. Class Guideline: Autonomous and Remotely Operated Ships (DNV-CG-0264); Technical Report; DNV GL: Houston, TX, USA, 2021. [Google Scholar]










| Configuration Encoding | Parametrization |
|---|---|
| CS_1_Y | |
| CS_2_Y | |
| CS_3_Y | |
| CS_X_1 | |
| CS_X_2 | |
| CS_X_3 |
| Name | Goal Reached/% | Step Limit Reached/% | Exited Workspace/% |
|---|---|---|---|
C1_1_1 ![]() | 100.0 | 0.0 | 0.0 |
C1_2_1 ![]() | 21.0 | 79.0 | 0.0 |
C1_3_1 ![]() | 17.0 | 83.0 | 0.0 |
C1_1_2 ![]() | 100.0 | 0.0 | 0.0 |
C1_2_2 ![]() | 100.0 | 0.0 | 0.0 |
C1_3_2 ![]() | 45.0 | 55.0 | 0.0 |
C1_1_3 ![]() | 10.5 | 89.5 | 0.0 |
C1_2_3 ![]() | 99.5 | 0.5 | 0.0 |
C1_3_3 ![]() | 100.0 | 0.0 | 0.0 |
C1_1_1_D1 ![]() | 98.0 | 2.0 | 0.0 |
C1_1_1_D2 ![]() | 96.0 | 3.5 | 0.5 |
| Name | Goal Reached/% | Step Limit Reached/% | Exited Workspace/% |
|---|---|---|---|
C2_1_1 ![]() | 98.0 | 1.5 | 0.5 |
C2_2_1 ![]() | 44.5 | 55.5 | 0.0 |
C2_3_1 ![]() | 28.5 | 71.0 | 0.5 |
C2_1_2 ![]() | 99.0 | 1.0 | 0.0 |
C2_2_2 ![]() | 99.0 | 1.0 | 0.0 |
C2_3_2 ![]() | 82.0 | 13.5 | 4.5 |
C2_1_3 ![]() | 64.0 | 36.0 | 0.0 |
C2_2_3 ![]() | 72.0 | 28.0 | 0.0 |
C2_3_3 ![]() | 98.5 | 1.5 | 0.0 |
C1_1_1_D1 ![]() | 98.0 | 2.0 | 0.0 |
C1_1_1_D2 ![]() | 96.0 | 3.5 | 0.5 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Krautwig, B.; Wans, D.; Temmen, T.; Brinkmann, T.; Lee, S.-Y.; Kim, D.; Andert, J. Compensating Environmental Disturbances in Maritime Path Following Using Deep Reinforcement Learning. J. Mar. Sci. Eng. 2026, 14, 327. https://doi.org/10.3390/jmse14040327
Krautwig B, Wans D, Temmen T, Brinkmann T, Lee S-Y, Kim D, Andert J. Compensating Environmental Disturbances in Maritime Path Following Using Deep Reinforcement Learning. Journal of Marine Science and Engineering. 2026; 14(4):327. https://doi.org/10.3390/jmse14040327
Chicago/Turabian StyleKrautwig, Björn, Dominik Wans, Till Temmen, Tobias Brinkmann, Sung-Yong Lee, Daehyuk Kim, and Jakob Andert. 2026. "Compensating Environmental Disturbances in Maritime Path Following Using Deep Reinforcement Learning" Journal of Marine Science and Engineering 14, no. 4: 327. https://doi.org/10.3390/jmse14040327
APA StyleKrautwig, B., Wans, D., Temmen, T., Brinkmann, T., Lee, S.-Y., Kim, D., & Andert, J. (2026). Compensating Environmental Disturbances in Maritime Path Following Using Deep Reinforcement Learning. Journal of Marine Science and Engineering, 14(4), 327. https://doi.org/10.3390/jmse14040327












